Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,633–4,656 of 23,504 skills
Evaluates the ability of multimodal large language models (MLLMs) to perform in-context learning (ICL) in medical domains. It probes how effectively models leverage provided image-question-answer demonstrations to answer new clinical queries, while also measuring robustness to irrelevant examples, recency bias, and the gap between automated and expert clinical judgment. Use when the user wants to benchmark on SMMILE, SMMILE++, or asks about evaluating this task. Reports LLM-as-a-Judge.
Evaluates 3D medical image segmentation models on mesoscopic small vessel extraction from ultra-high-resolution (7T) Time-of-Flight Magnetic Resonance Angiography. It probes a model's ability to handle high noise, poor vessel-background contrast, and domain shifts across different MRI acquisition sources. Use when the user wants to benchmark on SMILE-UHURA Challenge Dataset, or asks about evaluating this task. Reports Dice coefficient (DICE).
Evaluates data efficiency and response generation quality in goal-oriented dialogue systems under few-shot conditions. It measures how well a model can generate contextually appropriate and entity-accurate responses using only a small fraction of in-domain dialogue data. Use when the user wants to benchmark on SMD, or asks about evaluating this task. Reports BLEU.
Evaluates neural marked spatio-temporal and temporal point process models on predicting the next event's time, location, and mark, while quantifying prediction uncertainty. It probes the model's ability to generate well-calibrated confidence regions for continuous variables and accurate probability estimates for discrete marks. Use when the user wants to benchmark on Earthquake, Crime, Football, StackOverflow, Retweet, MIMIC-II, Financial Transactions, or asks about evaluating this task. Repo...
Evaluates multi-agent reinforcement learning algorithms in a simulated urban driving environment, measuring scenario completion, episode duration, human-like driving fidelity, and traffic rule compliance. Use when the user wants to benchmark on SMARTS (NeurIPS Competition Track-1), or asks about evaluating this task. Reports Completion.
Evaluates large vision-and-language models on mathematical reasoning tasks from the Math Kangaroo Olympiad, testing their ability to solve grade-appropriate (K-12) multiple-choice problems that may require joint text and image interpretation. Use when the user wants to benchmark on SMART-840, or asks about evaluating this task. Reports accuracy.
Evaluates text-to-query systems across relational, document, and graph database models using four query languages (SQL, MQL, Cypher, SPARQL). It probes the ability of models to translate natural language medical questions into correct, executable database queries using standardized SNOMED-CT aligned synthetic patient data. Use when the user wants to benchmark on SM3-Text-to-Query, or asks about evaluating this task. Reports correctness.
Evaluates spoken language understanding capabilities across multiple tasks including intent classification, slot filling, emotion recognition, and dialogue act classification. It probes a model's ability to map raw audio inputs to semantic labels, test robustness to noise and low-resource settings, and assess the utility of pretrained ASR/NLU feature extractors in end-to-end speech processing pipelines. Use when the user wants to benchmark on FSC (Fluent Speech Commands), Snips, SLURP, IEMOCA...
This benchmark evaluates how different pose estimation models impact the quality of sign language translation. It probes the robustness of pose estimators to occlusion, temporal instability, and missing hand keypoints, measuring their downstream effect on translation metrics. Use when the user wants to benchmark on RWTH-PHOENIX-Weather 2014, Signsuisse, or asks about evaluating this task. Reports BLEU.
Evaluates monolingual, cross-lingual, and multilingual NLP models on a human- and machine-translated Slovene version of the SuperGLUE benchmark. It probes how well models handle morphological and grammatical challenges in low-resource language processing, and compares translation quality impacts on downstream task performance. Use when the user wants to benchmark on Slovene SuperGLUE, or asks about evaluating this task. Reports Avg.
Evaluates the sequential recommendation capability of a distilled small language model against traditional and LLM-based baselines. It measures ranking accuracy on user-item interaction histories and assesses computational efficiency (training/inference time and parameter count). Use when the user wants to benchmark on Amazon18, or asks about evaluating this task. Reports MRR.
Evaluates an agentic framework's ability to translate natural language scientific queries into executable Python pipelines for automated histopathology analysis on whole-slide images, requiring multi-step computational reasoning rather than simple knowledge recall or diagnosis. Use when the user wants to benchmark on SlideQuest, or asks about evaluating this task. Reports task_success_rate.
Evaluates multi-page visual document understanding and question answering. It probes a model's ability to retrieve relevant slides, perform spatial and layout reasoning, and accurately extract numeric values or generate lexical matches for open-ended answers. Use when the user wants to benchmark on SlideVQA, TechSlides, FinSlides, or asks about evaluating this task. Reports Num, Overall.
Evaluates the accuracy of simulation-based inference engines in recovering the true posterior distribution of model parameters given synthetic observational data. It probes the ability of implicit likelihood methods to handle complex, multimodal posteriors and varying simulation budgets. Use when the user wants to benchmark on SLCP (Simple Likelihood Complex Posterior), or asks about evaluating this task. Reports C2ST.
Evaluates the speech language model's performance across multiple benchmarks including linguistic acceptability, story completion, audio generation quality, and cross-domain text generation perplexity. Use when the user wants to benchmark on sBLIMP, StoryCloze, People Speech, or asks about evaluating this task. Reports MOSnet.
This evaluation protocol assesses end-to-end spoken dialogue models across three core capabilities: instruction understanding, logical reasoning, and open-ended oral conversation. It measures both the semantic quality of the generated responses and the acoustic fidelity of the synthesized speech. Use when the user wants to benchmark on Repeat, Summary, StoralEval, TruthfulEval, MLC, AlpacaEval, CommonEval, WildchatEval, or asks about evaluating this task. Reports ChatGPT Score.
Evaluates the ability of generative models to separate individual musical instrument stems from a mixed audio track. It probes how well the model captures inter-source dependencies and reconstructs clean waveforms for Bass, Drums, Guitar, and Piano. Use when the user wants to benchmark on Slakh2100, or asks about evaluating this task. Reports SI-SDR_i.
Evaluates medical visual question answering capabilities by testing a model's ability to reason over radiology images (CT/MRI/X-ray) to answer vision-only and knowledge-based questions in English and Chinese. It probes multimodal fusion, semantic segmentation utilization, and external medical knowledge graph integration for clinical reasoning. Use when the user wants to benchmark on SLAKE, or asks about evaluating this task. Reports Accuracy.
Evaluates bilingual foundation models on general knowledge, Chinese domain-specific reasoning, mathematical problem-solving, and language modeling capabilities using standardized benchmarks and custom held-out text corpora. Use when the user wants to benchmark on MMLU, CEVAL, CMMLU, GSM8K, Custom Chinese LM Testset, or asks about evaluating this task. Reports 5-shot accuracy.
Evaluates semantic segmentation models trained on synthetic aerial imagery for their ability to generalize to real-world UAV datasets and adapt to varying environmental conditions like weather, time of day, and camera viewpoint. Use when the user wants to benchmark on SKYSCENES, UAVid, AEROSCAPES, ICG DRONE, SYNDrone, or asks about evaluating this task. Reports mIoU.
Evaluates object detection and counting capabilities in densely packed scenes, specifically testing a model's ability to localize and count tightly overlapping items without false positives from standard non-maximum suppression. Use when the user wants to benchmark on SKU-110K, CARPK, PUCPR+, or asks about evaluating this task. Reports AP.
This benchmark evaluates continual learning methods for generating reusable procedural skills in LLM agents. It probes the quality of generated skills, their alignment with execution trajectories, and the ultimate task-solving accuracy and efficiency of a fixed solving agent. Use when the user wants to benchmark on SkillLearnBench, or asks about evaluating this task. Reports Acc..
Evaluates autonomous agents' ability to discover, patch, and evolve reusable skills over time in a sequential, lifelong learning setting. It probes whether models can consolidate successful execution traces into a compact, repairable skill library rather than merely accumulating fragmented task-specific traces. Use when the user wants to benchmark on SkillFlow, or asks about evaluating this task. Reports task completion rate (%comp.).
This evaluation probes a diffusion model's ability to generate pixel-level sketches that align with detailed text prompts while maintaining stylistic abstraction and human-like drawing characteristics. It measures both perceptual image quality and fine-grained text-to-image semantic alignment across multiple complementary metrics. Use when the user wants to benchmark on SketchDUO, or asks about evaluating this task. Reports TIFAScore.