Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,489–7,512 of 20,843 skills
Evaluates a model's ability to generate accurate, clinically correct chest X-ray reports while aligning its visual focus with radiologist gaze patterns. It probes both natural language generation quality and interpretable attention prediction to ensure diagnostic reasoning is visually grounded. Use when the user wants to benchmark on FG-CXR, or asks about evaluating this task. Reports C, F1_ex, fwIoU.
Evaluates large language models' harmlessness across three dimensions: factuality (resistance to misinformation and counterfactuals), fairness (prediction disparity across demographic groups), and toxicity (generation of harmful content when prompted with jailbreaks). Use when the user wants to benchmark on FFT, or asks about evaluating this task. Reports accuracy.
Evaluates in-processing group fairness methods by measuring the trade-off between model utility (acc) and various fairness metrics across multiple datasets and hyperparameter settings. Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports acc.
Evaluates the ability of genetic programming systems to discover exact symbolic mathematical expressions from numerical data points. It probes search efficiency, robustness to domain constraints (e.g., NaNs), and the impact of different fitness functions on expression discovery. Use when the user wants to benchmark on Feynman dataset, or asks about evaluating this task. Reports typically_solved.
This benchmark evaluates the robustness of few-shot image classifiers to spurious class-attribute correlations. It constructs evaluation tasks where support sets contain images with engineered spurious attributes, and query sets lack these attributes or contain attributes from other classes, then measures classification accuracy compared to standard random task sampling. Use when the user wants to benchmark on miniImageNet, tieredImageNet, CUB-200, or asks about evaluating this task. Reports ...
Evaluates the in-domain and out-of-domain generalization capabilities of large language models adapted via few-shot fine-tuning versus in-context learning across standard natural language inference and paraphrase detection benchmarks. Use when the user wants to benchmark on MNLI, RTE, QQP, or asks about evaluating this task. Reports accuracy.
Evaluates few-shot learning methods for decoding brain activation maps from fMRI data. It probes a model's ability to classify cognitive tasks using very limited labeled examples (1 or 5 per class) by randomly sampling novel classes and instances. Use when the user wants to benchmark on IBC dataset, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models on few-shot learning capabilities across nine diverse tasks. It probes the models' ability to leverage in-context demonstrations (0, 4, or 8 shots) and chain-of-thought reasoning under controlled retrieval settings, measuring performance relative to zero-shot baselines. Use when the user wants to benchmark on FewMMBench, or asks about evaluating this task. Reports accuracy.
Evaluates few-shot joint language understanding by measuring a model's ability to simultaneously predict dialogue intent and extract slots from query sentences using only a small support set from unseen domains. Use when the user wants to benchmark on FewJoint, or asks about evaluating this task. Reports Sentence Accuracy.
Evaluates Chinese NLP models on few-shot learning across nine tasks, including single-sentence classification, sentence-pair classification, and machine reading comprehension. It tests the ability of pre-trained language models and few-shot prompting/fine-tuning methods to generalize with limited labeled data. Use when the user wants to benchmark on FewCLUE, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify video actions from a small number of labeled examples per class. It probes few-shot learning capabilities by measuring accuracy across thousands of randomly sampled episodes on standard video benchmarks. Use when the user wants to benchmark on Kinetics, Something-Something V2, complete-Kinetics, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of a generative model to produce high-fidelity time series data under extreme data scarcity (few-shot fine-tuning). It probes cross-domain generalization and robustness to varying sequence lengths and channel dimensions by comparing generated samples against real test data. Use when the user wants to benchmark on ECG200, ETTh2, ETTm1, ETTm2, ILI, Weather, Synthetic sine wave, or asks about evaluating this task. Reports contextFID (c-FID).
Evaluates a model's ability to perform semantic segmentation on 3D point clouds using only a few labeled examples per category. It probes how well learned point embeddings generalize to unseen shapes when supervision is extremely limited, testing both few-shot (few labeled shapes) and few-point (few labeled points per shape) scenarios. Use when the user wants to benchmark on ShapeNet segmentation dataset, or asks about evaluating this task. Reports mIOU.
Evaluates few-shot image classification capability using a label-free, similarity-based approach. It probes how well self-supervised visual representations can classify novel classes with only a few key images per class, without any training or test labels. Use when the user wants to benchmark on miniImageNet, CIFAR-100FS, FC100, or asks about evaluating this task. Reports accuracy.
Evaluates the zero-shot and few-shot capabilities of large language models across a diverse suite of NLP, reasoning, and commonsense benchmarks. It measures how efficiently a model scales with compute and whether additional training objectives unlock emergent reasoning abilities. Use when the user wants to benchmark on GPT-3 suite, BigBench Emergent Suite, Commonsense QA benchmarks, Closed-book QA benchmarks, or asks about evaluating this task. Reports average score.
Evaluates parameter-efficient fine-tuning methods for few-shot natural language generation from structured data (knowledge graphs and semantic representations) to text. It probes the model's ability to adapt to data-scarce regimes while preserving generation fluency and factual alignment with the source structure. Use when the user wants to benchmark on WebNLG 2020, E2E, DART, or asks about evaluating this task. Reports BLEU.
Evaluates few-shot named entity recognition on document images, measuring how well models identify entity spans with limited training examples. It also probes model robustness to geometric image manipulations like rotation, scaling, and shifting during inference. Use when the user wants to benchmark on FUNSD, CORD, or asks about evaluating this task. Reports word-level F-1 score.
Evaluates few-shot classification performance of meta-learning algorithms on standard image datasets. It specifically probes robustness to distribution shift or difficulty by measuring accuracy on dynamically identified 'hard' episodes versus average episodic performance. Use when the user wants to benchmark on CIFAR-FS, mini-ImageNet, tieredImageNet, or asks about evaluating this task. Reports episodic accuracy.
Evaluates few-shot image classification models on semantically coherent versus uniformly sampled tasks. It probes the model's ability to generalize from limited support examples to query images across varying class coarseness and scale (5-way vs 100-way). Use when the user wants to benchmark on tieredImageNet, Danish Fungi 2020, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates the few-shot classification capability of deep learning models on histopathology images. It probes how well models generalize from extremely limited labeled examples (1, 5, or 10 per class) across disjoint medical imaging domains with varying resolutions and class distributions. Use when the user wants to benchmark on Komura & Ishikawa (2021), CRC-TP, NCT, LC25000, or asks about evaluating this task. Reports Accuracy(%).
Evaluates few-shot entity recognition in document images by measuring how well a model identifies and classifies entity spans using minimal labeled examples. It probes the model's ability to jointly leverage textual semantics and spatial layout information under low-data regimes. Use when the user wants to benchmark on FUNSD, CORD-Lv1, or asks about evaluating this task. Reports word-level F-1 score.
Evaluates few-shot classification performance under standard and explicit data-shift conditions. It probes a model's ability to generalize from a small number of labeled support samples to query samples across different image domains and distribution shifts. Use when the user wants to benchmark on miniImageNet, Office-Home, Easy-Office-Home, Hard-Office-Home, or asks about evaluating this task. Reports classification accuracy.
This benchmark evaluates few-shot audio classification capability across diverse acoustic domains including environmental sounds, musical instruments, bird species, and speaker recognition. It measures how well a model generalizes to novel classes with only a few labeled examples per class using a prototypical network framework. Use when the user wants to benchmark on ESC-50, FSD2018, NSynth, BirdCLEF 2020, VoxCeleb1, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to recognize human actions in video clips using only a few labeled examples per class. It probes the model's capacity to leverage motion dynamics and semantic cues for robust classification under data-scarce conditions. Use when the user wants to benchmark on Something-Something, Kinetics, UCF101, HMDB51, FineGym, or asks about evaluating this task. Reports average few-shot accuracy.