Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,201–7,224 of 20,839 skills
Evaluates the ability of pre-trained language models to perform a diverse set of natural language understanding tasks, including sentence classification, paraphrase detection, semantic similarity, natural language inference, and extractive question answering. Use when the user wants to benchmark on GLUE, SQuAD 1.1, SQuAD 2.0, or asks about evaluating this task. Reports Task-specific metrics (Accuracy, F1, Spearman, Matthews).
Evaluates the downstream generalization and data efficiency of pre-trained language models by fine-tuning them on standard natural language understanding and instruction-following benchmarks. Use when the user wants to benchmark on GLUE, SuperNatural-Instructions (SNI), or asks about evaluating this task. Reports GLUE.
This protocol evaluates whether language models memorize benchmark surface features or demonstrate true semantic robustness. It measures performance degradation when inputs are paraphrased, lexically/syntactically perturbed, or adversarially rewritten, contrasting deterministic greedy decoding with stochastic distributional evaluation. Use when the user wants to benchmark on GLUE (MNLI, QQP, QNLI, SST-2), or asks about evaluating this task. Reports GLUE robustness ratio.
Evaluates the generalization and stability of finetuned pretrained language models on low-resource NLP tasks. It probes how well models adapt to sentiment classification, natural language inference, paraphrasing, similarity assessment, and linguistic acceptability when trained on severely limited data (300–1000 examples). Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports averaged evaluation metrics.
Evaluates the quantization performance of pre-trained language models (both discriminative BERT-style and generative GPT-style) across standard NLP classification, regression, and language modeling tasks under data-free zero-shot quantization settings. Use when the user wants to benchmark on GLUE, WikiText2, Penn Treebank (PTB), WikiText103, or asks about evaluating this task. Reports GLUE Avg..
Evaluates few-shot text classification performance across multiple natural language understanding tasks. It probes a model's ability to generalize from extremely limited labeled examples (16 per class) by generating synthetic training data and fine-tuning a classifier. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports Average performance.
Evaluates the transfer learning capability of pre-trained language models across a diverse suite of natural language understanding tasks, including sentence classification, semantic textual similarity, and natural language inference. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports GLUE Average.
Evaluates the ability of efficient fine-tuning and structured sparsity methods to maintain performance across a diverse suite of natural language understanding tasks. It probes task-specific classification accuracy and correlation metrics under both full-data and limited-data regimes, as well as pre-training perplexity on a large-scale corpus. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports GLUE average score.
This benchmark evaluates session-based news recommendation systems by predicting the next article a user will click based on their recent interaction history. It probes a model's ability to capture short-term user intent, handle temporal dynamics, and leverage both content and contextual features in a streaming environment. Use when the user wants to benchmark on Globo.com, or asks about evaluating this task. Reports HR@5.
Evaluates the audio fidelity, speaker diversity, and transcript alignment of English multi-speaker TTS datasets. It also benchmarks how well zero-shot speaker-adaptive TTS models trained on these corpora generalize to unseen speakers and global accents. Use when the user wants to benchmark on GLOBE, VCTK, Common Voice, LibriTTS, LibriTTS-R, or asks about evaluating this task. Reports NMOS.
Evaluates AI music generation models for global and cultural bias by measuring how well generated tracks match reference tracks across different world regions and genres. It probes the models' out-of-distribution capabilities and tendency to default to mainstream styles over authentic regional ones. Use when the user wants to benchmark on GlobalDISCO, or asks about evaluating this task. Reports FAD.
This benchmark probes physical commonsense reasoning by testing whether models can distinguish correct from incorrect solutions to everyday physical tasks. It specifically evaluates cultural and linguistic grounding by using items constructed natively in 116 language varieties, avoiding translation artifacts that often skew multilingual evaluations. Use when the user wants to benchmark on Global PIQA, or asks about evaluating this task. Reports accuracy.
This evaluation probes a model's ability to recover and predict sparse latent fitness functions from observed sequence-fitness data distorted by global epistasis. It compares data efficiency and predictive accuracy under complete versus incomplete (subsampled) data regimes, highlighting the robustness of contrastive losses over mean-squared error. Use when the user wants to benchmark on NK model synthetic data, FLIP benchmark, or asks about evaluating this task. Reports Spearman correlation.
Evaluates a Retrieval-Augmented Generation system's ability to ground responses in retrieved legal documents and directly address jurisdiction-specific AI regulation queries. It tests the system's performance on both single-entity and multi-jurisdictional comparison tasks using automatic LLM-based scoring. Use when the user wants to benchmark on Global AI Regulation Test Queries, or asks about evaluating this task. Reports faithfulness.
Evaluates text-to-speech systems on pronunciation accuracy, speaker similarity, and emotional expressiveness across standard Chinese/English benchmarks and challenging internal datasets. It also assesses vocoder quality using objective and subjective audio metrics to measure overall synthesis fidelity. Use when the user wants to benchmark on Seed-TTS-eval, Libri & Chinese Dialects, or asks about evaluating this task. Reports CER.
This evaluation protocol assesses the zero-shot and few-shot capabilities of large bilingual language models across diverse English and Chinese benchmarks. It probes language modeling, multi-choice question answering, reasoning, commonsense, and cross-lingual transfer abilities. Use when the user wants to benchmark on LAMBADA, Pile, MMLU, BIG-bench-lite, CLUE, FewCLUE, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to predict glioma IDH mutation status (mutant vs. wild-type) by integrating multi-modal MRI data, including anatomical sequences, tumor geometry, and reconstructed brain networks. It probes the model's capacity for cross-modal feature alignment and patient-level binary classification under data-scarce conditions. Use when the user wants to benchmark on TCIA & In-house Glioma Cohort, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to predict glioblastoma molecular subtypes using paired MRI and histopathology data. It probes the model's capacity to fuse heterogeneous imaging modalities, preserve topological structures, and handle missing data scenarios. Use when the user wants to benchmark on Ivy GAP + Cancer Stem Cells ISH Survey, or asks about evaluating this task. Reports accuracy.
Evaluates a weakly supervised segmentation model's ability to accurately delineate glandular structures in colorectal histopathology images using sparse annotations. It also probes cross-domain generalization across different institutional cohorts with varying staining protocols and scanner characteristics. Use when the user wants to benchmark on GlaS, or asks about evaluating this task. Reports mIoU.
Evaluates semi-supervised gland segmentation performance on histopathology images under limited labeled data (5% or 10%). It probes the model's ability to disentangle stain color and tissue structure while maintaining boundary precision and shape preservation with minimal annotations. Use when the user wants to benchmark on GlaS, CRAG, or asks about evaluating this task. Reports Dice.
Evaluates visual-linguistic relation understanding by requiring models to classify image-text matching, ground matched objects with bounding boxes, and identify mismatched relations from candidates. It specifically probes data efficiency and length generalization capabilities in out-of-distribution settings. Use when the user wants to benchmark on GITM-MR, or asks about evaluating this task. Reports Match%.
Evaluates LLMs' ability to extract and verify user interests from interaction histories, focusing on factual grounding, specificity, and strict instruction following across heterogeneous engagement types. Use when the user wants to benchmark on Unspecified real-world engagement datasets, or asks about evaluating this task. Reports IG.
Evaluates automatic speech recognition (ASR) models on low-resource languages (Thai, Indonesian, Vietnamese) to measure transcription accuracy against reference texts. It probes the model's ability to handle domain-shifted audio and varying linguistic structures using character-level or word-level error metrics. Use when the user wants to benchmark on GigaSpeech 2, Common Voice 17.0, FLEURS, or asks about evaluating this task. Reports CER/WER.
Evaluates automatic speech recognition (ASR) systems on a large-scale, multi-domain English corpus containing both read and spontaneous speech. It benchmarks transcription accuracy across tiered training subsets and professionally re-transcribed evaluation sets using word error rate. Use when the user wants to benchmark on GigaSpeech, or asks about evaluating this task. Reports WER.