Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,993–8,016 of 20,853 skills
Evaluates link recommendation models on social networks by measuring how well they predict future connections while respecting individual users' diversity preferences across profile dimensions. Use when the user wants to benchmark on Large-scale social network datasets, or asks about evaluating this task. Reports F1 Score.
Evaluates vision-language models' ability to perform fine-grained low-level visual perception by identifying both the specific type of image distortion and its severity level from a single image. It probes whether models rely on direct perceptual pattern matching or struggle with subtle severity discrimination. Use when the user wants to benchmark on DistortBench, or asks about evaluating this task. Reports Acc..
Probes a model's ability to perform dynamic aspect-based summarization on disordered, non-sequential texts where sentences from multiple sources are shuffled. It tests whether the model can cluster fragmented content by underlying topics or aspects and generate precise, coherent summaries without relying on original sentence order. Use when the user wants to benchmark on D-CnnDM, D-WikiHow, or asks about evaluating this task. Reports Human Evaluation (Coherence, Consistency, Fluency, Relevanc...
Evaluates pre-trained language models' ability to reason about diseases by mapping symptoms, treatments, tests, procedures, and terminology to disease names. It isolates medical reasoning types and uses adversarial negative examples to prevent knowledge leakage. Use when the user wants to benchmark on DisKnE, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to predict user check-in behavior (CTR) in location-based recommendation by disentangling sequential and geographical influences. It probes how well the model handles data sparsity and cold-start scenarios using real-world POI interaction logs. Use when the user wants to benchmark on Foursquare Tokyo, Foursquare New York, Meituan, or asks about evaluating this task. Reports AUC.
Evaluates machine translation systems on discourse-level coherence and terminological precision in expert domains. It probes the model's ability to maintain long-form text consistency and handle domain-specific language beyond sentence-level translation. Use when the user wants to benchmark on DiscoX, or asks about evaluating this task. Reports Metric-S.
Evaluates whether text-to-speech systems can correctly realize discourse-dependent word-level stress based on contrasting contexts. It probes the model's ability to adapt prosodic emphasis dynamically rather than relying on fixed sentence-internal stress patterns. Use when the user wants to benchmark on CAST, or asks about evaluating this task. Reports Pair-Correct.
This benchmark evaluates the ability of discrete speech models to unsupervisedly discover phoneme inventories and capture phonemic contrasts. It measures how well predicted units align with gold phonemes through phonetic similarity, recognition error, and temporal segmentation accuracy across multiple languages. Use when the user wants to benchmark on discoPhon, or asks about evaluating this task. Reports PNMI.
Evaluates speech enhancement models in extremely low SNR conditions by measuring noise suppression, speech quality preservation, and intelligibility using both objective metrics and subjective listening tests. Use when the user wants to benchmark on Low-SNR Dataset, VB-DMD Dataset, DNS Non-Reverb Test Dataset, DNS Real Recordings, or asks about evaluating this task. Reports PESQ.
Evaluates whether sentence representations capture discourse-aware semantics by testing performance on tasks involving sentence ordering, discourse relations, and coherence across multiple domains. Use when the user wants to benchmark on DiscoEval, or asks about evaluating this task. Reports accuracy.
This protocol evaluates how well a metamodel can predict the benchmark performance of unseen models using a highly condensed subset of test samples. It probes the trade-off between evaluation cost reduction and the fidelity of accuracy estimation and model ranking preservation across language and vision benchmarks. Use when the user wants to benchmark on MMLU, HellaSwag, Winogrande, ARC, ImageNet-1k, or asks about evaluating this task. Reports MAE, Spearman rank correlation.
Evaluates the ability of fine-tuned LLMs to generate clinically accurate, complete, and readable discharge summaries for cardiac patients from raw medical records. It probes domain-specific medical summarization, factual consistency, and adherence to clinical documentation standards. Use when the user wants to benchmark on Cardiology Clinical Dataset, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates image copy detection systems under realistic, adversarial conditions. It probes a model's ability to match transformed query images against a large reference database while resisting geometric, color, overlay, and deepfake manipulations. The setup emphasizes scalability and robustness in a high-false-positive-rate, needle-in-haystack search regime. Use when the user wants to benchmark on DISC21, or asks about evaluating this task. Reports micro Average Precision.
Evaluates large vision-language models' ability to assess disaster damage from remote sensing imagery. It probes capabilities in multi-sensor (optical/SAR) understanding, object counting, relational reasoning, and generating professional disaster response reports. Use when the user wants to benchmark on DisasterM3, or asks about evaluating this task. Reports accuracy (%).
Evaluates whether the model refuses or safely handles requests for disallowed content (e.g., hate speech, illicit advice, personal data, self-harm, sexual/exploitative material) across standard and production-like multiturn conversations. Use when the user wants to benchmark on Production Benchmarks, or asks about evaluating this task. Reports not_unsafe.
Evaluates distant speech recognition and noise robustness by measuring word error rate on close-talk and far-field recordings of natural dinner conversations. It tests the model's ability to handle uncontrolled acoustic conditions, background music, and spatially diverse microphone placements. Use when the user wants to benchmark on DiPCo, or asks about evaluating this task. Reports WER.
Evaluates machine translation systems on four discourse phenomena: anaphora resolution, lexical consistency, coherence/readability, and discourse connectives. It probes whether context-aware models can maintain discourse-level quality and consistency across different language pairs beyond standard n-gram overlap. Use when the user wants to benchmark on DiP Benchmark, or asks about evaluating this task. Reports BLEU.
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
Evaluates the cross-domain generalization and scaling behavior of a natural-image pre-trained vision transformer (DINOv3) across diverse medical imaging modalities, including 2D/3D classification and segmentation tasks. Use when the user wants to benchmark on NIH-14, RSNA-Pneumonia, Camelyon16, Camelyon17, BCNB, Kvasir-Capsule, AutoLaparo, EndoVis18, EDD 2020, CT-RATE, Medical Segmentation Decathlon (MSD), CREMI, AC3/4, AutoPET-II, HECKTOR 2022, or asks about evaluating this task. Reports AUC...
Evaluates the cross-task generalizability of the DINOv2 vision foundation model on medical image analysis tasks, specifically disease classification and organ segmentation across X-ray, CT, and MRI modalities. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, SARS-CoV-2, Brain Tumor, Montgomery County (MC), AMOS, MSD Heart, MSD Hipp, MSD Spleen, or asks about evaluating this task. Reports AUROC.
This benchmark evaluates Vision-Language Models on fine-grained visual discrimination, nutritional quantification from images, and complex food-related visual question answering. It probes the models' ability to fuse multi-view imagery, perform volumetric reasoning, and avoid parametric knowledge biases when identifying dishes and estimating macronutrients. Use when the user wants to benchmark on DiningBench, or asks about evaluating this task. Reports Accuracy, MAPE.
This benchmark evaluates the robustness of network intrusion detection models against distribution shifts between different network environments. It specifically probes cross-domain generalization by training on one NetFlow dataset and testing on another, measuring how well domain-invariant feature extraction mitigates performance degradation when facing unseen attack distributions. Use when the user wants to benchmark on NFv2-UNSW-NB15, NFv2-CIC-2018, or asks about evaluating this task. Repo...
Evaluates a model's ability to generate syntactically and semantically correct SQL queries from natural language questions across diverse database schemas. It probes schema linking, handling of complex query structures (joins, nested subqueries, aggregations), and the capacity for iterative self-correction when initial generations fail. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates multilingual and multidomain aspect-based sentiment analysis by predicting continuous valence-arousal (VA) scores alongside aspect, opinion, and category extraction. It probes a model's ability to perform fine-grained dimensional sentiment regression and structured information extraction across diverse languages and domains. Use when the user wants to benchmark on DimABSA, or asks about evaluating this task. Reports RMSE_VA, cF1.