Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,073–9,096 of 21,245 skills
This benchmark evaluates how large language models exhibit social bias across different tasks, bias types, and social groups. It systematically measures bias through direct classification tasks and indirect text generation tasks, using standardized metrics to enable cross-dataset fairness comparisons. Use when the user wants to benchmark on CEB, or asks about evaluating this task. Reports Micro-F1.
Evaluates the quality, sparsity, and computational efficiency of counterfactual explanation methods for recommender systems. It probes how effectively explanations can alter recommendation rankings, how interpretable the generated explanations are, and the cost of generating them across different input formats and perturbation scopes. Use when the user wants to benchmark on Recommender system interaction datasets, or asks about evaluating this task. Reports POS-P@K.
Evaluates clinical information retrieval systems on their ability to rank relevant medical documents for clinical decision support queries. It probes the effectiveness of query and document processing techniques such as negation detection, concept extraction, and pseudorelevance feedback in a standardized biomedical search setting. Use when the user wants to benchmark on TREC CDS'16, or asks about evaluating this task. Reports infNDCG.
Evaluates a multi-modal deep learning framework's ability to predict drug-target binding interactions across standard, cross-domain, and cold-start scenarios. It probes the model's capacity to integrate textual, structural, and functional biological features for robust binary classification under distribution shifts and unseen entities. Use when the user wants to benchmark on BindingDB, Davis, or asks about evaluating this task. Reports AUROC.
This benchmark probes a vision-language model's ability to maintain visual fidelity when explicit visual evidence conflicts with strong commonsense priors. It specifically measures whether models override entrenched prior-driven expectations with counterfactual image-grounded claims, isolating hallucination from generic perception errors. Use when the user wants to benchmark on CDH-Bench, or asks about evaluating this task. Reports CFAD.
Cross-domain few-shot video action recognition. It probes a model's ability to adapt to new video domains using a large source dataset and unlabeled target videos, with evaluation restricted to a 5-way 5-shot setting where only five target classes are tested with five labeled support examples each. Use when the user wants to benchmark on Kinetics-100, Kinetics-400, UCF101, HMDB51, Something-SomethingV2, Diving48, RareAct, or asks about evaluating this task. Reports 5-way 5-shot accuracy.
Evaluates cross-domain facial expression recognition (CD-FER) models by measuring how well they transfer learned features from a labeled source dataset to an unlabeled target dataset. It probes the model's ability to learn domain-invariant representations and adapt to distribution shifts across different facial expression datasets. Use when the user wants to benchmark on RAF-DB, AFE, CK+, JAFFE, SFEW2.0, FER2013, ExpW, or asks about evaluating this task. Reports accuracy.
Probes multimodal LLMs' ability to answer binary questions about traffic videos while maintaining logical consistency across counterfactual video-question pairs. It specifically diagnoses failure modes like positive omission, negative hallucination, and mutual-exclusivity violations by enforcing a strict quadruple-level decision rule. Use when the user wants to benchmark on CCTVBench, or asks about evaluating this task. Reports QuadAcc.
Evaluates the quality of a large-scale monolingual web corpus by measuring downstream performance on standard linguistic analogy tasks and a cross-lingual natural language inference benchmark. Use when the user wants to benchmark on CCNet, XNLI, or asks about evaluating this task. Reports XNLI.
Evaluates the quality of mined parallel sentence pairs by training machine translation systems on them and measuring translation performance. It probes the effectiveness of global, margin-based bitext mining in a multilingual embedding space. The benchmark measures how well the mined data generalizes across different language families and scripts. Use when the user wants to benchmark on CCMatrix, or asks about evaluating this task. Reports BLEU.
Evaluates intent classification robustness under noisy ASR conditions by measuring how well a model aligns noisy transcripts with clean references and preserves semantic consistency. Use when the user wants to benchmark on SLURP, Timers, FSC, SNIPS, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to identify five types of clinical named entities (diseases, symptoms, exams, treatments, body parts) in Chinese medical texts. It specifically probes character-level sequence labeling performance and the impact of integrating external dictionary features. Use when the user wants to benchmark on CCKS-2017 Task 2, or asks about evaluating this task. Reports F1-score.
Evaluates the effectiveness of a high-quality Chinese pre-training corpus by training a 0.5B language model from scratch and measuring its zero-shot generalization across standard English and Chinese knowledge benchmarks. The protocol also includes a comparative analysis of data quality classifiers using macro F1-score on a held-out test set. Use when the user wants to benchmark on Standard Benchmarks (ARC-C, ARC-E, HellaSwag, Winograd, MMLU, OpenbookQA, PIQA, SIQA, CEval, CMMLU), or asks abo...
This benchmark evaluates the factual accuracy and consistency of multimodal large language models (MLLMs) when answering questions in text or speech modalities across eight languages. It specifically probes cross-lingual transfer capabilities and cross-modal alignment by measuring how well models maintain factual correctness when switching between languages or between text and audio inputs. Use when the user wants to benchmark on CCFQA, or asks about evaluating this task. Reports F1 score.
Evaluates speech restoration models on realistic, multi-stage degradations including acoustic noise/reverberation, codec compression artifacts, and secondary processing artifacts from upstream enhancement models. Probes the trade-off between signal fidelity/intelligibility and perceptual quality while measuring computational efficiency. Use when the user wants to benchmark on CCF AATC 2025 Test Set, or asks about evaluating this task. Reports WAcc.
Evaluates continual semi-supervised learning on crowd counting by measuring how well a model adapts to evolving unlabeled data streams across sequential sessions. Use when the user wants to benchmark on Continual Crowd Counting (CCC), or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Evaluates whether large language models can effectively learn and utilize spatial coordinate information versus categorical/compositional data for property prediction. It quantifies the systematic performance degradation (the 'Coordinate-Category Cliff') when tasks require geometric reasoning rather than simple type matching, and tests whether scaling model size or dataset volume mitigates this deficit. Use when the user wants to benchmark on Synthetic coordinate-category datasets, Materials ...
Evaluates LLMs on cognitive behavioral therapy (CBT) assistance across basic knowledge recall and cognitive model understanding. The benchmark probes multiple-choice knowledge acquisition and multi-label classification of cognitive distortions and core beliefs to measure therapeutic applicability and fine-grained clinical reasoning. Use when the user wants to benchmark on CBT-QA, CBT-CD, CBT-PC, CBT-FC, or asks about evaluating this task. Reports Accuracy, F1.
Evaluates an ANN-based adaptive classifier's ability to classify encrypted network traffic into known categories such as malware families, operating systems, browsers, and applications. It specifically probes the model's capacity to dynamically adapt to new or out-of-distribution classes without retraining, while measuring any performance degradation on existing classes compared to traditional baselines. Use when the user wants to benchmark on BOA, MTA, or asks about evaluating this task. Rep...
Evaluates whether Concept Bottleneck Models learn semantically meaningful concept representations from input images under varying annotation granularity and concept correlation structures. Measures how well the model predicts intermediate concepts and downstream tasks compared to standard neural networks. Use when the user wants to benchmark on Playing cards, CheXpert, or asks about evaluating this task. Reports concept accuracy.
Evaluates the computational performance and resource efficiency of edge computing platforms for connected and autonomous vehicle workloads. It probes how well hardware handles real-time vision, deep learning, and diagnostic tasks under varying resource constraints. Use when the user wants to benchmark on CAVBench, or asks about evaluating this task. Reports Matching Factor (MF).
Probes the ability of causal representation learning (CRL) models to recover ground-truth latent variables from high-fidelity visual simulations. It evaluates both component-wise and block-wise identifiability under realistic conditions where theoretical assumptions may be violated. Use when the user wants to benchmark on CausalVerse, or asks about evaluating this task. Reports Mean Correlation Coefficient (MCC).
This benchmark evaluates a model's ability to discover and reason about causal structures from visual and tabular data. It probes whether models can infer correct causal graphs from observational data, learn disentangled representations from images, and perform valid causal interventions with limited visual samples. Use when the user wants to benchmark on Causal3D, or asks about evaluating this task. Reports correctness of inferred causal structures.
Evaluates Video-Language Models' ability to perform joint retrieval and causal reasoning over two causally separated video clips connected by a 'bridge entity'. It also probes causal world modeling by asking models to identify cause-effect relationships in human behaviors within long videos. Use when the user wants to benchmark on Causal2Needles, or asks about evaluating this task. Reports accuracy.