Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,305–8,328 of 20,898 skills
Evaluates a lightweight CNN's ability to classify chest radiology images (CXR and CT scans) as positive or negative for COVID-19, and in a three-class setting. It also tests cross-modality generalization by training on CT and testing on CXR data. Use when the user wants to benchmark on CXR/CT Chest Radiology Dataset, or asks about evaluating this task. Reports accuracy.
Binary classification of chest X-ray images to detect SARS-CoV-2 infection. It probes a model's ability to distinguish COVID-19 positive cases from negative cases (including no pneumonia and non-SARS-CoV-2 pneumonia) using a large, multinational dataset. Use when the user wants to benchmark on COVID-Net CXR-2 benchmark dataset, or asks about evaluating this task. Reports Sensitivity.
This benchmark evaluates deep graph generative models (JT-VAE and DQN) for their ability to design novel molecular structures optimized for high predicted potency against the SARS-CoV-2 3CL-protease, while balancing drug-likeness, lipophilicity, and synthesizability. It also assesses structural novelty relative to known antivirals and predicted binding affinity using in silico classifiers. Use when the user wants to benchmark on ChEMBL/BindingDB/ToxCat pharmacology dataset, or asks about eval...
Evaluates the impact of five image enhancement techniques (histogram equalization, CLAHE, complement, gamma correction, BCET) on six CNN architectures for three-class classification (COVID-19, lung opacity, normal) using chest X-ray images. It also assesses whether lung segmentation improves classification accuracy and model interpretability. Use when the user wants to benchmark on COVQU-20, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates AI models for detecting COVID-19 infection and assessing lung severity using lung ultrasound videos, clinical variables, and blood count data. It measures how well zero-shot and fine-tuned models generalize to real-world, heterogeneous clinical data compared to human annotators and tabular baselines. Use when the user wants to benchmark on COVID-BLUeS, or asks about evaluating this task. Reports accuracy.
Measures the alignment between a proxy benchmark's ranking and a target reward modeling benchmark's ranking at the top-k positions. It quantifies how many of the highest-performing models on a reward benchmark are also identified as top performers on a given proxy benchmark. Use when the user has predictions and gold and needs to compute coverage_at_top_k.
Evaluates LLM safety guardrails and policy-adaptation frameworks on their ability to correctly identify harmful, toxic, or policy-violating content across diverse attack vectors. It probes robustness against automated jailbreaks, over-refusal in benign contexts, and zero-shot adaptability to out-of-domain policy enforcement. Use when the user wants to benchmark on AdvBenchM, WildGuard, HarmBench, JailJudge, PKU-SafeRLHF, ToxicChat, BeaverTails, XSTest, PAN Wikipedia Vandalism Corpus 2010, Hum...
Tests a model's ability to predict binary case outcomes (accepted/denied) and generate human-readable explanations by citing relevant sentences from the input document. This probes joint reasoning, outcome forecasting, and justification generation in legal contexts. Use when the user wants to benchmark on LegalEval CJPE Dataset, or asks about evaluating this task. Reports standard F1 score.
This benchmark evaluates the object counting and spatial individuation capabilities of multimodal large language models (MLLMs) on real-world images characterized by high density, clutter, and occlusion. It probes whether generalist models can perform precise, fine-grained visual grounding and numerical reasoning out-of-the-box without specialized training. Use when the user wants to benchmark on CountQA, or asks about evaluating this task. Reports Exact Match (EM).
This benchmark evaluates the effectiveness and linguistic quality of counterfactual text generation methods. It probes a model's ability to modify input text to flip a target classifier's predicted label while preserving grammatical correctness, fluency, and coherence, highlighting the trade-off between label-flipping success and text quality. Use when the user wants to benchmark on IMDB, SNLI, or asks about evaluating this task. Reports flip rate (FR).
Evaluates a fairness auditing framework's ability to detect individual discrimination in decision-making systems by comparing factual outcomes against counterfactual or similar-group outcomes. It probes whether protected attributes causally influence decisions beyond legitimate factors. Use when the user wants to benchmark on Synthetic Loan Application, Law School Admissions, or asks about evaluating this task. Reports individual discrimination cases.
Evaluates how well counterfactual representations (CFRs) in high-dimensional embedding space mimic true text counterfactuals, measuring prediction consistency, probability alignment, and downstream fairness improvements across synthetic and real-world biased datasets. Use when the user wants to benchmark on EEEC+, BiasInBios, or asks about evaluating this task. Reports PIP.
Evaluates the capability of generative models to produce counterfactual images that preserve causal consistency, maintain realism, and minimally alter non-target attributes under specified interventions. It probes composition stability, attribute manipulation effectiveness, and distributional fidelity across varying dataset complexities and causal graphs. Use when the user wants to benchmark on MorphoMNIST, CelebA, ADNI, or asks about evaluating this task. Reports FID, CLD.
Evaluates whether LLM-based contact center QA systems exhibit systematic bias when agent identity (gender, ethnicity, religion, disability) or contextual factors (past performance, behavioral style) are counterfactually altered. Measures if model judgments change disproportionately based on these attributes rather than transcript content. Use when the user wants to benchmark on Contact-Center QA Transcripts, or asks about evaluating this task. Reports Counterfactual Flip Rate (CFR).
Evaluates the trade-off between predictive accuracy and counterfactual fairness on real-world datasets. It measures how well a model's predictions remain invariant to sensitive attributes (race, gender) while maintaining performance on regression or classification tasks. Use when the user wants to benchmark on LSAC, Compas, Adult, or asks about evaluating this task. Reports Balanced Accuracy.
Evaluates a model's ability to detect counterfactual statements in product reviews. It probes robustness to selection bias from clue phrases, cross-lingual transfer via machine translation, and the effectiveness of different sentence encoders and classifiers on imbalanced binary classification tasks. Use when the user wants to benchmark on Multilingual Counterfactual Detection Dataset (Amazon Reviews), or asks about evaluating this task. Reports F1.
Evaluates the reliability of counterfactual trajectory estimation in chaotic versus non-chaotic dynamical systems under parameter uncertainty and observational noise. It probes whether Bayesian filtering and particle-based smoothing can accurately recover 'what-if' scenarios when small initial perturbations lead to divergent outcomes. Use when the user wants to benchmark on Lorenz System, Rössler System, Logistic Growth, or asks about evaluating this task. Reports RMSE_t.
Evaluates the effectiveness and stability of sequential LLM knowledge editing methods. It measures how well a model updates a specific fact while preserving related paraphrases, neighboring facts, generation fluency, and overall general capabilities over thousands of edits. Use when the user wants to benchmark on CounterFact, GLUE_MMLU_GSM8K_HumanEval_MBPP, or asks about evaluating this task. Reports Efficacy.
Evaluates the human-centric quality of counterfactual explanations across eight explanatory virtues. It probes how well explanations convey desired outcomes, remain feasible, consistent, complete, trustworthy, understandable, fair, and appropriately complex. Use when the user wants to benchmark on CounterEval, or asks about evaluating this task. Reports Overall Satisfaction, Feasibility, Consistency, Completeness, Trust, Understandability, Fairness, Complexity.
Evaluates a model's ability to predict the next item in a user session based on historical click sequences. It probes the model's capacity to capture sequential dependencies and session-level patterns in sparse e-commerce or media interaction data. Use when the user wants to benchmark on Tmall, RetailRocket, Diginetica, or asks about evaluating this task. Reports P@10.
Evaluates how well AI models can assist human moderators in content moderation by prioritizing comments for review. It probes the model's ability to estimate uncertainty accurately and guide human review capacity to maximize collaborative accuracy and efficiency under constraints. Use when the user wants to benchmark on CoToMoD, or asks about evaluating this task. Reports OC-Acc.
Evaluates whether reasoning models explicitly acknowledge external hint injections within their chain-of-thought reasoning traces. It probes model transparency and the alignment between internal reasoning tokens and final output disclosures. Use when the user wants to benchmark on MMLU, GPQA Diamond, or asks about evaluating this task. Reports faithfulness.
Evaluates the zero-shot and few-shot reasoning capabilities of language models, specifically probing their ability to generate step-by-step chain-of-thought rationales and produce correct answers across classification and generation tasks. Use when the user wants to benchmark on BigBench Hard (BBH), P3, MGSM, or asks about evaluating this task. Reports accuracy.
Evaluates zero-shot text-to-speech synthesis quality, focusing on content consistency (how well generated speech matches input text) and speaker similarity (how well the cloned voice matches the reference speaker) across English and Chinese. It also probes emotion controllability and the utility of synthesized speech for augmenting ASR training data. Use when the user wants to benchmark on LibriTTS, AISHELL-3, or asks about evaluating this task. Reports WER (%), CER (%).