Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,867
skills in category
870
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,281–8,304 of 20,867 skills

Court Judgment Prediction EvalA

Tests a model's ability to predict binary case outcomes (accepted/denied) and generate human-readable explanations by citing relevant sentences from the input document. This probes joint reasoning, outcome forecasting, and justification generation in legal contexts. Use when the user wants to benchmark on LegalEval CJPE Dataset, or asks about evaluating this task. Reports standard F1 score.

researchpythongo
0
3
Countqa EvalA

This benchmark evaluates the object counting and spatial individuation capabilities of multimodal large language models (MLLMs) on real-world images characterized by high density, clutter, and occlusion. It probes whether generalist models can perform precise, fine-grained visual grounding and numerical reasoning out-of-the-box without specialized training. Use when the user wants to benchmark on CountQA, or asks about evaluating this task. Reports Exact Match (EM).

researchpythonperformance
0
3
Counterfactual Text Gen EvalA

This benchmark evaluates the effectiveness and linguistic quality of counterfactual text generation methods. It probes a model's ability to modify input text to flip a target classifier's predicted label while preserving grammatical correctness, fluency, and coherence, highlighting the trade-off between label-flipping success and text quality. Use when the user wants to benchmark on IMDB, SNLI, or asks about evaluating this task. Reports flip rate (FR).

researchpythongo
0
3
Counterfactual Situation Testing EvalA

Evaluates a fairness auditing framework's ability to detect individual discrimination in decision-making systems by comparing factual outcomes against counterfactual or similar-group outcomes. It probes whether protected attributes causally influence decisions beyond legitimate factors. Use when the user wants to benchmark on Synthetic Loan Application, Law School Admissions, or asks about evaluating this task. Reports individual discrimination cases.

researchpythontesting
0
3
Counterfactual Rep EvalA

Evaluates how well counterfactual representations (CFRs) in high-dimensional embedding space mimic true text counterfactuals, measuring prediction consistency, probability alignment, and downstream fairness improvements across synthetic and real-world biased datasets. Use when the user wants to benchmark on EEEC+, BiasInBios, or asks about evaluating this task. Reports PIP.

researchpythongit
0
3
Counterfactual Image Gen EvalA

Evaluates the capability of generative models to produce counterfactual images that preserve causal consistency, maintain realism, and minimally alter non-target attributes under specified interventions. It probes composition stability, attribute manipulation effectiveness, and distributional fidelity across varying dataset complexities and causal graphs. Use when the user wants to benchmark on MorphoMNIST, CelebA, ADNI, or asks about evaluating this task. Reports FID, CLD.

researchpythongo
0
3
Counterfactual Fairness Qa EvalA

Evaluates whether LLM-based contact center QA systems exhibit systematic bias when agent identity (gender, ethnicity, religion, disability) or contextual factors (past performance, behavioral style) are counterfactually altered. Measures if model judgments change disproportionately based on these attributes rather than transcript content. Use when the user wants to benchmark on Contact-Center QA Transcripts, or asks about evaluating this task. Reports Counterfactual Flip Rate (CFR).

researchpythonperformance
0
3
Counterfactual Fairness EvalA

Evaluates the trade-off between predictive accuracy and counterfactual fairness on real-world datasets. It measures how well a model's predictions remain invariant to sensitive attributes (race, gender) while maintaining performance on regression or classification tasks. Use when the user wants to benchmark on LSAC, Compas, Adult, or asks about evaluating this task. Reports Balanced Accuracy.

researchpythonperformance
0
3
Counterfactual Detection EvalA

Evaluates a model's ability to detect counterfactual statements in product reviews. It probes robustness to selection bias from clue phrases, cross-lingual transfer via machine translation, and the effectiveness of different sentence encoders and classifiers on imbalanced binary classification tasks. Use when the user wants to benchmark on Multilingual Counterfactual Detection Dataset (Amazon Reviews), or asks about evaluating this task. Reports F1.

researchpythonexpress
0
3
Counterfactual Chaos EvalA

Evaluates the reliability of counterfactual trajectory estimation in chaotic versus non-chaotic dynamical systems under parameter uncertainty and observational noise. It probes whether Bayesian filtering and particle-based smoothing can accurately recover 'what-if' scenarios when small initial perturbations lead to divergent outcomes. Use when the user wants to benchmark on Lorenz System, Rössler System, Logistic Growth, or asks about evaluating this task. Reports RMSE_t.

researchpythonperformance
0
3
Counterfact Edit EvalA

Evaluates the effectiveness and stability of sequential LLM knowledge editing methods. It measures how well a model updates a specific fact while preserving related paraphrases, neighboring facts, generation fluency, and overall general capabilities over thousands of edits. Use when the user wants to benchmark on CounterFact, GLUE_MMLU_GSM8K_HumanEval_MBPP, or asks about evaluating this task. Reports Efficacy.

researchpythongo
0
3
Countereval EvalA

Evaluates the human-centric quality of counterfactual explanations across eight explanatory virtues. It probes how well explanations convey desired outcomes, remain feasible, consistent, complete, trustworthy, understandable, fair, and appropriately complex. Use when the user wants to benchmark on CounterEval, or asks about evaluating this task. Reports Overall Satisfaction, Feasibility, Consistency, Completeness, Trust, Understandability, Fairness, Complexity.

researchpythonrust
0
3
Cotrec EvalA

Evaluates a model's ability to predict the next item in a user session based on historical click sequences. It probes the model's capacity to capture sequential dependencies and session-level patterns in sparse e-commerce or media interaction data. Use when the user wants to benchmark on Tmall, RetailRocket, Diginetica, or asks about evaluating this task. Reports P@10.

researchpythongo
0
3
Cotomod EvalA

Evaluates how well AI models can assist human moderators in content moderation by prioritizing comments for review. It probes the model's ability to estimate uncertainty accurately and guide human review capacity to maximize collaborative accuracy and efficiency under constraints. Use when the user wants to benchmark on CoToMoD, or asks about evaluating this task. Reports OC-Acc.

researchpythongo
0
3
Cot Faithfulness EvalA

Evaluates whether reasoning models explicitly acknowledge external hint injections within their chain-of-thought reasoning traces. It probes model transparency and the alignment between internal reasoning tokens and final output disclosures. Use when the user wants to benchmark on MMLU, GPQA Diamond, or asks about evaluating this task. Reports faithfulness.

researchpythongo
0
3
Cot EvalA

Evaluates the zero-shot and few-shot reasoning capabilities of language models, specifically probing their ability to generate step-by-step chain-of-thought rationales and produce correct answers across classification and generation tasks. Use when the user wants to benchmark on BigBench Hard (BBH), P3, MGSM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cosyvoice Tts EvalA

Evaluates zero-shot text-to-speech synthesis quality, focusing on content consistency (how well generated speech matches input text) and speaker similarity (how well the cloned voice matches the reference speaker) across English and Chinese. It also probes emotion controllability and the utility of synthesized speech for augmenting ASR training data. Use when the user wants to benchmark on LibriTTS, AISHELL-3, or asks about evaluating this task. Reports WER (%), CER (%).

researchpythonshell
0
3
Costnav EvalA

Economic viability and cost-aware performance of embodied agents in urban sidewalk delivery navigation. It evaluates how technical metrics like collision rate and arrival success translate into real-world financial outcomes, including maintenance costs, energy usage, revenue, and break-even points. Use when the user wants to benchmark on CostNav Urban Sidewalk Navigation Simulation, or asks about evaluating this task. Reports Profit/run.

researchpythongit
0
3
Cosql EvalA

Evaluates conversational text-to-SQL systems on cross-domain database querying. It probes dialogue state tracking via SQL grounding, response generation from query results, and user intent/dialogue act prediction under real-world ambiguity and clarification dynamics. Use when the user wants to benchmark on CoSQL, or asks about evaluating this task. Reports Question Match.

researchpythongo
0
3
Cosql Cg EvalA

This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on CoSQL-CG, or asks about evaluating this task. Reports question match (QM).

researchpythongo
0
3
Cosmos Qa EvalA

This benchmark evaluates a model's ability to perform contextual commonsense reasoning in machine reading comprehension. It probes whether systems can make non-literal, implicit inferences about causes, effects, and counterfactuals based on personal narratives, rather than relying on explicit textual evidence or simple semantic matching. Use when the user wants to benchmark on Cosmos QA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Cosmos Drive Dreams EvalA

Evaluates the effectiveness of a synthetic driving data generation pipeline by measuring performance gains in downstream autonomous driving perception tasks, including 3D lane detection, 3D object detection, and LiDAR-based detection, particularly under challenging conditions like extreme weather and nighttime. Use when the user wants to benchmark on Waymo Open Dataset, RDS-HQ, RDS-HQ-HL, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Cosmoflow Hpc ScalingA

Evaluates the compute efficiency and horizontal scalability of a 3D convolutional neural network framework on supercomputers, measuring sustained floating-point throughput and parallel scaling efficiency across thousands of nodes. Use when the user has predictions and gold and needs to compute Pflop/s.

researchpythongo
0
3
Cosmic Symmetry Benchmark EvalA

Evaluates the ability of graph neural networks to extract cosmological parameters and local velocity fields from large-scale 3D point clouds of dark matter halos. It probes both global long-range correlation capture (via cosmological parameter regression) and local geometric dependency modeling (via per-node velocity prediction). Use when the user wants to benchmark on Quijote BSQ Point Cloud, or asks about evaluating this task. Reports MSE.

researchpythonnode
0
3