All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,385
- 892
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,865–9,888 of 21,385 skills
- Active Look Hallucination EvalEvaluates the ability of large vision-language models to mitigate object-existence hallucinations by dynamically allocating visual computation based on uncertainty. It probes fine-grained perception, object counting, spatial reasoning, and color recognition under adaptive visual grounding. Use when the user wants to benchmark on POPE, MME, CHAIR, or asks about evaluating this task. Reports POPE Accuracy.Votes: 0GitHub stars: 3
- Actionbench EvalEvaluates a text-to-image model's ability to customize generated images with a specific subject's appearance while accurately transferring a target action from an exemplar image, without appearance leakage or subject deformation. Use when the user wants to benchmark on ActionBench, or asks about evaluating this task. Reports total accuracy.Votes: 0GitHub stars: 3
- Action Prediction EvalTests translating a user's natural language command into a structured executable action (function, arguments, status) for GUI interaction. It bridges high-level intent with precise, structured API-like calls. Use when the user wants to benchmark on GUI-360°-Bench, or asks about evaluating this task. Reports Step success rate.Votes: 0GitHub stars: 3
- Acronym Id Disamb EvalEvaluates a model's ability to identify acronym boundaries and their corresponding long forms in scientific text, and to disambiguate ambiguous acronyms by selecting the correct long form from a candidate set. Use when the user wants to benchmark on SciAI, SciAD, or asks about evaluating this task. Reports Macro F1.Votes: 0GitHub stars: 3
- Acpbench Hard EvalEvaluates language and reasoning models on open-ended, generative planning tasks derived from PDDL domains. It probes capabilities like action applicability, reachability, progression, justification, and next-action prediction without predefined answer choices. Use when the user wants to benchmark on ACPBench Hard, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Acol Interpreter Benchmark EvalEvaluates the execution speed and dispatching efficiency of different interpreter implementations (AST vs. bytecode variants) for a simple imperative language (ACOL) in Prolog. Use when the user wants to benchmark on ACOL Interpreter Benchmarks, or asks about evaluating this task. Reports geometric_mean_runtime.Votes: 0GitHub stars: 3
- Aces EvalEvaluates machine translation metrics on their ability to correctly rank good translations above incorrect ones across specific linguistic error phenomena. It probes metric robustness to fine-grained translation errors like hallucination, omission, and real-world knowledge failures. Use when the user wants to benchmark on ACES, or asks about evaluating this task. Reports Kendall's tau-like correlation.Votes: 0GitHub stars: 3
- Ace2005 Ner EvalEvaluates the ability of pretrained models to recognize named entities in open-domain text, specifically probing generalization when name regularity and mention coverage are manipulated. It measures how well models rely on contextual patterns versus superficial spelling cues. Use when the user wants to benchmark on ACE2005, or asks about evaluating this task. Reports Micro-F1.Votes: 0GitHub stars: 3
- Ace Reason Nemotron EvalThis evaluation protocol assesses the mathematical reasoning and code generation capabilities of large language models. It probes the model's ability to solve competitive math problems and implement algorithms for coding contests under strict generation constraints. Use when the user wants to benchmark on AIME2024, AIME2025, MATH500, HMMT2025 Feb, BRUMO2025, LiveCodeBench v5, LiveCodeBench v6, Codeforces (LiveCodeBench Pro), EvalPlus, or asks about evaluating this task. Reports avg@k.Votes: 0GitHub stars: 3
- Ace Mol EvalProbes molecular representation models on their ability to predict chemical properties and classify molecular structures. It evaluates how well pre-trained embeddings capture task-relevant chemical motifs when probed with linear classifiers or regressors. Use when the user wants to benchmark on MoleculeNet, Photoswitch, Synthetic Toxicity Benchmark, or asks about evaluating this task. Reports %AUCROC, MAE.Votes: 0GitHub stars: 3
- Ace MetricProbes the physical plausibility and spatiotemporal flow preservation of data-driven weather forecasting models by directly measuring errors in horizontal (advection) and vertical (convection) atmospheric motions, rather than relying on pixel-wise accuracy metrics that reward blurriness. Use when the user has predictions and gold and needs to compute ACE.Votes: 0GitHub stars: 3
- AccuracyProbes the pairwise ranking accuracy of an AI judge system when evaluating generated commit messages against a heuristic ground truth derived from multiple automatic text generation metrics. Use when the user has predictions and gold and needs to compute accuracy.Votes: 0GitHub stars: 3
- Accflow EvalEvaluates a model's ability to estimate long-range dense optical flow between distant video frames, specifically testing robustness to large motions and severe occlusions. It measures how well the model accumulates flow over multiple steps while correcting misalignments and occlusion artifacts. Use when the user wants to benchmark on CVO, HS-Sintel, or asks about evaluating this task. Reports EPE.Votes: 0GitHub stars: 3
- Accesseeval EvalThis benchmark probes systematic disability bias in large language models by comparing responses to neutral queries versus disability-aware queries. It measures whether explicitly mentioning a disability degrades response quality across sentiment, social perception, and factual accuracy dimensions. Use when the user wants to benchmark on AccessEval, or asks about evaluating this task. Reports Bias Degradation Rate ($\Delta_M$).Votes: 0GitHub stars: 3
- Accented Clinical Asr EvalEvaluates ASR models on African-accented clinical speech to measure how well they transcribe medical named entities (MNEs) like drug names, diagnoses, and lab results. It specifically probes the gap between standard word-level accuracy and clinically relevant entity recognition. Use when the user wants to benchmark on AfriSpeech, or asks about evaluating this task. Reports M-WER.Votes: 0GitHub stars: 3
- Accentbox EvalEvaluates a zero-shot text-to-speech system's ability to generate speech with high-fidelity target accents while preserving the reference speaker's voice. It probes the model's capacity to disentangle accent characteristics from speaker identity using continuous embeddings, measuring both objective acoustic similarity and subjective listener preference across inherent and cross-accent generation tasks. Use when the user wants to benchmark on Common Voice v17.0 (English), VCTK, LibriTTS-R (cle...Votes: 0GitHub stars: 3
- Acappella Separation EvalThis benchmark evaluates audio-visual singing voice separation models by measuring how accurately they isolate target singing voices from mixed audio accompanied by video. It probes the model's ability to leverage visual motion cues (face landmarks) to separate overlapping or low-volume singing voices across different languages and volume conditions. Use when the user wants to benchmark on Acappella, or asks about evaluating this task. Reports SDR.Votes: 0GitHub stars: 3
- Academiceval EvalEvaluates an LLM's ability to perform long-context summarization and synthesize related work sections by retrieving and reasoning over heterogeneous academic memory chunks. Use when the user wants to benchmark on AcademicEval-abstract, AcademicEval-related, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Aca Sentiment EvalEvaluates the ability of sentiment analysis models to classify short social media posts and movie reviews into discrete sentiment categories. It probes how well supervised and unsupervised word representations capture task-specific sentiment orientation. Use when the user wants to benchmark on ACA, Stanford Sentiment Treebank (SST), or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Ac Vrnn Trajectory Prediction EvalEvaluates a model's ability to predict multi-modal future trajectories of agents given historical positions. It probes the model's capacity to capture social interactions, scene constraints, and long-term motion dynamics across diverse environments. Use when the user wants to benchmark on ETH, UCY, Stanford Drone Dataset (SDD), STATS SportVU NBA, Intersection Drone Dataset (inD), TrajNet++, or asks about evaluating this task. Reports TopK ADE, TopK FDE.Votes: 0GitHub stars: 3
- Abstract Image Visual Reasoning EvalEvaluates multimodal models' ability to comprehend and reason over synthetic abstract images, including charts, tables, road maps, dashboards, relation graphs, flowcharts, visual puzzles, and planar layouts. Use when the user wants to benchmark on Synthetic Abstract Image Benchmark, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Abstain Qa EvalEvaluates large language models' ability to abstain from answering when questions are unanswerable or when uncertain, while maintaining accuracy on answerable questions. It measures how well models balance abstention with correct answer selection under different prompting strategies and uncertainty calibration methods. Use when the user wants to benchmark on Abstain-QA, or asks about evaluating this task. Reports Abstention Rate (AR).Votes: 0GitHub stars: 3
- Absa Sentiment EvalProbes an LLM's ability to identify granular sentiment toward specific topics within a text, including custom labels like 'not mentioned' and 'wished for'. It tests fine-grained aspect-level classification rather than overall review sentiment. Use when the user wants to benchmark on TravelBench ABSA, or asks about evaluating this task. Reports F1-score.Votes: 0GitHub stars: 3
- Abs Rel Depth ClassProbes the ability of depth estimation models to accurately predict distances for specific semantic classes, particularly focusing on thin structures like wires and cables. It measures class-specific absolute relative error to highlight performance on challenging, low-pixel-count obstacles relevant to drone navigation. Use when the user has predictions and gold and needs to compute AbsRel_class.Votes: 0GitHub stars: 3