All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,697–9,720 of 21,377 skills
- Androidworld EvalThis benchmark evaluates the ability of autonomous multimodal agents to navigate and interact with real-world Android applications to complete programmatic user instructions. It probes UI understanding, precise touch interaction, state tracking, and error recovery in a dynamic mobile environment. Use when the user wants to benchmark on AndroidWorld, or asks about evaluating this task. Reports Success Rate (SR).Votes: 0GitHub stars: 3
- Androidcontrol EvalEvaluates the ability of UI control agents to execute mobile app tasks by predicting correct actions given textual screen representations and instruction history. It probes in-domain and out-of-domain generalization, and measures how model performance scales with the volume of training demonstrations. Use when the user wants to benchmark on AndroidControl, or asks about evaluating this task. Reports step-wise accuracy.Votes: 0GitHub stars: 3
- Android Malware Classification EvalEvaluates the ability of Graph Neural Networks to classify Android applications as benign or malicious, and to identify specific malware families or categories, by learning topological patterns from function call graphs. Use when the user wants to benchmark on Malnet-Tiny, Drebin, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Ancholik Ner EvalEvaluates Named Entity Recognition (NER) capabilities across five regional dialects of the Bangla language. It probes a model's ability to correctly identify and classify entities (Person, Location, Organization, Role, Food) in dialect-specific text where linguistic features and vocabulary differ significantly from standard Bangla. Use when the user wants to benchmark on ANCHOLIK-NER, or asks about evaluating this task. Reports F1-score.Votes: 0GitHub stars: 3
- Analytic Score EvalEvaluates the scoring accuracy of interpretable automated scoring frameworks on educational assessment items across three domains. It measures how well LLM-extracted features and ordinal logistic regression align with human raters while adhering to strict interpretability constraints. Use when the user wants to benchmark on Educational Assessment Items (Science, Reading Informational Text, Reading Literature), or asks about evaluating this task. Reports QWK.Votes: 0GitHub stars: 3
- Analogy Multiple Choice EvalEvaluates a language model's ability to perform analogical reasoning and select the correct word pair from multiple choices under temperature scaling. Use when the user wants to benchmark on Analogy Multiple Choice, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Amp Motion Control EvalEvaluates a physics-based character's ability to learn stylized locomotion and complex task execution (e.g., navigating targets, avoiding obstacles) by imitating unstructured motion datasets. It probes the model's capacity to compose disparate skills, generalize across gaits, and maintain high-fidelity motion tracking without manual motion planning. Use when the user wants to benchmark on AMP Motion Datasets, or asks about evaluating this task. Reports normalized task return.Votes: 0GitHub stars: 3
- Amp Classification EvalEvaluates the ability of reprogrammed language models to classify antimicrobial peptide (AMP) sequences into binary categories (toxic vs. non-toxic, or AMP vs. non-AMP) using limited labeled data. Use when the user wants to benchmark on AMP Dataset, or asks about evaluating this task. Reports Test Accuracy.Votes: 0GitHub stars: 3
- Amos Downstream EvalEvaluates the downstream performance of pretrained text encoders on a suite of natural language understanding and reading comprehension benchmarks via standard single-task fine-tuning. Use when the user wants to benchmark on GLUE, SQuAD 2.0, or asks about evaluating this task. Reports AVG.Votes: 0GitHub stars: 3
- Amodal Optical Flow EvalEvaluates a model's ability to predict multi-layered pixel-level motion fields that explicitly account for both visible and occluded regions of objects (amodal optical flow), along with associated masks and semantic labels. It also assesses the utility of these predictions for downstream panoptic tracking. Use when the user wants to benchmark on AmodalSynthDrive, or asks about evaluating this task. Reports AFQ.Votes: 0GitHub stars: 3
- Amodal 3d Reconstruction EvalEvaluates a model's ability to infer occluded (amodal) 3D geometry and predict physically stable configurations in cluttered tabletop scenes. It further tests downstream robotic manipulation success (grasping, pushing, rearranging) under varying levels of visual occlusion. Use when the user wants to benchmark on ShapeNet, MuJoCo Cluttered Tabletop Benchmark, or asks about evaluating this task. Reports Chamfer distance.Votes: 0GitHub stars: 3
- Amo Bench EvalEvaluates large language models' ability to solve high school and IMO-level mathematics competition problems. It probes complex mathematical reasoning, problem-solving under strict constraints, and the model's capacity to scale reasoning effort with test-time compute. Use when the user wants to benchmark on AMO-Bench, or asks about evaluating this task. Reports AVG@32.Votes: 0GitHub stars: 3
- Amigo EvalProbes long-horizon agentic planning, cross-image grounding, and uncertainty-driven question selection. Models must iteratively ask constrained Yes/No/Unsure questions to identify a hidden target from a gallery of visually similar dress images while strictly tracking constraints and avoiding prohibited attributes. Use when the user wants to benchmark on AMIGO, or asks about evaluating this task. Reports identification success.Votes: 0GitHub stars: 3
- Amharicstoryqa EvalEvaluates long-sequence narrative understanding and cultural variation in Amharic using story-based question answering. Probes both multiple-choice and generative QA capabilities across different Ethiopian regional folktales. Use when the user wants to benchmark on AmharicStoryQA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Amharic Asr EvalEvaluates fine-tuned Whisper models for Amharic speech-to-text recognition by measuring transcription accuracy at word and character levels, alongside n-gram overlap. It also probes the impact of homophone normalization and zero-shot generalization on low-resource language ASR performance. Use when the user wants to benchmark on FLEURS Amharic, BDU Speech Corpus, Mozilla Common Voice v17.0 Amharic, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Amega Clinical Reasoning EvalEvaluates the clinical reasoning capabilities and on-device runtime efficiency of various LLMs using the AMEGA benchmark. It measures response accuracy via an LLM-as-a-judge scoring system and tracks inference throughput and thermal throttling effects across different mobile hardware configurations. Use when the user wants to benchmark on AMEGA, or asks about evaluating this task. Reports AMEGA score.Votes: 0GitHub stars: 3
- Amble EvalEvaluates archival domain adaptation capabilities across four distinct tasks. It probes a model's ability to predict document retention periods, classify open access status, determine confidentiality levels, and correct post-OCR text errors in Chinese archival records. Use when the user wants to benchmark on AMBLE, or asks about evaluating this task. Reports F1 score, Levenshtein Distance.Votes: 0GitHub stars: 3
- Ambisql EvalEvaluates a Text-to-SQL system's ability to generate correct SQL from ambiguous natural language queries when integrated with an interactive ambiguity resolution module. It also measures the system's precision, recall, and F1 in detecting and classifying specific types of schema-mapping and reasoning ambiguities. Use when the user wants to benchmark on AmbiSQL Constructed Dataset, or asks about evaluating this task. Reports Exact Match accuracy.Votes: 0GitHub stars: 3
- Ambiqt EvalEvaluates a model's ability to generate diverse, semantically valid SQL queries when a natural language question is ambiguous. It specifically probes whether the model can cover multiple valid interpretations of the same query within a limited set of top-k outputs. Use when the user wants to benchmark on AmbiQT, SPIDER, Kaggle DBQA, or asks about evaluating this task. Reports BothInTopK.Votes: 0GitHub stars: 3
- Ambiguous Emotion Recognition EvalEvaluates audio-language models' ability to recognize ambiguous emotions in speech by predicting full emotion probability distributions and dominant class labels. It specifically probes how test-time scaling (TTS) strategies and model capacity interact with varying levels of emotional ambiguity to improve or degrade recognition performance. Use when the user wants to benchmark on IEMOCAP, MSP-Podcast, CREMA-D, or asks about evaluating this task. Reports JS divergence.Votes: 0GitHub stars: 3
- Ambigqa EvalThis benchmark evaluates a model's ability to identify ambiguous open-domain questions, generate multiple plausible answer spans, and produce disambiguated question rewrites that distinguish between different interpretations of the same query. Use when the user wants to benchmark on AMBIGNQ, or asks about evaluating this task. Reports F1ans.Votes: 0GitHub stars: 3
- Amazon Stark Skb EvalEvaluates the ability of neural retriever-reranker pipelines to accurately retrieve relevant product entities from semi-structured e-commerce knowledge graphs using natural language queries. It probes semantic matching, cross-encoder reranking effectiveness, and the impact of graph-based augmentation on retrieval precision and recall. Use when the user wants to benchmark on Amazon STaRK SKB, or asks about evaluating this task. Reports Hit@1.Votes: 0GitHub stars: 3
- Amazon Review Summarization EvalEvaluates the ability of abstractive summarization models to generate concise, informative summaries of product reviews while preserving aspect and opinion details. It measures lexical overlap with human-written reference summaries. Use when the user wants to benchmark on Amazon Reviews (Healthcare & Electronics), or asks about evaluating this task. Reports ROUGE-1.Votes: 0GitHub stars: 3
- Amazon M2 EvalEvaluates session-based recommendation and text generation capabilities across multiple languages and locales. It probes a model's ability to predict the next product in a shopping session, transfer knowledge across domain-shifted locales, and generate product titles from session context. Use when the user wants to benchmark on Amazon-M2, or asks about evaluating this task. Reports next-product prediction.Votes: 0GitHub stars: 3