All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,649–9,672 of 21,377 skills
- Araspider EvalEvaluates the quality of English-to-Arabic translations for text-to-SQL tasks and measures the execution accuracy of generated SQL queries on Arabic natural language questions. Use when the user wants to benchmark on AraSpider, or asks about evaluating this task. Reports Execution accuracy.Votes: 0GitHub stars: 3
- Arahahealthqa EvalThis benchmark evaluates Arabic language models on healthcare-related question answering, specifically probing their ability to classify mental health conditions and generate culturally appropriate medical advice. It tests both discriminative capabilities (multi-label classification and multiple-choice selection) and generative capabilities (open-ended response generation) in clinical and mental health contexts. Use when the user wants to benchmark on AraHealthQA, or asks about evaluating thi...Votes: 0GitHub stars: 3
- Aradice EvalEvaluates LLMs' capabilities in understanding and generating dialectal Arabic (Levantine, Egyptian, Gulf) and assessing cultural awareness. It probes dialect identification, text generation, cognitive reasoning, and machine translation across dialects diverging from Modern Standard Arabic. Use when the user wants to benchmark on AraDiCE, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Arabicmmlu EvalArabicMMLU probes massive multitask language understanding in Modern Standard Arabic across 40 educational subjects spanning STEM, social sciences, humanities, Arabic language, and culturally specific domains. It evaluates models on cross-lingual transfer, cultural localization, and robustness to linguistic phenomena like negation across primary, middle, high school, and university levels. The benchmark specifically tests how well models handle Arabic-specific knowledge and exam-style multipl...Votes: 0GitHub stars: 3
- Arabic Sts Mteb EvalEvaluates the semantic textual similarity (STS) capability of Arabic text embedding models, specifically testing how Matryoshka Representation Learning and hybrid loss training preserve semantic alignment across different embedding dimensions. Use when the user wants to benchmark on MTEB Arabic STS (STS17, STS22, STS22-v2), or asks about evaluating this task. Reports STS correlation (Pearson/Spearman, scaled 0-100).Votes: 0GitHub stars: 3
- Arabic Evidence Retrieval EvalThis benchmark tests a system's ability to retrieve relevant evidence snippets from a large pool of web pages for a given Arabic claim. It evaluates ranking performance in a retrieval setting where only a small fraction of snippets actually contain verifying evidence. Use when the user wants to benchmark on Arabic Evidence Retrieval Dataset, or asks about evaluating this task. Reports P@10.Votes: 0GitHub stars: 3
- Arabic Claim Verification EvalThis benchmark evaluates a model's ability to classify the veracity of Arabic social media claims as true or false. It probes factual consistency and reasoning against reliable sources in a binary classification setting. Use when the user wants to benchmark on Arabic Claim Verification Dataset, or asks about evaluating this task. Reports Macro-F1.Votes: 0GitHub stars: 3
- Arabic Check Worthiness EvalThis benchmark evaluates a model's ability to identify check-worthy factual claims within Arabic social media posts. It probes the system's capacity to filter out non-factual or irrelevant content and prioritize claims that require verification based on public interest and potential impact. Use when the user wants to benchmark on Arabic Check-Worthiness Dataset, or asks about evaluating this task. Reports P@30.Votes: 0GitHub stars: 3
- Arabic Call Asr EvalEvaluates the ability of Automatic Speech Recognition (ASR) models to accurately transcribe spoken Arabic from real-world telephonic calls. It probes robustness to dialectal diversity, variable audio quality, and background noise typical of call-domain environments. Use when the user wants to benchmark on Arabic Call Domain Benchmark, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Arabculture EvalEvaluates large language models' ability to perform commonsense reasoning within specific Arab cultural contexts. It probes region-specific grounding by testing models across 13 Arab countries and 12 cultural domains, measuring how well they understand local norms, habits, and social scenarios in both Arabic and English. Use when the user wants to benchmark on ArabCulture, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Ar Mot EvalEvaluates a model's ability to track multiple visual objects in video sequences using auditory referring expressions instead of text. It probes cross-modal alignment, robustness to varying weather and video quality conditions, and handling of complex, unconstrained traffic dynamics. Use when the user wants to benchmark on Echo-KITTI, Echo-KITTI+, Echo-BDD, or asks about evaluating this task. Reports HOTA.Votes: 0GitHub stars: 3
- Ar Bench EvalEvaluates large language models on appellate review tasks for criminal judgments, specifically detecting, classifying, and correcting legal errors in finalized court decisions. It probes fine-grained legal reasoning, diagnostic accuracy, and the ability to generate legally valid corrections. Use when the user wants to benchmark on AR-Bench, or asks about evaluating this task. Reports Accuracy (Acc), Macro F1 (MaF1).Votes: 0GitHub stars: 3
- Aquila Vl EvalEvaluates a 2B-parameter vision-language model's visual understanding, knowledge reasoning, and text reading capabilities across a comprehensive suite of standard multimodal benchmarks. It measures how well the model handles general VQA, mathematical reasoning, and document comprehension tasks. Use when the user wants to benchmark on MMBench, MMStar, MMMU, MathVista, HallusionBench, AI2D, OCRBench, MMVet, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aqua20 EvalEvaluates deep learning models' ability to classify marine species from underwater images under challenging environmental conditions like turbidity, low illumination, and occlusion. It probes robustness to visual distortions, class imbalance, and fine-grained feature discrimination in complex aquatic scenes. Use when the user wants to benchmark on AQUA20, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Aqua Bench EvalThis benchmark evaluates audio question answering models on their ability to correctly answer standard multiple-choice questions and, crucially, to detect and reject unanswerable cases. It specifically probes three failure modes: missing correct options, categorical mismatches between questions and answers, and questions irrelevant to the audio input. Use when the user wants to benchmark on AQUA-Bench, or asks about evaluating this task. Reports conditional accuracy (CA).Votes: 0GitHub stars: 3
- Apst Safety EvalThis evaluation probes the operational reliability and safety alignment of large language models under repeated inference. It specifically measures how stochastic decoding and sampling depth expose intermittent safety failures, refusal inconsistencies, and guardrail instability that single-generation benchmarks typically mask. Use when the user wants to benchmark on APST Safety Prompt Set (AIR-BENCH Equivalent), or asks about evaluating this task. Reports empirical failure probability.Votes: 0GitHub stars: 3
- Apres Paper Revision EvalThis protocol evaluates an LLM's ability to predict a paper's future scientific impact based on its text and peer reviews, and its ability to iteratively revise the manuscript to maximize that predicted impact. It probes the model's capacity for rubric discovery, agentic text editing, and alignment with human expert preferences. Use when the user wants to benchmark on ICLR & NeurIPS Peer Review Dataset, or asks about evaluating this task. Reports MAE, Improvement Score ($\Delta S$).Votes: 0GitHub stars: 3
- Apr Plausible Patch EvalEvaluates the ability of code language models to automatically generate correct patches for real-world Java bugs. It probes whether pre-trained or fine-tuned models can produce syntactically valid and semantically correct code that passes developer-written test suites and survives manual verification. Use when the user wants to benchmark on Defects4J v1.2, Defects4J v2.0, QuixBugs, HumanEval-Java, or asks about evaluating this task. Reports plausible_patch.Votes: 0GitHub stars: 3
- Apptek Callcenter Asr EvalThis benchmark evaluates automatic speech recognition (ASR) systems on their ability to transcribe long-form, spontaneous call-center dialogues across 14 English accents. It specifically probes robustness to non-standard accents, conversational speech patterns, and sensitivity to audio segmentation strategies. Use when the user wants to benchmark on AppTek Call-Center Dialogues, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Apps EvalEvaluates a model's ability to generate correct Python code from natural language problem descriptions. It measures functional correctness by executing generated programs against a large bank of automated test cases, rather than relying on text-similarity metrics like BLEU. Use when the user wants to benchmark on APPS, or asks about evaluating this task. Reports strict_accuracy.Votes: 0GitHub stars: 3
- Appearance Free Action EvalEvaluates zero-shot generalization of action recognition models to appearance-free videos generated by warping noise or random dots with optical flow, testing reliance on motion cues over static shape and texture. Use when the user wants to benchmark on UCF5, AFD5, AFF5, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Appear2meaning EvalThis benchmark probes vision-language models' ability to infer non-observable, culturally grounded structured metadata (culture, period, origin, creator) from images of heritage artifacts. It evaluates whether models can go beyond visual perception to perform semantic alignment with museum annotations, revealing culture-dependent reasoning capabilities and potential biases. Use when the user wants to benchmark on Appear2Meaning, or asks about evaluating this task. Reports exact match accuracy.Votes: 0GitHub stars: 3
- Apistox EvalEvaluates the ability of molecular graph machine learning models and fingerprint-based methods to predict binary pesticide toxicity to honey bees. It specifically probes domain generalization by testing performance on structurally novel compounds and temporally separated data rather than random splits. Use when the user wants to benchmark on ApisTox, or asks about evaluating this task. Reports MCC.Votes: 0GitHub stars: 3
- Apibench Q EvalEvaluates the retrieval accuracy and ranking quality of query-based API recommendation systems for Java APIs at both class and method levels. It also measures how query reformulation techniques impact recommendation performance. Use when the user wants to benchmark on APIBench-Q, or asks about evaluating this task. Reports Success Rate@k.Votes: 0GitHub stars: 3