All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,793–9,816 of 21,377 skills
- Agentehr EvalEvaluates autonomous clinical decision-making agents on Electronic Health Record (EHR) data. It probes multi-step reasoning, long-context dependency preservation, and robustness to distribution shifts across different hospital databases and clinical event types. Use when the user wants to benchmark on MIMIC-IV / MIMIC-III, or asks about evaluating this task. Reports average score.Votes: 0GitHub stars: 3
- Agentds EvalThis benchmark evaluates AI agents and human-AI collaboration on domain-specific data science tasks across six industries. It probes the ability to perform feature engineering, integrate multimodal data (images, text, PDFs, JSON), and build predictive models that require genuine domain reasoning rather than generic pipelines. Use when the user wants to benchmark on AgentDS, or asks about evaluating this task. Reports quantile_score.Votes: 0GitHub stars: 3
- Agentboard EvalEvaluates an agent's ability to complete multi-round interactive tasks and achieve target goals across diverse task types. It simulates real-world environments where the model must navigate sequential decision-making to reach a defined endpoint. Use when the user wants to benchmark on AgentBoard, or asks about evaluating this task. Reports target achievement rate.Votes: 0GitHub stars: 3
- Agent Spec EvalEvaluates the cross-framework portability and reusability of declarative agent specifications by executing identical agentic workflows across four different runtime frameworks (AutoGen, CrewAI, LangGraph, WayFlow) on three distinct task benchmarks. Use when the user wants to benchmark on SimpleQA Verified, BIRD-SQL, $\tau^{2}$-Bench, or asks about evaluating this task. Reports F1 score, EX%, Passˆk.Votes: 0GitHub stars: 3
- Agent Red Teaming EvalEvaluates the security and robustness of frontier AI agents against adversarial prompt injections and policy violations across multiple models and realistic deployment scenarios. It measures how effectively attacks transfer between models and whether model capability or inference compute correlates with safety. Use when the user wants to benchmark on Agent Red Teaming (ART) benchmark, or asks about evaluating this task. Reports ASR.Votes: 0GitHub stars: 3
- Age Prediction Fairness EvalEvaluates the accuracy and demographic fairness of deep learning models for age prediction from facial images. It probes the model's ability to generalize across diverse camera settings and ethnicities/genders while mitigating bias through distribution-aware curation and augmentation. Use when the user wants to benchmark on APPA-REAL, MORPH-2, UTKFace, Mega Asian, AFAD, CACD, or asks about evaluating this task. Reports MAE.Votes: 0GitHub stars: 3
- Agad Medical Auc EvalEvaluates a generative anomaly detection model's ability to distinguish normal from abnormal medical images using pseudo-anomaly generation and self-contrast learning. It probes robustness on fine-grained, real-world medical imaging data with limited anomaly supervision. Use when the user wants to benchmark on Alzheimer's Dataset Dubey (2019), ChestXray Kermany et al. (2018), Lung Histopathology (LC25000 subset), Retinal OCT Kermany et al. (2018), or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Afromt EvalThis benchmark evaluates machine translation capabilities across eight morphologically rich African languages translated from English. It specifically probes how well models handle complex morphosyntactic features like noun classification and verb extensions in low-resource settings. Use when the user wants to benchmark on AFROMT, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Afro Nmt EvalEvaluates neural machine translation performance across five low-resource African languages (Swahili, Amharic, Tigrigna, Oromo, Somali) paired with English. It probes model robustness to domain shifts and compares single-language, semi-supervised, transfer-learning, and multilingual training strategies. Use when the user wants to benchmark on AfroNMT, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Afrispeech 200 EvalEvaluates automatic speech recognition (ASR) models on pan-African accented English speech across clinical and general domains. It probes out-of-distribution generalization, zero-shot performance on unseen accents, and domain-specific robustness. Use when the user wants to benchmark on AfriSpeech-200, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Afrisenti Sentiment EvalEvaluates sentiment classification capabilities across 14 low-resource African languages using Twitter data. Tests both monolingual and multilingual transfer, as well as zero-shot adaptation via parameter-efficient fine-tuning. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports weighted F1 score.Votes: 0GitHub stars: 3
- Afrisenti EvalEvaluates multilingual and cross-lingual sentiment classification capabilities on low-resource African languages using Twitter data. It probes how well pre-trained language models handle dialectal variation, code-switching, and mixed scripts in fine-tuning and zero-shot transfer settings. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports F1.Votes: 0GitHub stars: 3
- Afriqa EvalEvaluates cross-lingual open-retrieval question answering systems across 10 African languages. It probes the pipeline's ability to translate low-resource queries, retrieve relevant passages, and accurately extract or generate answers. Use when the user wants to benchmark on AFRIQA, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Afrimteb EvalEvaluates text embedding models on a wide range of African language tasks, including classification, retrieval, semantic similarity, clustering, and bitext mining. It probes cross-lingual transfer, language coverage, and the ability of embeddings to capture semantic and discriminative signals across 59 African languages. Use when the user wants to benchmark on AfriMTEB, AfriMTEB-Lite, or asks about evaluating this task. Reports macro average score.Votes: 0GitHub stars: 3
- Afrimte EvalEvaluates machine translation quality for under-resourced African languages using human-annotated Direct Assessment (DA) scores and error-span annotations. It probes a model's ability to preserve meaning across 13 diverse language pairs. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports Direct Assessment (DA) score.Votes: 0GitHub stars: 3
- African Llm Benchmark EvalThis evaluation probes the cross-lingual reasoning and domain knowledge capabilities of large language models across low-resource African languages. It measures how well models perform on translated benchmarks compared to English, and assesses the impact of cultural appropriateness and fine-tuning data quality on model accuracy. Use when the user wants to benchmark on Winogrande, MMLU (Clinical Sections), Belebele, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Afri Mcqa EvalEvaluates multimodal large language models' ability to answer visual questions about African cultural contexts in both native African languages and English, across text and audio input modalities. Use when the user wants to benchmark on Afri-MCQA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aerr Continuous EvalEvaluates a model's ability to recognize spontaneous apparent emotional reactions from video by predicting continuous arousal and valence dimensions per frame. Use when the user wants to benchmark on SEWA, RECOLA, or asks about evaluating this task. Reports ccc.Votes: 0GitHub stars: 3
- Aeropath Airway Segmentation EvalThis benchmark evaluates 3D medical image segmentation models on contrast-enhanced CT scans containing severe airway pathologies. It probes a model's ability to accurately segment complex, distorted bronchial trees and maintain topological completeness down to small airway generations despite anatomical anomalies like tumors and emphysema. Use when the user wants to benchmark on AeroPath, or asks about evaluating this task. Reports DSC.Votes: 0GitHub stars: 3
- Aerial D Res EvalThis benchmark evaluates a model's ability to perform referring expression segmentation on aerial imagery, testing its capacity to localize objects or regions based on natural language instructions. It specifically probes robustness to domain-specific challenges such as densely packed targets, varying object scales, and simulated historical image degradation (monochrome, sepia, and grainy conditions). Use when the user wants to benchmark on Aerial-D, RRSIS-D, NWPU-Refer, RefSegRS, Urban1960Sa...Votes: 0GitHub stars: 3
- Aeria Edge Ai EvalEvaluates an auction-based dynamic pricing and resource allocation mechanism for on-demand DNN inference at the edge. It probes the system's ability to jointly optimize model partitioning, pricing, and resource distribution under varying user requirements and real-world trace-driven conditions. Use when the user wants to benchmark on Multi30K, ImageNet-1K, CIFAR-100, CIFAR-10, Shanghai Telecom, or asks about evaluating this task. Reports revenue.Votes: 0GitHub stars: 3
- Aepc Qa EvalThis benchmark evaluates large language models' ability to retain and apply specialized domain knowledge in Quebec's civil law insurance sector. It probes both closed-book knowledge retention and the effectiveness of retrieval-augmented generation (RAG) pipelines under jurisdiction-specific, high-stakes regulatory scenarios. Use when the user wants to benchmark on AEPC-QA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Ae Nerf 3d Manipulation EvalEvaluates a model's ability to reconstruct 3D objects from single 2D images and disentangle/manipulate specific 3D attributes (shape, appearance, camera pose) while preserving high visual fidelity. Use when the user wants to benchmark on CARLA, Photoshapes, or asks about evaluating this task. Reports FID.Votes: 0GitHub stars: 3
- Advrace EvalEvaluates the robustness of machine reading comprehension models against four adversarial perturbations (AddSent, CharSwap, Distractor Extraction, and Distractor Generation) applied to passages, questions, and answer options. It measures how much model accuracy degrades when faced with label-preserving but semantically altered inputs compared to clean data. Use when the user wants to benchmark on AdvRACE, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3