All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,673–9,696 of 21,377 skills
- Apex EvalEvaluates models on complex, real-world professional reasoning tasks across four domains (investment banking, management consulting, big law, primary care). It probes document analysis, multi-step reasoning, and domain-specific judgment under practical constraints. Use when the user wants to benchmark on APEX-v1.0, or asks about evaluating this task. Reports autograded scores.Votes: 0GitHub stars: 3
- Apeach EvalEvaluates the ability of NLP models to detect hate speech in Korean text. It specifically probes domain-agnostic generalizability and resistance to common inductive biases like text length or topic distribution. Use when the user wants to benchmark on APEACH, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Ape Prompt EvalEvaluates the effectiveness of automatically generated prompts (instructions) from the APE framework compared to human-designed or baseline prompts across various natural language processing tasks. Use when the user wants to benchmark on Instruction Induction, BIG-Bench Instruction Induction (BBII), MultiArith, GSM8K, or asks about evaluating this task. Reports zero-shot execution accuracy.Votes: 0GitHub stars: 3
- Ape EvalEvaluates automatic post-editing (APE) models by measuring how effectively they correct machine-translated text to align with human references. It probes the system's ability to fix translation artifacts, preserve source semantics, and adapt to different domains and translation technologies. Use when the user wants to benchmark on WMT'18 SMT, SubEdits, MLQE-PE, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Anything To Audio EvalEvaluates a unified diffusion transformer model's ability to generate high-fidelity audio and music conditioned on diverse modalities (text, video, image, audio). It measures acoustic similarity, generation quality/diversity, and cross-modal semantic alignment across multiple standard audio generation benchmarks. Use when the user wants to benchmark on AudioCaps, VGGSound, AVVP, MusicCaps, V2M-bench, or asks about evaluating this task. Reports FAD.Votes: 0GitHub stars: 3
- Anytext Benchmark EvalEvaluates the ability of text-to-image models to accurately render specified multilingual text (English and Chinese) in arbitrary shapes and positions while maintaining visual realism and seamless background integration. Use when the user wants to benchmark on AnyText-benchmark, or asks about evaluating this task. Reports Sen. ACC.Votes: 0GitHub stars: 3
- Anyedit EvalEvaluates the ability of image editing models to follow natural language instructions to modify images while preserving unedited regions and maintaining semantic/visual consistency. It probes alignment with complex editing intents, content preservation, and robustness across diverse editing types including implicit and visual-conditioned tasks. Use when the user wants to benchmark on Emu Edit Test, MagicBrush, AnyEdit-Test, or asks about evaluating this task. Reports CLIPim.Votes: 0GitHub stars: 3
- Anticancer Drug Response EvalEvaluates a model's ability to predict anticancer drug responses (IC50 scores) between drugs and cell lines. It probes the model's capacity to perform weighted link prediction/regression on a multimodal graph combining drug and cell line similarities. Use when the user wants to benchmark on CCLE Dataset, or asks about evaluating this task. Reports MSE.Votes: 0GitHub stars: 3
- Antibody Domainbed EvalEvaluates out-of-distribution generalization of protein language models and sequence CNNs for therapeutic antibody design across different antigen targets and generative models. It probes robustness to covariate shifts, label shifts, and assay biases in molecular sequence data. Use when the user wants to benchmark on Antibody DomainBed, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Answersumm EvalEvaluates multi-perspective answer summarization for community question answering, probing content selection, perspective clustering, abstractive summarization, and factual consistency/coverage. Use when the user wants to benchmark on AnswerSumm, or asks about evaluating this task. Reports F1, ROUGE-1/2/L.Votes: 0GitHub stars: 3
- Answer Switching RateMeasures the causal influence of activation-based interventions (linear directions or multidimensional cones) on an LLM's factual reasoning. It quantifies how effectively steering or ablating specific neural subspaces switches model outputs from truthful to untruthful across a set of propositional prompts. Use when the user has predictions and gold and needs to compute Answer Switching Rate (ASR).Votes: 0GitHub stars: 3
- Answer Leakage Robustness EvalThis benchmark evaluates the robustness of LLM-based tutoring models against adversarial student agents designed to elicit final answers. It measures how easily tutors disclose solutions under various attack strategies and tracks the dialogue length required for answer leakage across math, multiple-choice, and coding domains. Use when the user wants to benchmark on GSM8K, MMLU, HumanEval, or asks about evaluating this task. Reports answer leakage rate.Votes: 0GitHub stars: 3
- Anomalymatch EvalThis evaluation probes a semi-supervised anomaly detection model's ability to identify rare or visually distinct objects in highly imbalanced image datasets. It measures how effectively the model ranks anomalies at the top of its predictions using limited initial labels and iterative active learning cycles. Use when the user wants to benchmark on miniImageNet, GalaxyMNIST, Galaxy Zoo 2 (Kaggle Challenge subset), or asks about evaluating this task. Reports AUROC.Votes: 0GitHub stars: 3
- Anomalygen EvalThis benchmark evaluates log-based anomaly detection models by measuring their ability to classify log sequences as normal or anomalous. It specifically probes how well different model paradigms (classical ML, supervised/unsupervised deep learning, and LLM-based) generalize when trained on code-guided synthetic data augmentation across varying augmentation ratios. Use when the user wants to benchmark on HDFS, Zookeeper, or asks about evaluating this task. Reports F1-score.Votes: 0GitHub stars: 3
- Anomaly Detection Meta EvalEvaluates anomaly detection algorithms on a large corpus of synthetic datasets systematically varied along four dimensions: point difficulty, semantic variation, relative frequency, and feature relevance. It probes algorithm robustness, generalization across diverse anomaly-generating processes, and the impact of experimental design choices on reported performance. Use when the user wants to benchmark on Synthetic Anomaly Detection Corpus, or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Anomaly Detection EvalEvaluates unsupervised and semi-supervised time series anomaly detection pipelines across multiple real-world and benchmark datasets. It measures how well different models identify known anomalous segments in telemetry, production traffic, and synthetic signals. Use when the user wants to benchmark on NAB, NASA, YAHOO, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Anomaly Detection Benchmark EvalThis benchmark evaluates the detection accuracy and computational efficiency of classical machine learning, tree-based, and deep learning anomaly detection algorithms across diverse multivariate and univariate datasets. It probes how well different models handle class imbalance, varying anomaly prevalence, and differing requirements for labeled anomaly data during training. The evaluation also measures training time and resource consumption to assess real-world deployment feasibility. Use whe...Votes: 0GitHub stars: 3
- Annbatch Data Loading EvalMeasures the data loading throughput and epoch iteration time for large-scale biological datasets. It benchmarks how efficiently a loader can fetch and prepare mini-batches from disk compared to existing frameworks. Use when the user wants to benchmark on Tahoe100M, 1000 Genomes GRCh38, Single-cell microscopy images, or asks about evaluating this task. Reports samples/sec.Votes: 0GitHub stars: 3
- Ann Benchmarks EvalThis benchmark evaluates approximate nearest neighbor (ANN) search algorithms by measuring the trade-off between search quality (recall) and computational efficiency (queries per second, index size, and build time). It probes how well different algorithmic families perform across diverse high-dimensional datasets and distance metrics, revealing robustness and approximation capabilities. Use when the user wants to benchmark on SIFT, GIST, GLOVE, NYTimes, Rand-Euclidean, SIFT-Hamming, Word2Bits...Votes: 0GitHub stars: 3
- Animationbench EvalThis benchmark evaluates video generation models on character-centric animation capabilities, specifically probing IP preservation, motion expressiveness, deformation accuracy, and multi-angle consistency. It operationalizes animation principles into measurable dimensions to identify gaps missed by standard realism-focused benchmarks. Use when the user wants to benchmark on AnimationBench, or asks about evaluating this task. Reports AnimationBench score.Votes: 0GitHub stars: 3
- Animal3d EvalThis benchmark evaluates the ability of deep learning models to estimate 3D pose and shape of diverse mammal species from single images. It probes cross-species generalization, synthetic-to-real transfer, and the adaptation of human-centric pose estimation architectures to non-human anatomies. Use when the user wants to benchmark on Animal3D, or asks about evaluating this task. Reports S-MPJPE.Votes: 0GitHub stars: 3
- Anim 400k EvalEvaluates automated end-to-end video dubbing systems by testing their ability to generate synchronized English audio from Japanese source video, specifically probing prosody matching, timing alignment, and multi-speaker isolation capabilities. Use when the user wants to benchmark on Anim-400K, or asks about evaluating this task. Reports MUSHRA.Votes: 0GitHub stars: 3
- Anetqa EvalEvaluates fine-grained compositional reasoning over untrimmed videos by requiring models to interpret spatio-temporal scene graphs and answer complex questions involving attributes, actions, and temporal relationships. Use when the user wants to benchmark on ANetQA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Androidworld Generalization EvalProbes the zero-shot generalization capability of mobile agents trained via online reinforcement learning across increasingly challenging unseen scenarios in Android environments, including new task instances, UI templates, and entirely new applications. It measures how well learned interaction policies transfer to novel contexts without additional supervised fine-tuning. Use when the user wants to benchmark on AndroidWorld-Generalization, or asks about evaluating this task. Reports Success R...Votes: 0GitHub stars: 3