All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,625–9,648 of 21,377 skills
- Artseek EvalEvaluates a multimodal retrieval-augmented generation pipeline for deep artwork understanding. It probes the system's ability to retrieve relevant art-historical context from a large corpus, classify artwork attributes (style, genre, artist), and generate grounded, interpretable captions/explanations from image input alone. Use when the user wants to benchmark on WikiFragments, WikiArt/ArtGraph, ArtPedia, SemArt v2.0, PaintingForm, or asks about evaluating this task. Reports NDCG@5, Top-1 Acc...Votes: 0GitHub stars: 3
- Artistmus EvalThis benchmark evaluates the factual accuracy and contextual reasoning capabilities of LLMs in music question answering. It specifically probes how well models can retrieve and utilize artist-centric knowledge from a domain-specific database versus relying on parametric memory, comparing zero-shot, RAG, and reranked retrieval strategies. Use when the user wants to benchmark on ArtistMus, TrustMus, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Artifactbench EvalEvaluates the ability to detect AI-generated music by identifying irreversible residual artifacts from neural audio codecs. It probes robustness across diverse generators, lossy compression codecs, and adversarial source-separation attacks, while measuring false-positive rates on real-world music. Use when the user wants to benchmark on ArtifactBench v1, or asks about evaluating this task. Reports F1.Votes: 0GitHub stars: 3
- Artifact Understanding EvalEvaluates vision-language models on their ability to detect, spatially localize, and explain visual artifacts in AI-generated images. It probes the model's capacity for fine-grained visual reasoning and artifact-aware grounding beyond standard natural image understanding. Use when the user wants to benchmark on ArtiBench, LOKI, or asks about evaluating this task. Reports accuracy, mIoU, ROUGE.Votes: 0GitHub stars: 3
- Artefact EvalEvaluates the ability of segmentation models (CNNs, Transformers, diffusion models, and vision foundation models) to detect and classify diverse damage types on analogue media across different material and content categories. It probes cross-media generalization using a leave-one-out protocol and tests the effectiveness of zero-shot, supervised, and text-guided prompting strategies for pixel-level damage localization. Use when the user wants to benchmark on ARTeFACT, or asks about evaluating ...Votes: 0GitHub stars: 3
- Artbench 10 EvalEvaluates the quality and diversity of synthetic images generated by models across ten distinct artistic styles. It probes a model's ability to capture class-conditional and unconditional data distributions while measuring trade-offs between sample fidelity and variety. Use when the user wants to benchmark on ArtBench-10, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).Votes: 0GitHub stars: 3
- Art Redteam EvalEvaluates the safety vulnerabilities of text-to-image models by measuring how often benign, safe prompts trigger the generation of toxic or unsafe images. It also assesses the diversity and safety of the generated red-teaming prompts themselves. Use when the user wants to benchmark on MSCOCO, or asks about evaluating this task. Reports success ratio under safe prompts (%).Votes: 0GitHub stars: 3
- Art Audio Reasoning EvalThis benchmark evaluates multimodal large language models on their ability to perform cross-modal audio reasoning. It requires models to integrate multiple audio cues (e.g., speech, environmental sounds, speaker identity) and apply logical inference to answer questions, rather than just performing isolated audio tasks like transcription or classification. Use when the user wants to benchmark on ART (Audio Reasoning Tasks), or asks about evaluating this task. Reports Absolute accuracy.Votes: 0GitHub stars: 3
- Armor Pruning EvalEvaluates the effectiveness of semi-structured 2:4 pruning methods on large language models by measuring downstream task accuracy and language modeling perplexity. It specifically tests whether adaptive matrix factorization can preserve model capabilities better than direct weight removal while maintaining inference efficiency. Use when the user wants to benchmark on MMLU, GSM8K, BBH, GPQA, ARC-C, WinoGrande, HellaSwag, Wikitext2, C4, or asks about evaluating this task. Reports Task Accuracy ...Votes: 0GitHub stars: 3
- Armor EvalThis benchmark meta-evaluates objective music evaluation (OE) metrics by measuring how well their similarity scores and classification outputs align with human subjective judgments. It probes whether automated algorithms can reliably capture human perception of musical quality and distinguish human-composed from AI-generated music across diverse genres and generative models. Use when the user wants to benchmark on Armor, or asks about evaluating this task. Reports correlation coefficient.Votes: 0GitHub stars: 3
- Arm Cortex Ai Benchmark EvalThis benchmark evaluates how AI model size, pruning, and quantization affect deployment on bare-metal ARM Cortex-M0+/M4/M7 processors. It specifically probes the trade-offs between energy efficiency, inference latency, and model accuracy across different hardware architectures and application duty cycles. Use when the user wants to benchmark on Embedded AI Use Cases (e.g., Optical Digit Recognition, Visual Wake Words), or asks about evaluating this task. Reports inference cycle energy.Votes: 0GitHub stars: 3
- Arkitscenes EvalEvaluates 3D indoor scene understanding by testing object detection (single-frame and whole-scene) and color-guided depth upsampling on real-world mobile RGB-D data captured with consumer LiDAR devices. Use when the user wants to benchmark on ARKitScenes, or asks about evaluating this task. Reports mAP (mean average precision).Votes: 0GitHub stars: 3
- Arima Fraud Detection EvalEvaluates unsupervised anomaly detection models on credit card transaction time series to identify fraudulent spending deviations. It probes the ability of models to balance precision and recall in highly imbalanced, real-world financial data without relying on labeled fraud examples. Use when the user wants to benchmark on Credit card transaction time series, or asks about evaluating this task. Reports F-Measure.Votes: 0GitHub stars: 3
- Aria Nerf EvalEvaluates NeRF-based models on their ability to synthesize novel views from egocentric, multimodal sensor data captured in dynamic real-world environments. It probes how well current neural rendering methods handle temporal dynamics, lens distortion, and non-visual cues like IMU and gaze. Use when the user wants to benchmark on Aria-NeRF Dataset, or asks about evaluating this task. Reports PSNR.Votes: 0GitHub stars: 3
- Argoverse Shift EvalEvaluates the out-of-distribution generalization capability of trajectory prediction models on unseen HD map geometries. It measures how well models maintain prediction accuracy when transferred from seen to unseen domains without retraining. Use when the user wants to benchmark on argoverse-shift, or asks about evaluating this task. Reports minADE.Votes: 0GitHub stars: 3
- Arfake EvalEvaluates the ability of spoof-speech detection models to distinguish between real (bonafide) Arabic speech and synthetic speech generated by various Text-to-Speech (TTS) models across multiple Arabic dialects. It probes robustness against different voice-cloning systems and measures detection accuracy alongside perceptual realism and ASR fidelity. Use when the user wants to benchmark on ArFake, or asks about evaluating this task. Reports Equal Error Rate (EER).Votes: 0GitHub stars: 3
- Ares Safety EvalThis evaluation protocol assesses the safety alignment and general capability of large language models after undergoing an adaptive red-teaming and repair process. It probes the model's ability to refuse harmful or unsafe prompts while maintaining performance on standard knowledge and reasoning benchmarks, and measures the false refusal rate to ensure utility is preserved. Use when the user wants to benchmark on RedTeam, StrongReject, HarmBench, PKU-SafeRLHF, XSTest, MMLU, GSM8K, TruthfulQA, ...Votes: 0GitHub stars: 3
- Ares Android Testing EvalEvaluates the ability of automated black-box testing tools and reinforcement learning agents to explore Android applications effectively. It probes how well algorithms navigate complex UI states, maximize code/activity coverage, and trigger unique application crashes within a fixed time budget. Use when the user wants to benchmark on F-Droid top starred apps, AndroTest, Synthetic FATE models, or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Arctic Extract EvalEvaluates a multimodal document understanding model across visual QA, multilingual text comprehension, English reading comprehension, and table extraction tasks. It probes the model's ability to process long-context documents, answer questions from images or text, and extract structured tabular data from unstructured layouts. Use when the user wants to benchmark on SQuAD2.0, DocVQA, Arctic-TILT, MLQA, xQuAD, or asks about evaluating this task. Reports ANLS*.Votes: 0GitHub stars: 3
- Arcdeck EvalEvaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references. Use when the user wants to benchmark on ArcBench, or asks about evaluating this task. Reports VLM-based Q/A Quiz Accuracy.Votes: 0GitHub stars: 3
- Arc EvalMeasures general fluid intelligence and developer-aware generalization by requiring systems to infer abstract transformation rules from few input-output grid demonstrations and apply them to novel test cases. It explicitly avoids measuring task-specific memorization or crystallized knowledge, focusing instead on abstraction, reasoning, and broad generalization under strict prior constraints. Use when the user wants to benchmark on Abstraction and Reasoning Corpus (ARC), or asks about evaluati...Votes: 0GitHub stars: 3
- Arbitrary View Action EvalEvaluates human action recognition models under varying camera viewpoints and different subjects. It probes the model's ability to generalize across unseen subjects, unseen camera angles, and continuous 360-degree view changes. Use when the user wants to benchmark on Varying-view RGB-D Action Dataset, or asks about evaluating this task. Reports average recognition accuracy.Votes: 0GitHub stars: 3
- Arbibench EvalEvaluates the adversarial robustness of binarized neural networks (BNNs) against white-box and black-box attacks. It measures how well BNNs maintain prediction accuracy under controlled perturbation budgets compared to their clean accuracy, highlighting robustness trends across dataset scales. Use when the user wants to benchmark on CIFAR-10, ImageNet, or asks about evaluating this task. Reports ACC_norm.Votes: 0GitHub stars: 3
- Aratable EvalThis benchmark probes large language models' ability to reason over and understand Arabic tabular data across three core tasks: direct question answering, fact verification, and complex reasoning. It specifically tests whether models can extract, compare, and synthesize information from structured Arabic tables while adhering to linguistic and formatting nuances. Use when the user wants to benchmark on AraTable, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3