All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,769–9,792 of 21,377 skills
- Ai4arctic Sod EvalThis benchmark evaluates pixel-wise sea ice stage of development (SOD) segmentation using dual-polarized SAR imagery. It probes a model's ability to accurately classify ice types under varying quantization levels and measures hardware efficiency across different computing platforms. Use when the user wants to benchmark on AI4Arctic Sea Ice Dataset, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Ai Writing Assistance EvalEvaluates how source disclosure and perceived AI authorship influence human editing behavior and subsequent peer-review acceptance decisions for scientific abstracts. Use when the user wants to benchmark on CS-Conference-Abstracts, or asks about evaluating this task. Reports accept/reject decision.Votes: 0GitHub stars: 3
- Ai Review Detection EvalEvaluates a style-based classifier's ability to detect AI-generated text in academic peer reviews and measures temporal generalization by tracking detection rates across consecutive years. Use when the user wants to benchmark on ICLR Peer Reviews, Nature Communications Peer Reviews, or asks about evaluating this task. Reports percentage_ai_detected.Votes: 0GitHub stars: 3
- Ai Red Teaming Ctf EvalEvaluates adversarial AI red-teaming capabilities by measuring participant success rates in bypassing LLM guardrails, manipulating model outputs, and extracting sensitive data through prompt injection and jailbreaking techniques. Use when the user wants to benchmark on AI Red Teaming CTF (CTF ID: 2604), or asks about evaluating this task. Reports solve_rate.Votes: 0GitHub stars: 3
- Ai QualityEvaluates a feature-hierarchical edge inference framework's ability to dynamically allocate communication and computation resources to maximize AI quality under strict latency and energy constraints. Use when the user has predictions and gold and needs to compute AI quality (mAP).Votes: 0GitHub stars: 3
- Ai Paper Error Audit EvalEvaluates an LLM-based auditing system's ability to detect, categorize, and quantify objective mistakes in published AI research papers. It measures the system's precision against human verification and its recall against injected ground-truth errors across mathematical, textual, tabular, and cross-reference categories. Use when the user wants to benchmark on Published AI Papers (ICLR, NeurIPS, TMLR), or asks about evaluating this task. Reports precision.Votes: 0GitHub stars: 3
- Ai Genbench EvalEvaluates the ability of AI-generated image detectors to generalize to novel, temporally subsequent generative models under realistic post-processing conditions. It measures how well detectors maintain performance when incrementally trained on historically ordered synthetic data and tested on unseen future generators. Use when the user wants to benchmark on AI-GenBench, or asks about evaluating this task. Reports AUROC, Accuracy.Votes: 0GitHub stars: 3
- Ai Face Fairness Bench EvalEvaluates the fairness and utility of AI-generated face detectors across demographic attributes (skin tone, gender, age) and intersectional groups. It measures how well detectors distinguish real vs. AI-generated faces while ensuring equitable performance across demographic subgroups. Use when the user wants to benchmark on AI-Face, or asks about evaluating this task. Reports $F_{MEO}$.Votes: 0GitHub stars: 3
- Ai Accelerator Training EvalEvaluates the computational performance and energy efficiency of various AI accelerators (CPUs, GPUs, TPUs) across standard deep learning workloads, including CNNs and NLP models. It measures how hardware architecture, numerical precision, and batch size impact training throughput and power consumption. Use when the user wants to benchmark on Standard DNN Workloads (ResNet50, Inception v3, Vgg16, LSTM, Deep Speech 2, Transformer), or asks about evaluating this task. Reports throughput.Votes: 0GitHub stars: 3
- Ahup 3d Pose EvalEvaluates monocular 3D human pose estimation models trained exclusively on synthetic 3D data and real 2D images, testing their ability to generalize to real-world 3D pose benchmarks without using any real 3D pose annotations during training. It probes domain adaptation capabilities, cross-dataset generalization, and the effectiveness of skeletal pose alignment strategies. Use when the user wants to benchmark on Human3.6M, MuPoTS, SURREAL, ScanAva+, MSCOCO, MPII Human Pose, or asks about evalu...Votes: 0GitHub stars: 3
- Ahat Planning EvalEvaluates an agent's ability to decompose abstract, long-horizon household instructions into feasible, constraint-satisfying action plans. It probes intent inference, subgoal grounding, and robustness to environmental clutter and instruction ambiguity. Use when the user wants to benchmark on AHAT, Human Tasks, PARTNR, Behavior-1K, or asks about evaluating this task. Reports Success Rate (SR).Votes: 0GitHub stars: 3
- Agrigpt Omni EvalEvaluates a unified speech-vision-text model's capability in multilingual agricultural reasoning, covering text generation, vision-language QA, and multimodal speech understanding across open-ended and multiple-choice formats. Use when the user wants to benchmark on AgriBench-13K, AgriBench-VL-4K, AgriBench-Omni-2K, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Agriculture Asr EvalEvaluates Automatic Speech Recognition (ASR) models on real-world agricultural field recordings across three Indian languages (Hindi, Telugu, Odia). It probes the models' ability to transcribe domain-specific terminology under challenging acoustic conditions like wind noise and multi-speaker overlap. Use when the user wants to benchmark on Agricultural Field Recordings, or asks about evaluating this task. Reports AWWER.Votes: 0GitHub stars: 3
- Agnn Citation EvalEvaluates semi-supervised node classification on citation networks using an attention-based graph neural network. Tests performance under fixed benchmark splits, random node sampling, and larger training sets to measure classification accuracy and attention interpretability. Use when the user wants to benchmark on CiteSeer, Cora, PubMed, or asks about evaluating this task. Reports classification accuracy.Votes: 0GitHub stars: 3
- Agieval EvalThis benchmark evaluates foundation models on human-level cognitive abilities and general reasoning by testing them on a diverse collection of standardized admission and qualification exams. It probes domain-specific knowledge, analytical reasoning, and problem-solving across subjects like mathematics, law, logic, and languages. Use when the user wants to benchmark on AGIEval, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aghi Qa EvalEvaluates the perceptual quality and text-image correspondence of AI-generated human images, while also benchmarking the ability of models to identify visible and semantically distorted human body parts. Use when the user wants to benchmark on AGHI-QA, or asks about evaluating this task. Reports SRCC.Votes: 0GitHub stars: 3
- Agentvista EvalEvaluates the ability of multimodal agents to perform long-horizon, multi-step tool use in complex, realistic visual environments. It probes cross-image reasoning, constraint tracking, and robust grounding when interacting with dynamic tools like web search and code execution. Use when the user wants to benchmark on AgentVista, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Agentsynth EvalEvaluates the ability of multimodal language models to execute long-horizon, multi-step computer-use tasks on a desktop environment. It probes visual grounding, precise GUI interaction, state tracking, and error recovery across varying task complexities and software domains. Use when the user wants to benchmark on AgentSynth, or asks about evaluating this task. Reports success rate.Votes: 0GitHub stars: 3
- Agentsafe EvalEvaluates the safety of embodied vision-language model agents when executing hazardous instructions in simulated indoor environments. It probes four key capabilities across perception, planning, and execution stages: object recognition accuracy, refusal to plan harmful actions, success in generating harmful plans, and success in physically executing them under both direct and jailbroken conditions. Use when the user wants to benchmark on AGENTSAFE, or asks about evaluating this task. Reports ...Votes: 0GitHub stars: 3
- Agentrewardbench EvalThis benchmark evaluates the effectiveness of LLM-based judges in automatically assessing web agent trajectories. It probes the judges' ability to correctly predict task success, detect side effects, and identify repetitive actions by comparing their outputs against expert human annotations. Use when the user wants to benchmark on AgentRewardBench, or asks about evaluating this task. Reports precision.Votes: 0GitHub stars: 3
- Agentrecbench EvalThis benchmark evaluates LLM-based agentic recommender systems across three scenarios: classic, evolving-interest, and cold-start recommendation. It probes the agents' ability to dynamically plan, utilize textual interaction environments, and adapt to user preference shifts or data sparsity using structured user/item profiles and reviews. Use when the user wants to benchmark on Amazon, GoodReads, Yelp, or asks about evaluating this task. Reports Hit Rate@$N.Votes: 0GitHub stars: 3
- Agentquest EvalThis evaluation protocol measures LLM agent performance on multi-step reasoning tasks by tracking step-wise progress toward goal completion and the frequency of repetitive actions or states. It enables fine-grained debugging and architectural refinement beyond simple pass/fail success rates. Use when the user wants to benchmark on ALFWorld, Sudoku, or asks about evaluating this task. Reports progress rate.Votes: 0GitHub stars: 3
- AgentprmevalEvaluates LLM agents' ability to navigate simulated environments and execute multi-step plans to complete natural language instructions. It probes step-wise decision-making, goal proximity tracking, and sequential task execution across web shopping, grid-world navigation, and text-based crafting scenarios. Use when the user wants to benchmark on WebShop, BabyAI, TextCraft, or asks about evaluating this task. Reports success rate.Votes: 0GitHub stars: 3
- Agentfuel EvalEvaluates LLM-based data analysis agents on their ability to execute domain-specific time-series queries, particularly focusing on stateful logic, temporal dependencies, and incident pattern detection. It probes whether agents can correctly interpret schemas, track state across sequential events, and identify anomalous behavior without relying on predefined time windows. Use when the user wants to benchmark on AgentFuel Benchmark, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3