Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,585–7,608 of 20,843 skills

Facial Emotion Analysis EvalA

Evaluates vision-language models on facial emotion analysis tasks, including fine-grained action unit detection, categorical emotion recognition, and grounded natural language reasoning over facial expressions. The protocol tests both recognition accuracy and the model's ability to generate interpretable, AU-grounded explanations. Use when the user wants to benchmark on DISFA, BP4D, RAF-AU, FER2013, AffectNet, RAF-DB, FABA-Instruct, FEA-20K, or asks about evaluating this task. Reports F1 scor...

researchpythongo
0
3
Facial Attribute Prediction EvalA

Tests the model's capability to predict multiple facial attributes (e.g., gender, hairstyle) from a single facial image as a multilabel classification task. It probes fine-grained visual feature extraction and attribute-level alignment. Use when the user wants to benchmark on CelebA, LFWA, or asks about evaluating this task. Reports Average Precision (AP).

researchpythongo
0
3
Facet Fairness EvalA

This benchmark probes the intersectional fairness of computer vision models by evaluating their performance across diverse demographic attributes (e.g., skin tone, gender presentation, hair type) and person-related categories (e.g., occupations, hobbies). It measures whether models exhibit systematic performance disparities when detecting, classifying, or segmenting individuals with different attribute combinations. Use when the user wants to benchmark on FACET, or asks about evaluating this ...

researchpythongo
0
3
Facescape EvalA

Evaluates single-view 3D face reconstruction models on their ability to predict high-fidelity, expression-specific dynamic details (displacement maps) and generate riggable 3D face meshes from a single image. It probes geometric accuracy, detail synthesis, and generalization across multiple expressions and real-world sequences. Use when the user wants to benchmark on FaceScape, Volker Sequence, or asks about evaluating this task. Reports mean absolute point-to-surface distance.

researchpythonexpress
0
3
Faceid 6m EvalA

This evaluation protocol assesses the effectiveness of a large-scale dataset for training conditional diffusion models in FaceID customization. It probes the model's ability to preserve facial identity from a reference image while adhering to textual prompts and maintaining overall image quality. Use when the user wants to benchmark on COCO2017, Unsplash-50, or asks about evaluating this task. Reports Face Sim.

researchpythongo
0
3
Facebook Hateful Memes EvalA

This benchmark evaluates multimodal hate speech detection by classifying image-caption pairs (memes) as hateful or non-hateful. It probes a model's ability to align visual and textual cues while resisting spurious correlations, particularly under different prompt structures and data augmentation strategies. Use when the user wants to benchmark on Facebook Hateful Memes, or asks about evaluating this task. Reports weighted-F1 score.

researchpythongo
0
3
Facebehaviornet EvalA

Evaluates a multi-task facial analysis model's ability to jointly predict continuous affect (valence/arousal), discrete facial expressions, and facial action units from in-the-wild and lab-controlled face images. It probes the model's capacity for task-coupled learning and generalization across heterogeneous annotation schemes and domains. Use when the user wants to benchmark on Aff-Wild, AffectNet, AFEW, RAF-DB, EmotioNet, DISFA, BP4D, BP4D+, or asks about evaluating this task. Reports CCC.

researchpythongo
0
3
Face Mtl EvalA

Evaluates a multi-task learning framework for face analysis across heterogeneous tasks including valence-arousal estimation, action unit detection, expression classification, and face recognition. It probes the model's ability to jointly learn from diverse, in-the-wild and lab-controlled facial datasets while mitigating negative transfer through distribution matching. Use when the user wants to benchmark on Aff-Wild, AffectNet, RAF-DB, DISFA, GFT, BP4D, CelebA, or asks about evaluating this t...

researchpythongo
0
3
Fabl EvalA

Evaluates a joint learning framework (FABL) for real-time human behavior recognition using 3D skeletal data from depth sensors. It tests the method's ability to simultaneously select discriminative body parts and features for action classification across public benchmarks and a custom robot-interaction task. Use when the user wants to benchmark on MSR Action3D Dataset, Cornell Activity Dataset 60 (CAD-60), Baxter Robot Serving Drinks Task, or asks about evaluating this task. Reports average r...

researchpythontesting
0
3
F Siol 310 EvalA

Evaluates few-shot incremental learning (FSIL) capabilities in robotic vision, specifically testing a model's ability to learn new object classes sequentially with very limited examples (5 or 10 per class) while resisting catastrophic forgetting of previously learned classes. Use when the user wants to benchmark on F-SIOL-310, or asks about evaluating this task. Reports classification accuracy (%).

researchpythongo
0
3
Ezsql Sql To Text EvalA

Evaluates a model's capability to generate fluent natural language descriptions from SQL queries (SQL-to-text) and measures how well the generated text can augment training data for Text-to-SQL parsers. Use when the user wants to benchmark on WikiSQL, Spider, or asks about evaluating this task. Reports BLEU-4.

researchpythongo
0
3
Extremenerf EvalA

Evaluates few-shot novel view synthesis and depth estimation under unconstrained, varying illumination. It probes a model's ability to maintain geometric consistency and produce photorealistic images when trained on only a few sparse views with different lighting conditions. Use when the user wants to benchmark on Phototourism F^3, NeRF Extreme, LLFF, or asks about evaluating this task. Reports SSIM.

researchpythonperformance
0
3
Extractbench EvalA

Evaluates LLMs on complex, schema-driven PDF-to-JSON structured extraction, testing their ability to handle nested objects, arrays, heterogeneous field types, and strict correctness criteria across enterprise-scale documents. Use when the user wants to benchmark on ExtractBench, or asks about evaluating this task. Reports Pass Rate.

researchpythongo
0
3
Extra Cot EvalA

Evaluates the ability of large language models to generate mathematically reasoned chain-of-thought outputs that are compressed to a target token budget while preserving logical fidelity and answer accuracy. Use when the user wants to benchmark on GSM8K, MATH-500, AMC2023, or asks about evaluating this task. Reports Acc@all.

researchpythongo
0
3
Expressive Voice Cloning EvalA

Evaluates zero-shot and adapted voice cloning models on their ability to preserve speaker identity, transfer expressive style (pitch and rhythm), and generate natural-sounding speech. It measures speaker similarity, style fidelity, and perceptual quality across text-to-speech, imitation, and style transfer tasks. Use when the user wants to benchmark on VCTK, Libri-TTS, or asks about evaluating this task. Reports Speaker Classification Accuracy.

researchpythonexpress
0
3
Express Emotion Recognition EvalA

Evaluates language models' ability to recognize and decompose fine-grained human emotions from self-disclosed narratives. It probes whether models can align with human emotional expressions across 10 Plutchik-based dimensions (8 basic emotions + 2 sentiments) rather than just predicting surface-level emotion words. Use when the user wants to benchmark on EXPRESS, or asks about evaluating this task. Reports F1_V.

researchpythonrust
0
3
Express Bench EvalA

Evaluates an agent's ability to actively explore 3D environments to gather visual evidence and answer questions accurately, while measuring exploration efficiency and navigation performance. It specifically probes whether the agent's final answer is grounded in the actual visual observations collected during its exploration path, detecting hallucinations and ungrounded reasoning. Use when the user wants to benchmark on EXPRESS-Bench, or asks about evaluating this task. Reports C.

researchpythongo
0
3
Explanation Disagreement EvalA

Evaluates the consistency of post-hoc explanation methods by quantifying how much feature attributions differ across algorithms for identical model predictions. It probes whether local explanations are reliable and whether practitioners have principled ways to resolve conflicts when different methods yield conflicting importance scores. Use when the user wants to benchmark on COMPAS, German Credit, News text dataset, PASCAL VOC 2012, or asks about evaluating this task. Reports L2 distance of ...

researchpythongo
0
3
Explainseg Segmentation EvalA

Evaluates the ability of an XAI-driven classification model to generate clinically meaningful segmentation masks without pixel-level annotations. It probes spatial coherence, boundary precision, and generalization across diverse medical imaging modalities (mammography, histopathology, endoscopy). Use when the user wants to benchmark on CBIS-DDSM, NuInsSeg, Kvasir-SEG, or asks about evaluating this task. Reports Dice.

researchpythonperformance
0
3
Explaincpe EvalA

This benchmark evaluates large language models on Chinese medical multiple-choice questions, specifically probing their ability to select correct answers and generate faithful, logically consistent free-text explanations. It measures both factual accuracy and the quality of interpretability in high-stakes healthcare domains. Use when the user wants to benchmark on ExplainCPE, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Exp Bench EvalA

Evaluates AI agents' end-to-end capability to conduct real AI research experiments, including designing methodologies, implementing code, executing experiments, and drawing conclusions. Use when the user wants to benchmark on EXP-Bench, or asks about evaluating this task. Reports All·E✓.

researchpythondocker
0
3
Exoplanet Vit Temporal EvalA

Evaluates a Vision Transformer's ability to classify exoplanet transits by processing temporal light curve data transformed into image representations (Recurrence Plots and Gramian Angular Fields). It probes the model's capacity to capture long-range temporal dependencies and handle class imbalance in astronomical time-series data. Use when the user wants to benchmark on Kepler Light Curve Exoplanet Candidates, or asks about evaluating this task. Reports F1-score.

researchpythonangular
0
3
Exoplanet Spectra Model BenchmarkA

Evaluates the consistency of 1D radiative-convective atmospheric models in predicting exoplanet emission and transmission spectra, specifically probing how differences in opacity linelists, line shape treatments, and chemical equilibrium assumptions affect spectral predictions relative to JWST observational uncertainties. Use when the user wants to benchmark on Exoplanet Atmospheric Test Cases, or asks about evaluating this task. Reports spectral_resolution_and_jwst_error_bar_comparison.

researchpython
0
3
Exoplanet Imaging Challenge Ii EvalA

This benchmark evaluates the ability of high-contrast imaging post-processing pipelines to accurately estimate the astrometric position of injected exoplanet signals in multispectral astronomical data. It probes how well algorithms handle varying signal-to-noise ratios, complex residual backgrounds (e.g., diffraction patterns, coronagraphic inner working angles), and different observing conditions. Use when the user wants to benchmark on Exoplanet Imaging Data Challenge Phase II, or asks abou...

researchpythongo
0
3