Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,585–7,608 of 20,843 skills
Evaluates vision-language models on facial emotion analysis tasks, including fine-grained action unit detection, categorical emotion recognition, and grounded natural language reasoning over facial expressions. The protocol tests both recognition accuracy and the model's ability to generate interpretable, AU-grounded explanations. Use when the user wants to benchmark on DISFA, BP4D, RAF-AU, FER2013, AffectNet, RAF-DB, FABA-Instruct, FEA-20K, or asks about evaluating this task. Reports F1 scor...
Tests the model's capability to predict multiple facial attributes (e.g., gender, hairstyle) from a single facial image as a multilabel classification task. It probes fine-grained visual feature extraction and attribute-level alignment. Use when the user wants to benchmark on CelebA, LFWA, or asks about evaluating this task. Reports Average Precision (AP).
This benchmark probes the intersectional fairness of computer vision models by evaluating their performance across diverse demographic attributes (e.g., skin tone, gender presentation, hair type) and person-related categories (e.g., occupations, hobbies). It measures whether models exhibit systematic performance disparities when detecting, classifying, or segmenting individuals with different attribute combinations. Use when the user wants to benchmark on FACET, or asks about evaluating this ...
Evaluates single-view 3D face reconstruction models on their ability to predict high-fidelity, expression-specific dynamic details (displacement maps) and generate riggable 3D face meshes from a single image. It probes geometric accuracy, detail synthesis, and generalization across multiple expressions and real-world sequences. Use when the user wants to benchmark on FaceScape, Volker Sequence, or asks about evaluating this task. Reports mean absolute point-to-surface distance.
This evaluation protocol assesses the effectiveness of a large-scale dataset for training conditional diffusion models in FaceID customization. It probes the model's ability to preserve facial identity from a reference image while adhering to textual prompts and maintaining overall image quality. Use when the user wants to benchmark on COCO2017, Unsplash-50, or asks about evaluating this task. Reports Face Sim.
This benchmark evaluates multimodal hate speech detection by classifying image-caption pairs (memes) as hateful or non-hateful. It probes a model's ability to align visual and textual cues while resisting spurious correlations, particularly under different prompt structures and data augmentation strategies. Use when the user wants to benchmark on Facebook Hateful Memes, or asks about evaluating this task. Reports weighted-F1 score.
Evaluates a multi-task facial analysis model's ability to jointly predict continuous affect (valence/arousal), discrete facial expressions, and facial action units from in-the-wild and lab-controlled face images. It probes the model's capacity for task-coupled learning and generalization across heterogeneous annotation schemes and domains. Use when the user wants to benchmark on Aff-Wild, AffectNet, AFEW, RAF-DB, EmotioNet, DISFA, BP4D, BP4D+, or asks about evaluating this task. Reports CCC.
Evaluates a multi-task learning framework for face analysis across heterogeneous tasks including valence-arousal estimation, action unit detection, expression classification, and face recognition. It probes the model's ability to jointly learn from diverse, in-the-wild and lab-controlled facial datasets while mitigating negative transfer through distribution matching. Use when the user wants to benchmark on Aff-Wild, AffectNet, RAF-DB, DISFA, GFT, BP4D, CelebA, or asks about evaluating this t...
Evaluates a joint learning framework (FABL) for real-time human behavior recognition using 3D skeletal data from depth sensors. It tests the method's ability to simultaneously select discriminative body parts and features for action classification across public benchmarks and a custom robot-interaction task. Use when the user wants to benchmark on MSR Action3D Dataset, Cornell Activity Dataset 60 (CAD-60), Baxter Robot Serving Drinks Task, or asks about evaluating this task. Reports average r...
Evaluates few-shot incremental learning (FSIL) capabilities in robotic vision, specifically testing a model's ability to learn new object classes sequentially with very limited examples (5 or 10 per class) while resisting catastrophic forgetting of previously learned classes. Use when the user wants to benchmark on F-SIOL-310, or asks about evaluating this task. Reports classification accuracy (%).
Evaluates a model's capability to generate fluent natural language descriptions from SQL queries (SQL-to-text) and measures how well the generated text can augment training data for Text-to-SQL parsers. Use when the user wants to benchmark on WikiSQL, Spider, or asks about evaluating this task. Reports BLEU-4.
Evaluates few-shot novel view synthesis and depth estimation under unconstrained, varying illumination. It probes a model's ability to maintain geometric consistency and produce photorealistic images when trained on only a few sparse views with different lighting conditions. Use when the user wants to benchmark on Phototourism F^3, NeRF Extreme, LLFF, or asks about evaluating this task. Reports SSIM.
Evaluates LLMs on complex, schema-driven PDF-to-JSON structured extraction, testing their ability to handle nested objects, arrays, heterogeneous field types, and strict correctness criteria across enterprise-scale documents. Use when the user wants to benchmark on ExtractBench, or asks about evaluating this task. Reports Pass Rate.
Evaluates the ability of large language models to generate mathematically reasoned chain-of-thought outputs that are compressed to a target token budget while preserving logical fidelity and answer accuracy. Use when the user wants to benchmark on GSM8K, MATH-500, AMC2023, or asks about evaluating this task. Reports Acc@all.
Evaluates zero-shot and adapted voice cloning models on their ability to preserve speaker identity, transfer expressive style (pitch and rhythm), and generate natural-sounding speech. It measures speaker similarity, style fidelity, and perceptual quality across text-to-speech, imitation, and style transfer tasks. Use when the user wants to benchmark on VCTK, Libri-TTS, or asks about evaluating this task. Reports Speaker Classification Accuracy.
Evaluates language models' ability to recognize and decompose fine-grained human emotions from self-disclosed narratives. It probes whether models can align with human emotional expressions across 10 Plutchik-based dimensions (8 basic emotions + 2 sentiments) rather than just predicting surface-level emotion words. Use when the user wants to benchmark on EXPRESS, or asks about evaluating this task. Reports F1_V.
Evaluates an agent's ability to actively explore 3D environments to gather visual evidence and answer questions accurately, while measuring exploration efficiency and navigation performance. It specifically probes whether the agent's final answer is grounded in the actual visual observations collected during its exploration path, detecting hallucinations and ungrounded reasoning. Use when the user wants to benchmark on EXPRESS-Bench, or asks about evaluating this task. Reports C.
Evaluates the consistency of post-hoc explanation methods by quantifying how much feature attributions differ across algorithms for identical model predictions. It probes whether local explanations are reliable and whether practitioners have principled ways to resolve conflicts when different methods yield conflicting importance scores. Use when the user wants to benchmark on COMPAS, German Credit, News text dataset, PASCAL VOC 2012, or asks about evaluating this task. Reports L2 distance of ...
Evaluates the ability of an XAI-driven classification model to generate clinically meaningful segmentation masks without pixel-level annotations. It probes spatial coherence, boundary precision, and generalization across diverse medical imaging modalities (mammography, histopathology, endoscopy). Use when the user wants to benchmark on CBIS-DDSM, NuInsSeg, Kvasir-SEG, or asks about evaluating this task. Reports Dice.
This benchmark evaluates large language models on Chinese medical multiple-choice questions, specifically probing their ability to select correct answers and generate faithful, logically consistent free-text explanations. It measures both factual accuracy and the quality of interpretability in high-stakes healthcare domains. Use when the user wants to benchmark on ExplainCPE, or asks about evaluating this task. Reports Accuracy.
Evaluates AI agents' end-to-end capability to conduct real AI research experiments, including designing methodologies, implementing code, executing experiments, and drawing conclusions. Use when the user wants to benchmark on EXP-Bench, or asks about evaluating this task. Reports All·E✓.
Evaluates a Vision Transformer's ability to classify exoplanet transits by processing temporal light curve data transformed into image representations (Recurrence Plots and Gramian Angular Fields). It probes the model's capacity to capture long-range temporal dependencies and handle class imbalance in astronomical time-series data. Use when the user wants to benchmark on Kepler Light Curve Exoplanet Candidates, or asks about evaluating this task. Reports F1-score.
Evaluates the consistency of 1D radiative-convective atmospheric models in predicting exoplanet emission and transmission spectra, specifically probing how differences in opacity linelists, line shape treatments, and chemical equilibrium assumptions affect spectral predictions relative to JWST observational uncertainties. Use when the user wants to benchmark on Exoplanet Atmospheric Test Cases, or asks about evaluating this task. Reports spectral_resolution_and_jwst_error_bar_comparison.
This benchmark evaluates the ability of high-contrast imaging post-processing pipelines to accurately estimate the astrometric position of injected exoplanet signals in multispectral astronomical data. It probes how well algorithms handle varying signal-to-noise ratios, complex residual backgrounds (e.g., diffraction patterns, coronagraphic inner working angles), and different observing conditions. Use when the user wants to benchmark on Exoplanet Imaging Data Challenge Phase II, or asks abou...