Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,889–6,912 of 20,827 skills
Evaluates hardware-aware neural architecture search (HW-NAS) algorithms by measuring how effectively they discover network topologies that optimize the trade-off between classification accuracy and on-device inference latency for specific target hardware. Use when the user wants to benchmark on HW-NAS-Bench, or asks about evaluating this task. Reports top-1 accuracy.
This benchmark evaluates a model's ability to perform high-order, multistep visual question answering by integrating visual scene graphs with external commonsense knowledge. It explicitly probes the model's reasoning process by requiring it to predict intermediate knowledge triplets alongside the final answer, enforcing explainability and self-diagnosis capabilities. Use when the user wants to benchmark on HVQR, or asks about evaluating this task. Reports triplet recall.
Evaluates human-centric video anomaly detection models on continuously recorded real-world video streams. It measures how well models distinguish between normal and anomalous human activities across multiple camera views, particularly under continual learning conditions. Use when the user wants to benchmark on HuVAD, or asks about evaluating this task. Reports AUC-ROC.
Evaluates NLP models on patent-related tasks including binary classification of patent acceptance, multi-class subject area classification using IPC codes, and abstractive summarization of patent claims or descriptions into abstracts. Use when the user wants to benchmark on Harvard USPTO Patent Dataset (HUPD), or asks about evaluating this task. Reports accuracy.
Evaluates LLM humor generation by conducting pairwise preference judgments on joke outputs across 300 diverse prompts. It measures how well models master comedic mechanisms rather than relying on model scale, using an LLM-as-judge framework to produce global rankings. Use when the user wants to benchmark on SemEval-2026 MWAHAHA, or asks about evaluating this task. Reports Bradley-Terry Maximum Likelihood Estimation.
Evaluates large audio-language models on music understanding by testing their ability to answer multiple-choice questions about audio excerpts. It probes perceptual grounding, structural/harmonic/cultural reasoning, and robustness against text-only shortcuts or answer-position bias. Use when the user wants to benchmark on HumMusQA, or asks about evaluating this task. Reports accuracy.
Evaluates a humanoid robot's capability to learn and execute whole-body manipulation skills from robot-free demonstrations, assessing manipulation precision, dynamic coordination, generalization to unseen environments/objects, and data-collection efficiency. Use when the user wants to benchmark on Squatting, Tossing, Bimanual Actions, Walking, Dynamic Motions, or asks about evaluating this task. Reports success rate.
Evaluates text embedding models against human baselines across 16 MTEB datasets, probing semantic similarity, classification, clustering, and reranking capabilities. It specifically measures cross-lingual performance and identifies task ambiguities where model scores may reflect label pattern reproduction rather than genuine understanding. Use when the user wants to benchmark on MTEB (16 datasets, 26 task-language pairs), or asks about evaluating this task. Reports accuracy.
This benchmark evaluates human-like spoken dialogue systems across two core capabilities: emotional intelligence (multi-turn emotion tracking, causal reasoning, and empathetic response generation) and full-duplex interaction (natural turn-taking, interruption handling, and noise rejection during concurrent listening and speaking). It uses authentic real-world conversations to measure long-term emotional consistency and cognitive synchronization. Use when the user wants to benchmark on HumDial...
Evaluates the emotional intelligence of audio language models across multi-turn dialogues. It probes four core capabilities: tracking emotional trajectories over time, reasoning about implicit emotional causes, generating empathetic responses, and resolving conflicts between acoustic and textual emotional signals. Use when the user wants to benchmark on HumDial-EIBench, or asks about evaluating this task. Reports Accuracy (%).
Probes the generalizability, diversity, and reconstruction accuracy of 3D human body expression models (gaze, face, hand, body) across multiple datasets and viewpoints. It evaluates how well models trained on HUMBI generalize to unseen datasets and how accurately they reconstruct 3D geometry from monocular images. Use when the user wants to benchmark on HUMBI, or asks about evaluating this task. Reports IoU.
This benchmark evaluates the human-centric video understanding capabilities of multimodal large language models (MLLMs). It specifically probes inner emotion perception, outer behavioral manifestations, and cross-modal speech-visual alignment through 16 fine-grained multiple-choice tasks. Use when the user wants to benchmark on HumanVBench, or asks about evaluating this task. Reports accuracy.
Evaluates the biomechanical plausibility and realism of human motion in AI-generated videos by measuring anatomical, kinematic, and kinetic correctness. It also assesses how well these automated metrics correlate with human preference judgments. Use when the user wants to benchmark on HumanScore Benchmark, or asks about evaluating this task. Reports Kinetic Correctness.
Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.
Evaluates a language-conditioned transformer model's ability to generate physically plausible and text-aligned 3D humanoid poses from text commands. It probes motion quality, diversity, and multimodal alignment on a retargeted human motion benchmark, as well as real-world deployment success rates. Use when the user wants to benchmark on HumanoidML3D, Humanoid-X, or asks about evaluating this task. Reports FID.
Evaluates a unified state-action policy's ability to perform dexterous manipulation tasks on humanoid robots. It specifically probes in-distribution (I.D.) task execution and out-of-distribution (O.O.D.) generalization across varying backgrounds, object placements, and cross-embodiment transfers. Use when the user wants to benchmark on Robot & Human Manipulation Demonstrations, or asks about evaluating this task. Reports Success rate.
Evaluates imitation learning and vision-language-action policies on open-world humanoid manipulation. It probes robustness to high-dimensional action spaces, multimodal sensor fusion, and fine-grained visuospatial perception across locomotion, tool use, and precise manipulation tasks. Use when the user wants to benchmark on Humanoid Everyday, or asks about evaluating this task. Reports success rate.
Evaluates a model's ability to debug and fix buggy code by generating corrected implementations that pass provided unit tests. It probes code repair capabilities across multiple programming languages. Use when the user wants to benchmark on HumanEvalFix, or asks about evaluating this task. Reports pass rate.
This benchmark evaluates large multimodal models' ability to perform high-level visual reasoning over complex diagrams in coding contexts. It specifically probes spatial transformations, topological relationships, and dynamic pattern understanding by requiring models to translate visual information into executable code. Use when the user wants to benchmark on HumanEval-V, or asks about evaluating this task. Reports pass@1.
Evaluates a model's ability to generate correct, executable Python code from natural language function descriptions and signatures. It measures functional correctness by checking if generated code passes hidden unit tests. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports pass@1.
Evaluates instruction-based image editing models on their ability to modify source images according to textual prompts, with and without provided segmentation masks. It measures pixel-level fidelity, image quality, and text-image alignment across diverse editing categories such as add, remove, replace, action, counting, and relation. Use when the user wants to benchmark on HumanEdit, or asks about evaluating this task. Reports CLIP-T.
Evaluates the generalization and task-agnostic representation learning of human-centric vision models across six diverse downstream tasks. It probes how well a model trained on a large, multi-task human-centric corpus can adapt to in-distribution, out-of-distribution, and completely unseen human perception tasks. Use when the user wants to benchmark on HumanBench, or asks about evaluating this task. Reports mAP/mIoU/mA/Top1/pACC/MR/MSE/EPE.
Evaluates the ability of models to forecast future 3D human poses over a 1-second horizon given a short 40ms observation window. It probes long-term temporal prediction and personalization to individual-specific motion patterns. Use when the user wants to benchmark on Human3.6M, or asks about evaluating this task. Reports MPJE.
Evaluates a vision-language model's ability to understand and generate detailed descriptions of human-centric scenes, answer open- and closed-set questions about them, recognize facial attributes, and ground textual references to human objects in images. Use when the user wants to benchmark on HumanCaptionHQ, HumanVQA, FaceC, CelebA, LFWA, RefCOCO, or asks about evaluating this task. Reports semantic similarity.