Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,827
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,889–6,912 of 20,827 skills

Hw Nas Bench EvalA

Evaluates hardware-aware neural architecture search (HW-NAS) algorithms by measuring how effectively they discover network topologies that optimize the trade-off between classification accuracy and on-device inference latency for specific target hardware. Use when the user wants to benchmark on HW-NAS-Bench, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongo
0
3
Hvqr EvalA

This benchmark evaluates a model's ability to perform high-order, multistep visual question answering by integrating visual scene graphs with external commonsense knowledge. It explicitly probes the model's reasoning process by requiring it to predict intermediate knowledge triplets alongside the final answer, enforcing explainability and self-diagnosis capabilities. Use when the user wants to benchmark on HVQR, or asks about evaluating this task. Reports triplet recall.

researchpythongo
0
3
Huvad EvalA

Evaluates human-centric video anomaly detection models on continuously recorded real-world video streams. It measures how well models distinguish between normal and anomalous human activities across multiple camera views, particularly under continual learning conditions. Use when the user wants to benchmark on HuVAD, or asks about evaluating this task. Reports AUC-ROC.

researchpythongit
0
3
Hupd EvalA

Evaluates NLP models on patent-related tasks including binary classification of patent acceptance, multi-class subject area classification using IPC codes, and abstractive summarization of patent claims or descriptions into abstracts. Use when the user wants to benchmark on Harvard USPTO Patent Dataset (HUPD), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Humorrank EvalA

Evaluates LLM humor generation by conducting pairwise preference judgments on joke outputs across 300 diverse prompts. It measures how well models master comedic mechanisms rather than relying on model scale, using an LLM-as-judge framework to produce global rankings. Use when the user wants to benchmark on SemEval-2026 MWAHAHA, or asks about evaluating this task. Reports Bradley-Terry Maximum Likelihood Estimation.

researchpython
0
3
Hummusqa EvalA

Evaluates large audio-language models on music understanding by testing their ability to answer multiple-choice questions about audio excerpts. It probes perceptual grounding, structural/harmonic/cultural reasoning, and robustness against text-only shortcuts or answer-position bias. Use when the user wants to benchmark on HumMusQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Hummi Wholebody Manipulation EvalA

Evaluates a humanoid robot's capability to learn and execute whole-body manipulation skills from robot-free demonstrations, assessing manipulation precision, dynamic coordination, generalization to unseen environments/objects, and data-collection efficiency. Use when the user wants to benchmark on Squatting, Tossing, Bimanual Actions, Walking, Dynamic Motions, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Hume EvalA

Evaluates text embedding models against human baselines across 16 MTEB datasets, probing semantic similarity, classification, clustering, and reranking capabilities. It specifically measures cross-lingual performance and identifies task ambiguities where model scores may reflect label pattern reproduction rather than genuine understanding. Use when the user wants to benchmark on MTEB (16 datasets, 26 task-language pairs), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Humdial EvalA

This benchmark evaluates human-like spoken dialogue systems across two core capabilities: emotional intelligence (multi-turn emotion tracking, causal reasoning, and empathetic response generation) and full-duplex interaction (natural turn-taking, interruption handling, and noise rejection during concurrent listening and speaking). It uses authentic real-world conversations to measure long-term emotional consistency and cognitive synchronization. Use when the user wants to benchmark on HumDial...

researchpythongo
0
3
Humdial Eibench EvalA

Evaluates the emotional intelligence of audio language models across multi-turn dialogues. It probes four core capabilities: tracking emotional trajectories over time, reasoning about implicit emotional causes, generating empathetic responses, and resolving conflicts between acoustic and textual emotional signals. Use when the user wants to benchmark on HumDial-EIBench, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Humbi EvalA

Probes the generalizability, diversity, and reconstruction accuracy of 3D human body expression models (gaze, face, hand, body) across multiple datasets and viewpoints. It evaluates how well models trained on HUMBI generalize to unseen datasets and how accurately they reconstruct 3D geometry from monocular images. Use when the user wants to benchmark on HUMBI, or asks about evaluating this task. Reports IoU.

researchpythonexpress
0
3
Humanvbench EvalA

This benchmark evaluates the human-centric video understanding capabilities of multimodal large language models (MLLMs). It specifically probes inner emotion perception, outer behavioral manifestations, and cross-modal speech-visual alignment through 16 fine-grained multiple-choice tasks. Use when the user wants to benchmark on HumanVBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Humanscore EvalA

Evaluates the biomechanical plausibility and realism of human motion in AI-generated videos by measuring anatomical, kinematic, and kinetic correctness. It also assesses how well these automated metrics correlate with human preference judgments. Use when the user wants to benchmark on HumanScore Benchmark, or asks about evaluating this task. Reports Kinetic Correctness.

researchpython
0
3
Humanref EvalA

Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.

researchpythongo
0
3
Humanoid Pose Control EvalA

Evaluates a language-conditioned transformer model's ability to generate physically plausible and text-aligned 3D humanoid poses from text commands. It probes motion quality, diversity, and multimodal alignment on a retargeted human motion benchmark, as well as real-world deployment success rates. Use when the user wants to benchmark on HumanoidML3D, Humanoid-X, or asks about evaluating this task. Reports FID.

researchpython
0
3
Humanoid Policy EvalA

Evaluates a unified state-action policy's ability to perform dexterous manipulation tasks on humanoid robots. It specifically probes in-distribution (I.D.) task execution and out-of-distribution (O.O.D.) generalization across varying backgrounds, object placements, and cross-embodiment transfers. Use when the user wants to benchmark on Robot & Human Manipulation Demonstrations, or asks about evaluating this task. Reports Success rate.

researchpythongo
0
3
Humanoid Everyday EvalA

Evaluates imitation learning and vision-language-action policies on open-world humanoid manipulation. It probes robustness to high-dimensional action spaces, multimodal sensor fusion, and fine-grained visuospatial perception across locomotion, tool use, and precise manipulation tasks. Use when the user wants to benchmark on Humanoid Everyday, or asks about evaluating this task. Reports success rate.

researchpython
0
3
Humanevalfix EvalA

Evaluates a model's ability to debug and fix buggy code by generating corrected implementations that pass provided unit tests. It probes code repair capabilities across multiple programming languages. Use when the user wants to benchmark on HumanEvalFix, or asks about evaluating this task. Reports pass rate.

researchpythongo
0
3
Humaneval V EvalA

This benchmark evaluates large multimodal models' ability to perform high-level visual reasoning over complex diagrams in coding contexts. It specifically probes spatial transformations, topological relationships, and dynamic pattern understanding by requiring models to translate visual information into executable code. Use when the user wants to benchmark on HumanEval-V, or asks about evaluating this task. Reports pass@1.

researchpythongit
0
3
Humaneval EvalA

Evaluates a model's ability to generate correct, executable Python code from natural language function descriptions and signatures. It measures functional correctness by checking if generated code passes hidden unit tests. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports pass@1.

researchpythonperformance
0
3
Humanedit EvalA

Evaluates instruction-based image editing models on their ability to modify source images according to textual prompts, with and without provided segmentation masks. It measures pixel-level fidelity, image quality, and text-image alignment across diverse editing categories such as add, remove, replace, action, counting, and relation. Use when the user wants to benchmark on HumanEdit, or asks about evaluating this task. Reports CLIP-T.

researchpythongo
0
3
Humanbench EvalA

Evaluates the generalization and task-agnostic representation learning of human-centric vision models across six diverse downstream tasks. It probes how well a model trained on a large, multi-task human-centric corpus can adapt to in-distribution, out-of-distribution, and completely unseen human perception tasks. Use when the user wants to benchmark on HumanBench, or asks about evaluating this task. Reports mAP/mIoU/mA/Top1/pACC/MR/MSE/EPE.

researchpythongo
0
3
Human36m Pose Forecasting EvalA

Evaluates the ability of models to forecast future 3D human poses over a 1-second horizon given a short 40ms observation window. It probes long-term temporal prediction and personalization to individual-specific motion patterns. Use when the user wants to benchmark on Human3.6M, or asks about evaluating this task. Reports MPJE.

researchpython
0
3
Human Scene Vlm EvalA

Evaluates a vision-language model's ability to understand and generate detailed descriptions of human-centric scenes, answer open- and closed-set questions about them, recognize facial attributes, and ground textual references to human objects in images. Use when the user wants to benchmark on HumanCaptionHQ, HumanVQA, FaceC, CelebA, LFWA, RefCOCO, or asks about evaluating this task. Reports semantic similarity.

researchpythongo
0
3