Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,057–7,080 of 20,839 skills
Evaluates multi-modal human activity recognition capabilities by classifying 300 action categories from synchronized RGB frames and asynchronous event streams under challenging real-world conditions such as low light, occlusion, and dynamic backgrounds. Use when the user wants to benchmark on HARDVS 2.0, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of lightweight convolutional neural networks to accurately classify human activities from wearable sensor time-series data. It probes the trade-off between model compression (parameter count and FLOPs) and classification performance on highly imbalanced, multi-class activity recognition tasks. Use when the user wants to benchmark on UCI-HAR, OPPORTUNITY, PAMAP2, UNIMIB-SHAR, WISDM, or asks about evaluating this task. Reports weighted F1 score.
Evaluates deep learning architectures (CNNs, LSTMs, DNNs) for frame-by-frame human activity recognition using wearable sensor time-series data. It probes the models' capacity to capture temporal dependencies and generalize across diverse domains (kitchen gestures, lifestyle/exercise, and medical gait analysis) while handling severe class imbalance. Use when the user wants to benchmark on Opportunity, PAMAP2, Daphnet Gait, or asks about evaluating this task. Reports mean f1-score.
Probes the ability of multimodal and unimodal models to recognize human activities from wearable sensor, pose, and video data. It evaluates classification performance, data efficiency across low-data regimes (1-100% training fractions), and zero-shot transfer capability to unseen real-world datasets. Use when the user wants to benchmark on MM-Fit, MHEALTH, MyoGym, MotionSense, or asks about evaluating this task. Reports Macro F1-Score.
This benchmark evaluates continual learning algorithms on sensor-based human activity recognition (HAR) datasets. It measures how well models balance plasticity (learning new activities) and stability (retaining old activities) while incrementally processing tasks, specifically probing robustness to class imbalance, sensor noise, and cross-user data leakage. Use when the user wants to benchmark on House A (HA), CASAS (WS, Milan, Twor, Aruba), PAMAP2, DSADS, HAPT, or asks about evaluating this...
Evaluates the ability of classical, deep learning, and generative models to accurately classify human activities from sensor data. Probes temporal pattern recognition, sensor fusion handling, and generalization across varying data complexities and sensor modalities. Use when the user wants to benchmark on UCI-HAR, Opportunity, PAMAP2, WISDM, Berkeley MHAD, or asks about evaluating this task. Reports Accuracy.
Evaluates Chinese entity linking models on few-shot and zero-shot scenarios, specifically probing their ability to link mentions to tail and emerging Wikidata entities without relying on head entity popularity or dataset-specific fine-tuning. Use when the user wants to benchmark on Hansel, TAC-KBP2015, or asks about evaluating this task. Reports R@1.
Probes whether neural NLI models rely on superficial syntactic heuristics (e.g., lexical overlap, subsequence matching) rather than genuine logical reasoning by presenting structurally similar counterexamples where heuristics lead to incorrect predictions. Use when the user wants to benchmark on HANS, or asks about evaluating this task. Reports accuracy.
Evaluates video foundation models' ability to understand fine-grained spatiotemporal dynamics in hand-object interactions. It probes spatial reasoning, motion tracking, and part-level geometric grounding through multiple-choice questions and video object segmentation tasks. Use when the user wants to benchmark on HanDyVQA, or asks about evaluating this task. Reports top-1 accuracy.
Evaluates a robot's ability to understand ambiguous human instructions and infer human subgoals in physically and socially complex household environments. It probes pragmatic reasoning, goal recognition, and collaborative task completion under partial observability. Use when the user wants to benchmark on HandMeThat, or asks about evaluating this task. Reports success rate.
Evaluates sequential dexterous manipulation by requiring a robot to first grasp a target object and then perform a specific downstream task (e.g., pushing, pressing, twisting, pulling, or picking a second object) while maintaining the grasp. It probes the policy's ability to allocate finger resources and maintain stable contacts to satisfy competing subtask constraints. Use when the user wants to benchmark on HANDFUL-Bench, or asks about evaluating this task. Reports terminal success rate ($p...
Evaluates the accuracy of 3D hand pose estimation methods by measuring joint localization error on isolated frame pairs. It probes how well different directional distance metrics handle orientation information and varying temporal offsets between frames. Use when the user wants to benchmark on Synthetic dataset, Realistic dataset, or asks about evaluating this task. Reports average joint error.
Evaluates the ability of a model to personalize a 3D hand avatar from a single RGB image and render it under novel poses and lighting conditions. It probes physically-based rendering accuracy, material/albedo recovery, and relighting generalization. Use when the user wants to benchmark on InterHand2.6M, HARP relit, or asks about evaluating this task. Reports PSNR.
Evaluates language models' ability to capture authorial style and stylometric patterns in scholarly text, independent of topical content. It probes whether models can distinguish documents by the same author across different topics and languages using triplet classification and document retrieval tasks. Use when the user wants to benchmark on HALvest-Contrastive, or asks about evaluating this task. Reports Accuracy.
Evaluates an agent's ability to learn humanlike abstractions and affordances for rapid problem solving in a structured visual game. It probes three levels of generalization: perceptual recognition, conceptual abstraction of semantics, and algorithmic strategy formation under limited training exposure. Use when the user wants to benchmark on HALMA, or asks about evaluating this task. Reports goal_reaching ($ ho_g$).
Evaluates LVLMs' ability to resist prompt-induced hallucinations by disentangling perception failures from instruction-induced presuppositions. It probes whether models rely on visual evidence or textual priors when answering questions that imply the presence of non-existent objects. Use when the user wants to benchmark on HalluScope, or asks about evaluating this task. Reports AdP.
Evaluates whether reinforcement finetuned language models appropriately refuse to answer unanswerable or ambiguous questions, and measures their accuracy on standard solvable math benchmarks to ensure performance is not degraded by the refusal training. Use when the user wants to benchmark on UWMP, SelfAware, Synthetic Unanswerable Math (SUM), GSM8K, Minerva, MATH-500, OlympiadBench, AMC23, or asks about evaluating this task. Reports refusal_rate.
Evaluates the ability of Large Vision-Language Models to generate factually aligned outputs by measuring object hallucination rates in captions and yes/no answers, as well as logical reasoning and attribute consistency across diverse visual prompts. Use when the user wants to benchmark on POPE, CHAIR, MMHal-Bench, or asks about evaluating this task. Reports POPE Average Accuracy.
Probes a model's ability to avoid generating factually incorrect statements about visual content. It measures alignment between model outputs and ground-truth visual facts using binary detection and scoring metrics. Use when the user wants to benchmark on POPE, AMBER-d, HallusionBench, or asks about evaluating this task. Reports Accuracy (Acc).
Evaluates the ability of various detection methods to identify hallucinations in financial question-answering systems augmented with knowledge graphs. It probes robustness to noisy or contradictory KG triplets by comparing performance with and without structured evidence. Use when the user wants to benchmark on HalluBench, or asks about evaluating this task. Reports F1.
This benchmark probes the hallucination detection capabilities of Large Audio-Language Models (LALMs) across speech, environmental sound, and music domains. It systematically induces hallucinations using adversarial prompts and mixed-audio inputs to evaluate response correctness, affirmative bias, and refusal behavior beyond standard accuracy. Use when the user wants to benchmark on HalluAudio, or asks about evaluating this task. Reports Accuracy.
Evaluates zero-shot text-to-speech synthesis capability for generating minute-long audio from text and a short reference prompt. It measures linguistic accuracy, speaker similarity, audio quality, and temporal dynamics against ground truth speech. Use when the user wants to benchmark on MinutesSpeech, LibriSpeech, or asks about evaluating this task. Reports WER.
Evaluates Large Vision-Language Models (LVLMs) for hallucinations by measuring their ability to generate faithful image descriptions (generative evaluation) and detect hallucinations in provided captions (discriminative evaluation). It specifically probes fine-grained hallucination categories: object, relation, attribute, and event hallucinations, while also analyzing the impact of output length and Chain-of-Thought prompting. Use when the user wants to benchmark on COCO 2014, or asks about e...
Evaluates language models' proficiency in Korean cultural knowledge and context. It probes capabilities across vocabulary (loan words, standard nomenclature, rare words), history, general knowledge, and reading comprehension, specifically highlighting the limitations of English-trained or non-Korean-tailored models. Use when the user wants to benchmark on HAE-RAE Bench, or asks about evaluating this task. Reports accuracy.