Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,937–6,960 of 20,827 skills
Probes a model's ability to perform many-hop fact verification by retrieving supporting evidence from multiple Wikipedia articles and determining whether a given claim is supported or not supported. It specifically tests long-range dependency reasoning, coreference resolution, and the ability to avoid semantic matching shortcuts that degrade as hop count increases. Use when the user wants to benchmark on HOVER, or asks about evaluating this task. Reports claim verification accuracy.
Evaluates retrieval and reasoning over housing statutes, requiring models to connect queries to lexically distant legal texts and answer standardized Yes/No or categorical questions. Use when the user wants to benchmark on Housing Statute QA, or asks about evaluating this task. Reports Recall@10.
Evaluates an agent's ability to detect misplacements of objects in indoor scenes and plan optimal rearrangement placements for carryable objects based on scene context and affordances. It probes commonsense reasoning about object-receptacle relationships and ranking quality under varying contextual cues. Use when the user wants to benchmark on Tidybot benchmark, Context-oriented benchmark (HSSD 200), or asks about evaluating this task. Reports NDCG@8.
This benchmark evaluates a model's ability to perform multi-hop question answering by reasoning across multiple documents. It specifically probes explainability through supporting fact prediction and tests robustness against distractor paragraphs and large-scale retrieval contexts. Use when the user wants to benchmark on HotpotQA, or asks about evaluating this task. Reports F1.
Probes long-horizon personalization and belief-update capability. It tests whether models can track evolving user preferences across ~6 months of conversation history and correctly select responses aligned with updated preferences, rather than anchoring on outdated values. Use when the user wants to benchmark on HorizonBench, or asks about evaluating this task. Reports accuracy.
Evaluates user behavior modeling capabilities across temporal generalization, cross-domain prediction, and unseen-user scenarios. It probes how well recommendation models and LLMs can generalize to out-of-distribution users and future time periods using real-world interaction sequences. Use when the user wants to benchmark on Amazon Reviews (HORIZON Benchmark), or asks about evaluating this task. Reports NDCG@K.
Evaluates few-shot node classification on text-attributed graphs using self-supervised preference tuning. It probes the model's ability to leverage graph topology and anchor labels at inference without any supervised training, measuring both classification accuracy and inference efficiency. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy (%).
Evaluates object detection models on chest X-rays for identifying critical retained foreign objects (RFOs) like sponges and needles. It probes both classification accuracy and precise localization of rare medical anomalies under data-scarce conditions. Use when the user wants to benchmark on Hopkins RFOs Bench, or asks about evaluating this task. Reports ACC.
Evaluates the accuracy, latency, and energy efficiency of a homomorphic inference accelerator performing SVM classification on encrypted data under intermittent power constraints. Use when the user wants to benchmark on MNIST, Human Activity Recognition, ADULT, or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to reconstruct complex object fields from in-line digital holograms, specifically testing its capacity to suppress twin-image artifacts and maintain reconstruction fidelity under various noise conditions. Use when the user wants to benchmark on Synthetic Inline Holography Dataset, or asks about evaluating this task. Reports reconstruction fidelity.
Evaluates the security resilience and hardware overhead of Higher-Order Logic Locking (HOLL) against a counterexample-guided inductive synthesis (CEGIS) attack on combinational circuits. It measures how long an attacker takes to recover the secret key relation and the area penalty incurred by the locking mechanism. Use when the user wants to benchmark on ISCAS'85 and MCNC benchmarks, or asks about evaluating this task. Reports attack_time.
Evaluates demographic and intersectional biases in language models by measuring disparities in token likelihoods, generation styles, and offensiveness across a curated set of demographic descriptor terms embedded in sentence templates. Use when the user wants to benchmark on HOLISTICBIAS, or asks about evaluating this task. Reports Full Gen Bias.
Evaluates the quality, text-motion alignment, and diversity of generated 2D whole-body human motion sequences conditioned on text prompts. It probes the model's ability to capture fine-grained spatial-temporal dynamics and handle occlusions or noisy 2D pose data. Use when the user wants to benchmark on Holistic-Motion2D, or asks about evaluating this task. Reports FID.
Evaluates a video generation model's ability to synthesize high-fidelity videos conditioned on multimodal inputs (reference images, audio, pose, and text) while maintaining reference consistency, audio-visual synchronization, and temporal coherence. Use when the user wants to benchmark on HOIVG-Bench, EMTD, or asks about evaluating this task. Reports NexusScore.
Evaluates a model's ability to detect Human-Object Interactions (HOIs) by predicting triplets of person, verb, and object along with their bounding boxes. It specifically probes the model's robustness to object bias by measuring performance on rare versus frequent interactions under both standard and object-conditional evaluation protocols. Use when the user wants to benchmark on HICO-DET, HOI-COCO, or asks about evaluating this task. Reports mAP.
This evaluation probes a model's capability to detect network intrusions by classifying traffic flows as benign or malicious. It assesses the system's ability to learn complex temporal and multi-scale features from network flow data to distinguish between normal activities and various attack types. Use when the user wants to benchmark on Hogzilla Dataset, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates the quality and robustness of heterogeneous network embedding (HNE) algorithms across diverse real-world graphs. It probes how well learned representations preserve multi-type structural and attribute information, measured via downstream node classification and link prediction tasks. Use when the user wants to benchmark on DBLP, Yelp, Freebase, PubMed, or asks about evaluating this task. Reports macro-F1.
Evaluates the trade-off between predictive performance and group fairness when applying causal pre-processing to approximate an unbiased data distribution. It probes whether debiasing techniques can simultaneously satisfy multiple fairness constraints without degrading model accuracy. Use when the user wants to benchmark on HMDA (Wisconsin, 2022), or asks about evaluating this task. Reports AUC.
Evaluates the accuracy of optical flow estimation models on synthetic and real-world video sequences. It specifically probes the model's ability to capture fine object contours, handle small or fast-moving targets, and maintain robustness under downscaling and occlusion. Use when the user wants to benchmark on Sintel, KITTI-2015, or asks about evaluating this task. Reports EPE.
Evaluates an embodied AI agent's ability to navigate indoor 3D environments to find specific object categories using RGB-D observations. It measures both navigation quality (success and path efficiency) and computational efficiency (latency, memory, and skip ratio) on a large-scale dataset. Use when the user wants to benchmark on HabitatMatterport3D (HM3D), or asks about evaluating this task. Reports SPL.
Evaluates the quality of self-supervised visual representations learned from histopathology images by measuring downstream classification performance at patch, slide, and patient levels. It probes the model's ability to capture hierarchical pathological structures and align them with clinical diagnostic categories. Use when the user wants to benchmark on OpenSRH, TCGA, or asks about evaluating this task. Reports kNN classification accuracy (ACC).
Evaluates large language models' multilingual comprehension of Hong Kong-specific knowledge, Cantonese linguistic capabilities, and reasoning across STEM, social sciences, and humanities in both Traditional and Simplified Chinese. Use when the user wants to benchmark on HKMMLU, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of LLM-based sequential recommendation models to predict the next item in a user's interaction history. It specifically probes how well models capture temporal dynamics by incorporating irregular time intervals between interactions, and assesses performance under warm and cold-start conditions. Use when the user wants to benchmark on Amazon Reviews (Video Games, CDs and Vinyl, Books), or asks about evaluating this task. Reports Hit Ratio@1.
This benchmark evaluates graph-based generative models for their ability to produce biologically relevant, drug-like hit molecules rather than just chemically valid structures. It probes the models' capacity to satisfy strict medicinal chemistry constraints, maintain distributional similarity to known bioactive compounds, and achieve strong predicted binding affinity to specific protein targets. Use when the user wants to benchmark on REINVENT Dataset, Hit-like Dataset, Target-Specific Ligand...