Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,721–6,744 of 20,827 skills
Evaluates the correlation and predictability of standard information retrieval metrics across multiple TREC test collections. It probes how well low-cost metrics can predict high-cost ones and how metric values vary between topic-wise and system-wise aggregations. Use when the user wants to benchmark on TREC Web & Robust Tracks (2000-2014), or asks about evaluating this task. Reports MAP.
Evaluates an agent's ability to navigate, perceive, and interact with dynamic 3D environments to answer visual questions. It probes spatial reasoning, long-term memory of object locations, and the ability to plan exploration based on question semantics. Use when the user wants to benchmark on iquad v1, or asks about evaluating this task. Reports Top-1 question answering accuracy.
This benchmark evaluates a model's ability to identify core user intents in personalized question answering. It probes whether systems can infer prioritized motivations from a user's historical Q&A interactions and a target question's narrative, rather than relying on explicit user statements. Use when the user wants to benchmark on IPQA, or asks about evaluating this task. Reports IPQA-Eval F1.
Evaluates a visualization agent's ability to reconstruct interactive charts from static images and answer binary questions about them. It probes the agent's capacity for spec-grounded introspection and view-grounded interaction to resolve visual ambiguities like overlapping geometries. Use when the user wants to benchmark on iPlotBench, or asks about evaluating this task. Reports Question-level accuracy.
Evaluates an AI agent's ability to solve complex, multi-part physics theory problems that require integrating visual data extraction, causal reasoning, and self-correction via tool use. The benchmark probes whether agentic architectures with domain-specific tools can match elite human performance on standardized physics competitions. Use when the user wants to benchmark on IPhO 2025 Theory Problems, or asks about evaluating this task. Reports score.
Evaluates the ability of a graph neural network to detect network intrusions in IoT environments by classifying traffic flows as benign or malicious, and identifying specific attack types. It probes the model's capacity to leverage topological graph structures and edge features for robust intrusion detection across imbalanced, real-world network traffic datasets. Use when the user wants to benchmark on BoT-IoT, NF-BoT-IoT, ToN-IoT, NF-ToN-IoT, or asks about evaluating this task. Reports F1-Sc...
Evaluates the ability of a lightweight CNN to classify IoT binary files as benign or belonging to specific DDoS malware families (Mirai, Linux.Gafgyt) by converting raw binaries into 64x64 grayscale images. Use when the user wants to benchmark on IoTPOT IoT DDoS Malware Dataset, or asks about evaluating this task. Reports accuracy.
This evaluation probes an intrusion detection model's ability to classify network traffic flows as benign or malicious across highly imbalanced IoT datasets. It specifically tests the model's robustness to extreme class imbalance and its capacity to leverage graph-structured representations of network flows for anomaly detection. Use when the user wants to benchmark on BoT-IoT, ToN-IoT, or asks about evaluating this task. Reports macro-F1.
Evaluates LLMs' ability to solve advanced astronomy and astrophysics problems, focusing on geometric/spatial reasoning, physical calculations, and multimodal data analysis. It benchmarks performance against human Olympiad participants using official scoring rubrics. Use when the user wants to benchmark on IOAA (International Olympiad on Astronomy and Astrophysics), or asks about evaluating this task. Reports score.
Evaluates the sequential financial decision-making capabilities of LLM-based agents across stock, cryptocurrency, and ETF trading environments. It probes the model's ability to process multi-modal market data, manage portfolio risk, and adapt to volatile market conditions over time. Use when the user wants to benchmark on INVESTORBENCH, or asks about evaluating this task. Reports SR (Sharpe Ratio).
Evaluates a model's capability to perform novel view synthesis, decompose scene properties (albedo, normals, roughness), and relight scenes under new lighting conditions using Gaussian surfels. It specifically probes the model's ability to model indirect illumination and inter-reflections without relying on pre-trained novel view synthesis data. Use when the user wants to benchmark on TensoIR*, Synthetic4Relight*, or asks about evaluating this task. Reports PSNR.
Tests a framework's ability to compress pairwise preference data into interpretable natural language principles (constitutions) and use them to reconstruct original annotations. It probes the model's adaptability to aligned, unaligned, individual, and demographic group preferences, as well as its capacity for bias detection. Use when the user wants to benchmark on Synthetic data, AlpacaEval, Chatbot Arena Conversations, PRISM, or asks about evaluating this task. Reports agreement.
Evaluates the fidelity of natural language explanations for sparse autoencoder (SAE) features by measuring how well an explanation predicts the downstream effects of directly intervening on the feature's activation, rather than just correlating with input contexts. Use when the user has predictions and gold and needs to compute intervention_scoring.
Evaluates LLM fairness and consistency across intersectional identity attributes (race, gender, socio-economic status) in both ambiguous and disambiguated contexts. It measures accuracy, stereotype alignment, subgroup disparity, and response stability across repeated runs. Use when the user wants to benchmark on Race_SES, Race_Gender, or asks about evaluating this task. Reports Accuracy.
Evaluates reinforcement learning agents' ability to navigate complex, un-signalized urban intersections under varying traffic conditions. It probes decision-making, collision avoidance, and route completion in dynamic environments with interacting social vehicles. Use when the user wants to benchmark on Intersection Scenarios (RL-CIS), or asks about evaluating this task. Reports Success rate(%).
Evaluates the runtime performance and memory efficiency of a speculatively staged Python interpreter against standard baselines like CPython and PyPy. It probes the interpreter's ability to eliminate dynamic type-checking overhead and optimize instruction dispatch through compile-time specialization. Use when the user wants to benchmark on Computer Language Benchmarks Game, or asks about evaluating this task. Reports speedup.
Evaluates multimodal large language models across general understanding, complex reasoning, mathematics, OCR, document comprehension, and agentic/GUI interaction tasks. Use when the user wants to benchmark on MMMU, MathVista, MMStar, MMVet, or asks about evaluating this task. Reports accuracy.
Evaluates the zero-shot sim-to-real transfer and generalization of a Vision-Language-Action (VLA) policy on diverse real-world and simulated manipulation tasks. It probes fundamental pick-and-place, articulated object manipulation, human-robot interaction, and long-horizon task composition capabilities. Use when the user wants to benchmark on InternData-A1 Real-World & Sim-to-Real Benchmarks, or asks about evaluating this task. Reports average success rate.
Evaluates a model's ability to edit two-person 3D motions according to text instructions, balancing semantic modification (instruction adherence) with content preservation (source fidelity) and motion realism. Use when the user wants to benchmark on InterEdit3D, or asks about evaluating this task. Reports Recall@1.
Evaluates the effectiveness of interactive document retrieval using user-identified Wikipedia concepts for query expansion and re-ranking. It also tests methods for selecting the most relevant Wikipedia concepts from a large pool based on semantic relevance and document ranking signals. Use when the user wants to benchmark on TREC Filtering-02, HARD-03, HARD-05, or asks about evaluating this task. Reports MAP, P@10.
Evaluates Large Audio Models (LAMs) on real-world, task-oriented voice assistant interactions by capturing user preferences through open-ended pairwise comparisons. It measures how well models align with actual user needs and preferences in an interactive setting, rather than relying on static reference-based benchmarks. Use when the user wants to benchmark on TalkArena Interactive User Preferences, or asks about evaluating this task. Reports Bradley-Terry model score.
Evaluates the ability of generative models to synthesize physically plausible, contact-consistent 3D human-object interaction sequences conditioned on text, actions, or object shapes. It probes motion realism, contact accuracy, and alignment between linguistic/action prompts and generated kinematics. Use when the user wants to benchmark on InterAct, or asks about evaluating this task. Reports FID.
Evaluates inter-rater variability among pathologists annotating histopathology images and measures how annotator conformity (agreement with an anchor) impacts downstream deep learning cell detection performance. Use when the user wants to benchmark on Histopathology Cell Annotation Dataset, or asks about evaluating this task. Reports mF1-score.
This benchmark evaluates the ability of pretrained sentence encoders and classifiers to correctly identify user intents from conversational utterances. It specifically probes few-shot generalization by testing models on severely limited training data (10 or 30 examples per intent) while maintaining a standard full test set. Use when the user wants to benchmark on BANKING77, CLINC150, HWU64, or asks about evaluating this task. Reports accuracy.