Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,441–4,464 of 23,503 skills
Evaluates stereotypical bias in pretrained language models across gender, profession, race, and religion using context-based association tests. It measures both language modeling capability and the model's tendency to prefer stereotypical over anti-stereotypical associations in natural language contexts. Use when the user wants to benchmark on StereoSet, or asks about evaluating this task. Reports icat.
Evaluates 6D object pose estimation methods under stereo vision conditions, specifically probing robustness to occlusion and scale ambiguity by leveraging dense 2D-3D correspondences and stereo disparity. Use when the user wants to benchmark on Stereo PBR YCB-V DS, or asks about evaluating this task. Reports ADD0.1.
Evaluates the ability of a stereo video deblurring algorithm to recover sharp frames from motion-blurred inputs, specifically testing its robustness to spatially-variant blur caused by independent 3D object motion and non-planar surfaces. Use when the user wants to benchmark on Custom synthetic raytraced & real stereo captures, or asks about evaluating this task. Reports PSNR.
Evaluates a vision-language model's ability to perceive, locate, and interact with graphical user interfaces across desktop and mobile environments. It also measures general multimodal reasoning and OCR capabilities to ensure the model retains broad foundational skills after GUI-specific training. Use when the user wants to benchmark on ScreenSpot-Pro, ScreenSpot-v2, OSWorld-G, MMBench-GUI-L2, VisualWebBench, OSWorld-Verified, AndroidWorld, AndroidDaily, or asks about evaluating this task. Re...
Evaluates audio language models on speech understanding, reasoning, and real-time interactive dialogue capabilities using raw acoustic signals rather than textual transcriptions. It measures both comprehension accuracy across multiple audio benchmarks and real-time generation fluency. Use when the user wants to benchmark on Big Bench Audio, Spoken MQA, MMSU, MMAU, Wild Speech, or asks about evaluating this task. Reports Average Score (%).
Probes a model's capability to accurately edit or synthesize audio with specific emotional tones, speaking styles, and paralinguistic elements (e.g., laughter, sighs) using zero-shot voice cloning or iterative refinement. It measures how well the model preserves linguistic content while transferring or inserting target audio attributes. Use when the user wants to benchmark on Step-Audio-Edit-Test, or asks about evaluating this task. Reports accuracy.
Evaluates sequential recommendation models that incorporate social influence and temporal dynamics. It probes the ability to predict the next item a user will interact with based on their historical behavior sequence, social connections, and event timestamps. Use when the user wants to benchmark on Delicious, Yelp, Ciao, or asks about evaluating this task. Reports Recall@10.
Evaluates academic search and recommendation systems in live production environments using A/B testing and user interaction logs, bridging the gap between offline test collections and real-world performance. Use when the user wants to benchmark on LIVIVO, GESIS Search, or asks about evaluating this task. Reports click-paths.
This framework evaluates the effectiveness of representation steering methods in modifying specific safety behaviors (harmfulness, hallucination, bias) while measuring cross-perspective entanglement. It probes whether steering interventions achieve their target behavioral changes without causing unintended degradation in other safety or reasoning capabilities. Use when the user wants to benchmark on SteeringSafety Benchmark (17 datasets), or asks about evaluating this task. Reports effectiven...
Evaluates the safety and helpfulness alignment of multimodal large language models under single-turn versus multi-turn interactive settings, specifically probing the static-to-dynamic generalization gap and the evolution of safety failure rates across conversation turns. Use when the user wants to benchmark on Steer-Bench, or asks about evaluating this task. Reports pass_rate.
This evaluation probes a decision-maker's ability to stealthily sample a subset of a dataset to artificially satisfy fairness metrics (Demographic Parity) while remaining statistically indistinguishable from the original data distribution. It measures how well biased sampling algorithms can evade detection by ideal auditors using distributional tests like Kolmogorov-Smirnov or Wasserstein distance. Use when the user wants to benchmark on Synthetic Loan Check, COMPAS, Adult, or asks about eval...
Evaluates a deep learning model's ability to directly estimate earthquake magnitude (both local ML and duration Md) from raw, unprocessed single-station seismograms without normalization or instrument response correction. It probes the model's robustness to site effects, regional calibration differences, and varying signal-to-noise ratios. Use when the user wants to benchmark on STEAD, or asks about evaluating this task. Reports mean error.
This benchmark evaluates whether deep learning models can accurately predict the epicentral distance of an earthquake from single-station ground motion waveforms. It specifically probes whether models learn intrinsic seismic features or merely exploit highly correlated auxiliary signals like P/S wave arrival times. Use when the user wants to benchmark on Stanford Earthquake Dataset (STEAD), or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Evaluates whether LLMs can track dynamic states over sequential update instructions. It probes the model's ability to maintain and update internal representations of an environment's state across multiple steps, testing sequential reasoning and input-window memory limits. Use when the user wants to benchmark on State-Tracking-Tasks (LinearWorld, HandSwap, Lights), or asks about evaluating this task. Reports accuracy.
Evaluates multimodal agents' ability to perceive current GUI states from screenshots, interpret natural language toggle instructions, and execute precise click actions. It specifically probes state-aware reasoning by measuring accuracy on both positive and negative toggle instructions, as well as grounding precision and false positive/negative rates. Use when the user wants to benchmark on state control benchmark, dynamic evaluation benchmark, or asks about evaluating this task. Reports O-AMR.
Evaluates a model's ability to retrieve relevant statistical data tables from a large corpus based on conversational dialogue history, and its ability to generate appropriate agent responses. It probes intent understanding, table-level grounding, and robustness to temporal distribution shifts. Use when the user wants to benchmark on StatCan Dialogue Dataset, or asks about evaluating this task. Reports recall@10.
Evaluates whether a single Vision-Language-Action model can generalize across diverse robotic manipulation benchmarks without task-specific fine-tuning. It probes the model's cross-embodiment generalization and robustness to varying action spaces and task distributions. Use when the user wants to benchmark on LIBERO, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, RoboChallenge, or asks about evaluating this task. Reports success_rate.
Evaluates retrieval models on their ability to find relevant entities in semi-structured knowledge bases using complex queries that combine textual descriptions and relational constraints. It probes joint reasoning over mixed textual-relational semantics and user-intent modeling across product, academic, and medical domains. Use when the user wants to benchmark on STaRK, or asks about evaluating this task. Reports Hit@k.
Evaluates vision-language models' ability to parse free-form workflow sketch images and generate structured JSON workflow definitions. It probes structural fidelity, trigger and component recognition, and hierarchical consistency in diagram-to-code translation. Use when the user wants to benchmark on StarFlow Dataset, or asks about evaluating this task. Reports FlowSim.
Evaluates code generation, completion, and bug-fixing capabilities across multiple programming languages and libraries. It probes a model's ability to write correct functions from prompts, translate code across languages, and fix existing buggy code using standard and enhanced benchmarks. Use when the user wants to benchmark on HumanEval, MBPP, EvalPlus, MultiPL-E, DS-1000, HumanEvalFix, or asks about evaluating this task. Reports pass@1.
Evaluates an LLM's ability to generate correct SQL queries from natural language questions over complex, multi-table database schemas. It probes the model's reasoning capabilities and schema generalization by requiring step-by-step rationales and testing on unseen databases. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports execution accuracy (EX).
Evaluates a model's ability to perform spatio-temporal reasoning and answer questions about dynamic scenes using compressed textual scene graph sequences. It probes the model's capacity to track object interactions, understand event ordering, and generalize to unseen temporal compositions without relying on raw visual inputs. Use when the user wants to benchmark on STAR, AGQA, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of models to recover 3D geometry and surface material from images, and to synthesize novel views or relight objects in unseen real-world environments. It probes inverse rendering capabilities under natural, uncontrolled lighting conditions where ground-truth material is unavailable. Use when the user wants to benchmark on Stanford-ORB, or asks about evaluating this task. Reports Bidirectional Chamfer Distance.
Evaluates 2D-to-3D semantic transfer pipelines for indoor scene understanding, focusing on per-point labeling accuracy for structural and furniture classes, and detection sensitivity for novel safety-critical objects in public safety contexts. Use when the user wants to benchmark on Stanford 2D-3D-S*, or asks about evaluating this task. Reports per-point accuracy.