Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,233–5,256 of 23,574 skills
This evaluation protocol assesses the cross-domain generalization, safety, and instruction-following capabilities of models after reasoning-focused supervised fine-tuning (SFT). It measures in-domain math performance, out-of-domain reasoning in coding and science, general instruction following, and resistance to harmful queries. Use when the user wants to benchmark on MATH500, AIME24, LiveCodeBench v2, GPQA-Diamond, MMLU-Pro, IFEval, AlpacaEval 2.0, HaluEval, TruthfulQA, HEx-PHI, or asks abou...
Evaluates language model reasoning and general knowledge across multiple standard benchmarks. It measures the impact of data mixture optimization on model performance using accuracy scores on commonsense, scientific, and factual QA tasks. Use when the user wants to benchmark on PIQA, ARC_C, ARC_E, HellaSwag, WinoGrande, SIQA, MMLU, or asks about evaluating this task. Reports test accuracy.
Evaluates the ability of 3D reconstruction models to synthesize novel views of smoke-degraded scenes with high photometric fidelity and structural preservation. It probes view-dependent medium modeling and multi-view consistency under severe scattering conditions. Use when the user wants to benchmark on RealX3D (NTIRE 2026 Track 2 Smoke Subset), or asks about evaluating this task. Reports PSNR.
Evaluates the scalability and performance of a simulation-checking algorithm for timed automata under fairness assumptions. It measures how efficiently the algorithm verifies liveness properties and handles state-space explosion across parameterized real-time system benchmarks. Use when the user wants to benchmark on Fischer's timed mutual exclusion algorithm, CSMA/CD, Timed consumer/producer, Network of TAs, or asks about evaluating this task. Reports CPU time.
Evaluates end-to-end simultaneous speech-to-speech translation quality and latency on long-form, multi-domain continuous speech. Probes the model's ability to maintain semantic accuracy, speaker voice characteristics, and low delay in real-time multilingual dialogue. Use when the user wants to benchmark on RealSI, or asks about evaluating this task. Reports VIP.
Evaluates the robustness of Out-of-Distribution (OOD) detection models under realistic distribution shifts caused by semantic-preserving transformations. It measures how well detectors distinguish between true out-of-distribution samples and inlier samples that have undergone common corruptions or augmentations, revealing performance gaps that standard benchmarks miss. Use when the user wants to benchmark on CIFAR-10-R, CIFAR-100-R, ImageNet-30-R, or asks about evaluating this task. Reports A...
Evaluates scientific chart question answering capabilities, specifically testing a model's ability to extract, reason over, and answer questions about real-world scientific charts. It probes first-order logic reasoning and neuro-symbolic capabilities by requiring formal verification of logical inferences from complex visual data. Use when the user wants to benchmark on RealCQA, or asks about evaluating this task. Reports accuracy.
Assesses multi-step agent capabilities including tool use, GUI grounding, compositional generalization, and long-horizon planning across real-world desktop and web applications. It evaluates whether agents can execute complex, cross-application workflows and self-evaluate their trajectories. Use when the user wants to benchmark on Real-World Cross-Application Benchmark Suite, or asks about evaluating this task. Reports Success.
Evaluates real-time video game control policies across programmatic and real-game environments, measuring task completion, combat effectiveness, human-like behavior, and instruction-following capability. Use when the user wants to benchmark on Hovercraft, Simple-FPS, Real Games (DOOM, Quake, Roblox), or asks about evaluating this task. Reports Hovercraft Loop Time.
Evaluates neural combinatorial optimization models on real-world vehicle routing problems, measuring their ability to generate high-quality routes under asymmetric travel constraints and generalizing to out-of-distribution city maps and location distributions. Use when the user wants to benchmark on Real-World Routing (RRNCO), or asks about evaluating this task. Reports Gap %.
Evaluates offline reinforcement learning and imitation learning algorithms on real-world dexterous manipulation tasks. It probes the ability to learn precise in-hand orientation and stable grasping from pre-collected robot data without online interaction, and measures transfer performance to physical hardware. Use when the user wants to benchmark on Real Robot Challenge 2022 TriFinger Datasets, or asks about evaluating this task. Reports overall score.
Evaluates industrial anomaly detection models under standard unsupervised and fully unsupervised (noisy training) settings. It probes image-level, pixel-level, and multi-view sample-level defect detection capabilities. Use when the user wants to benchmark on Real-IAD, or asks about evaluating this task. Reports AUROC.
Evaluates whether 3D-LLMs genuinely comprehend 3D spatial relationships rather than relying on linguistic shortcuts or text-only priors. It filters out 3D-independent questions and measures consistency across viewpoint rotations to penalize superficial pattern matching. Use when the user wants to benchmark on Real-3DQA, or asks about evaluating this task. Reports Viewpoint Rotation Score (VRS).
Evaluates a cross-domain representation learning framework for protein-molecule interactions by measuring prediction accuracy on regression and classification tasks across diverse biochemical benchmarks. Use when the user wants to benchmark on FreeSolv, CEP, BetaLactamase, Stability, BindingDB, PPIAffinity, BBBP, GO-CC, DrugBank, HumanPPI, YeastPPI, or asks about evaluating this task. Reports Root Mean Square Error (RMSE).
Evaluates LLM capabilities across the full academic peer review lifecycle, including predicting paper acceptance and scores, generating structured peer reviews, and simulating multi-turn author-reviewer rebuttal conversations. Use when the user wants to benchmark on Re$^2$, or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models' ability to comprehend long-form sequential manga narratives, focusing on story synthesis, character grounding, and temporal reasoning across non-linear, multi-panel sequences. Use when the user wants to benchmark on Re:Verse, or asks about evaluating this task. Reports BERTScore.
Evaluates multimodal LLMs' ability to reason across multiple images, including sequential and set-based consumption, interleaved text-image processing, and heterogeneous visual inputs like charts, equations, maps, and code. It probes cross-image contextual integration, precise visual reading, and step-by-step logical deduction. Use when the user wants to benchmark on ReMI, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates how well text-to-image models adhere to structured prompts by measuring the alignment between generated images and their corresponding captions. It probes the model's ability to preserve semantic details and follow prompt structure during fine-tuning. Use when the user wants to benchmark on Re-LAION-Caption 19M, or asks about evaluating this task. Reports VQA LLaVA.
Evaluates LLMs' zero-shot capability to classify outcome types and extract precise numerical values from randomized controlled trial reports. It probes the models' numerical reasoning, information extraction robustness, and suitability for automating meta-analysis pipelines. Use when the user wants to benchmark on RCT Numerical Extraction Dataset, or asks about evaluating this task. Reports exact_match_accuracy.
Evaluates whether direct classification of RAW sensor data achieves accuracy comparable to traditional RAW-to-RGB converted images, while measuring computational efficiency gains from skipping the conversion pipeline. Use when the user wants to benchmark on Custom RAW/RGB Dataset, or asks about evaluating this task. Reports top-1 classification accuracy.
Evaluates interpretability methods' ability to disentangle polysemantic language model representations by isolating causal attributes through activation interventions on residual stream features. Use when the user wants to benchmark on RAVEL, or asks about evaluating this task. Reports Disentanglescore.
Evaluates the fidelity of a simulated noisy speech generator against real VHF/UHF transmitted audio, and measures the downstream robustness of automatic speech recognition (ASR) models trained on the simulated data. Use when the user wants to benchmark on RATS Channel A, or asks about evaluating this task. Reports MSSL, WER.
This evaluation probes the robustness of end-to-end automatic speech recognition systems in noisy acoustic conditions. It measures how well a model preserves speech intelligibility and correctly transcribes utterances when background noise is present, specifically testing the mitigation of over-suppression artifacts during joint speech enhancement and recognition. Use when the user wants to benchmark on RATS Channel-A, or asks about evaluating this task. Reports WER(%).
This evaluation probes the robustness of automatic speech recognition (ASR) systems when trained on extremely limited in-domain noisy data. It measures how well a model can generalize to real-world noisy conditions by leveraging synthetic noisy data generated via a GAN, compared to traditional data augmentation and fine-tuning baselines. Use when the user wants to benchmark on RATS (Channel A), or asks about evaluating this task. Reports WER (%).