Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,905–8,928 of 21,231 skills
This benchmark evaluates LLM-based agents on autonomous, open-ended climate science problem-solving. It probes the model's ability to perform data-driven modeling, apply physics-aware constraints, and generate scientifically rigorous analysis reports without human intervention. Use when the user wants to benchmark on ClimaBench, or asks about evaluating this task. Reports Overall.
Evaluates a model's ability to accumulate knowledge across a sequence of NLP tasks (continual learning) while maintaining performance on previously seen tasks and generalizing to new few-shot tasks. Use when the user wants to benchmark on CLIF-26, CLIF-55, or asks about evaluating this task. Reports Final Accuracy.
This benchmark evaluates zero-shot detection of AI-generated images across general and domain-specific settings. It probes a model's ability to distinguish real from synthetic images without task-specific fine-tuning, measuring robustness to domain shifts (e.g., artistic styles, damaged cars, invoices) and resistance to 'flipped classification' where detectors misrank generated content as real. Use when the user wants to benchmark on General Image Benchmark (LAION + MS-COCO), ImaginET, CarDD,...
Evaluates the ability to detect fraudulent ad clicks (clickspam) by analyzing temporal reuse patterns in organic clickstreams. It tests both passive traffic analysis and active bait-click injection strategies to distinguish legitimate user behavior from automated or malware-driven fraud. Use when the user wants to benchmark on University Network Click Traffic Dataset, or asks about evaluating this task. Reports FPR.
This evaluation probes a recommender system's ability to mitigate clickbait by measuring performance exclusively on user interactions that result in positive post-click feedback (likes), rather than raw click-through rates. Use when the user wants to benchmark on Unspecified in provided section, or asks about evaluating this task. Reports post-click satisfaction (likes).
Evaluates language models' proficiency in Korean cultural knowledge (e.g., history, law, society, economy) and linguistic competence (e.g., grammar, functional usage) using multiple-choice questions sourced from official Korean exams and textbooks. Use when the user wants to benchmark on CLIcK, or asks about evaluating this task. Reports accuracy.
CliBench evaluates large language models on real-world clinical decision-making tasks, including diagnosis, procedure recommendation, lab test ordering, and medication prescribing. It probes the models' ability to process complex patient records, generate structured medical codes, and maintain coherence across multi-step clinical workflows in a zero-shot setting. Use when the user wants to benchmark on CliBench (MIMIC-IV derived), or asks about evaluating this task. Reports micro F1.
Evaluates 3D visual question answering capabilities on point cloud scenes, probing spatial reasoning, object recognition, and scene graph understanding without relying on common-sense spatial priors. Use when the user wants to benchmark on CLEVR3D, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates a model's ability to comprehend referring expressions in synthetic visual scenes. It probes compositional visual reasoning by measuring how well models localize objects based on text descriptions that vary in attribute complexity, spatial relationships, and reasoning topology. Use when the user wants to benchmark on CLEVR-Ref+, or asks about evaluating this task. Reports IoU.
Evaluates end-to-end formally verified code generation by requiring models to produce both a logically equivalent formal specification and a provably correct implementation in Lean 4. It probes the model's ability to reason about non-computable specifications, synthesize machine-checkable proofs, and ensure semantic correctness beyond syntactic compilation. Use when the user wants to benchmark on CLEVER, or asks about evaluating this task. Reports pass@600-seconds.
Evaluates Chinese large language models across six key dimensions: accuracy, robustness, fairness, calibration, bias, and diversity. The platform uses standardized prompts and dynamic test set sampling to mitigate train-test contamination while comparing open-source and limited-access models. Use when the user wants to benchmark on CLEVA benchmark suite, or asks about evaluating this task. Reports Accuracy.
Evaluates perception models' ability to estimate 6 DoF poses and complete depth maps for transparent and translucent objects. It specifically probes robustness to challenging real-world conditions such as heavy occlusion, cluttered backgrounds, varying lighting, and objects filled with liquid. Use when the user wants to benchmark on ClearPose, or asks about evaluating this task. Reports 6 DoF poses.
Evaluates a model's ability to continuously learn from non-stationary data streams without catastrophic forgetting, balancing stability and plasticity in regression tasks. Use when the user wants to benchmark on Artificial periodic dataset, Wind power generation dataset, or asks about evaluating this task. Reports prediction error.
Evaluates embodied cleaning agents in physics-accurate indoor simulations, probing their ability to perform sweeping and grasping tasks across diverse cluttered scenes. It measures task completion, spatial coverage efficiency, motion quality, and collision safety under strict time limits. Use when the user wants to benchmark on CleanUpBench, or asks about evaluating this task. Reports TCR.
This evaluation probes a model's ability to remove background noise from speech signals directly in the waveform domain. It measures perceptual quality, speech intelligibility, and computational efficiency using standardized objective metrics and crowdsourced subjective listening tests. Use when the user wants to benchmark on DNS dataset, Valentini dataset, Internal dataset, or asks about evaluating this task. Reports PESQ-MOS (SIG/BAK/OVRL).
Evaluates LLMs' ability to extract structured Causal Loop Diagrams (CLDs) from natural language system dynamics descriptions. It probes structured output generation, schema conformance, and iterative model updating under varying context lengths and prompt strategies. Use when the user wants to benchmark on CLD Leaderboard, or asks about evaluating this task. Reports exact_structured_match.
Evaluates the runtime performance and energy efficiency of different programming language implementations. It specifically compares Lua interpreters, LuaJIT JIT compilers, and C on computationally intensive benchmark programs from the Computer Language Benchmarks Game. Use when the user wants to benchmark on CLBG (Computer Language Benchmarks Game), or asks about evaluating this task. Reports Energy Consumption.
Measures end-to-end robotic manipulation performance and grasp robustness by clearing a bin of soft objects using a learned policy. It evaluates the ability to predict optimal grasping poses from RGB-D inputs and execute them across different hardware platforms. Use when the user wants to benchmark on Soft toy bin-clearing set, or asks about evaluating this task. Reports r_success.
Evaluates large language models' ability to generate complete Python classes with interdependent methods, rather than standalone functions. It probes long-context code reasoning, dependency modeling, and the effectiveness of holistic versus incremental generation strategies. Use when the user wants to benchmark on ClassEval, or asks about evaluating this task. Reports pass@1.
Predicts the outcome (win or lose) of U.S. class action lawsuits based on plaintiff complaint texts. It probes a model's ability to extract legally relevant allegations from long-form, unverified legal documents and make binary judgment predictions. Use when the user wants to benchmark on ClassActionPrediction, or asks about evaluating this task. Reports accuracy.
This benchmark probes multimodal large language models' ability to detect cross-modal contradictions between images and text. It evaluates whether models can identify inconsistencies when either modality contains errors or hallucinations, rather than assuming one modality is ground truth. The task reveals systematic modality biases and category-specific reasoning weaknesses. Use when the user wants to benchmark on CLASH, or asks about evaluating this task. Reports accuracy.
Evaluates large language models on European Portuguese across cultural alignment, safety safeguards, chain-of-thought reasoning, natural language understanding, and common NLU tasks. It probes how well models handle culture-specific implicit knowledge, refuse harmful requests, and perform multiple-choice or generative QA in Portuguese. Use when the user wants to benchmark on Tuguesice-PT, DoNotAnswer-PT, MuSR, AA-Omniscience-Public, GPQA Diamond, MMLU, MMLU Pro, CoPA, MRPC, RTE, or asks about...
Evaluates neural scene reconstruction and semantic segmentation capabilities on aerial UAV imagery. Probes the model's ability to generate high-fidelity 3D reconstructions, depth maps, and class-aware segmentation masks from multi-view inputs under varying scene complexities and viewpoint distributions. Use when the user wants to benchmark on ClaraVid, UAVid, or asks about evaluating this task. Reports reconstruction results.
This benchmark evaluates query-conditioned target sound extraction (TSE), testing a model's ability to isolate a target audio source from a mixture using language captions or reference audio queries. It probes multi-modal query processing and positive/negative query valence across diverse acoustic environments and musical instruments. Use when the user wants to benchmark on AudioCaps, AudioSet, ESC-50, FSDKaggle2018, MUSIC21, or asks about evaluating this task. Reports SDRi, SISDRi.