Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,569–5,592 of 23,732 skills
Evaluates vision models on prosthesis-specific video understanding, including instance segmentation of amputees and prosthetic limbs, 2D human pose estimation with focus on lower-body keypoints, and automated gait pattern classification from pose sequences. Use when the user wants to benchmark on ProGait, or asks about evaluating this task. Reports mIoU, AP@[0.5,0.95].
This protocol evaluates the causal impact of specific profile image features (smile, body-shot, and gender) on lender selection preferences in a simulated micro-lending marketplace. It uses a conjoint-style choice experiment with GAN-generated images to isolate how visual cues influence funding decisions independent of borrower creditworthiness. Use when the user wants to benchmark on Custom GAN-generated profile images, or asks about evaluating this task. Reports Average Treatment Effect (ATE).
Evaluates reinforcement learning agents on their ability to learn efficiently and generalize to unseen, procedurally generated environments. It measures how well algorithms adapt to novel level distributions under strict computational and timestep constraints. Use when the user wants to benchmark on Procgen Benchmark, or asks about evaluating this task. Reports mean normalized return.
This evaluation probes the training efficiency and multi-GPU scaling behavior of deep learning frameworks. It measures how quickly models process mini-batches and how effectively data parallelization affects model convergence across various network architectures and hardware configurations. Use when the user has predictions and gold and needs to compute processing_time.
Evaluates multimodal foundation models on open-ended, expert-level queries across 10 professional domains, probing visual perception, domain knowledge, and long-context reasoning in single-round, multi-lingual, and multi-turn settings. Use when the user wants to benchmark on ProBench, or asks about evaluating this task. Reports ELO rating.
Evaluates real-time probabilistic forecasting of financial and weather time series, probing a model's ability to quantify uncertainty via quantile modeling and maintain calibration over sequential submission rounds. Use when the user wants to benchmark on DAX, Wind, Temperature, or asks about evaluating this task. Reports skill score.
Evaluates the skill of a probabilistic Random Forest model in forecasting severe thunderstorms (tornadoes, large hail, damaging winds) 4–8 days in advance using ensemble meteorological data. It probes the model's calibration, discrimination, and spatial coverage compared to human-generated SPC outlooks. Use when the user wants to benchmark on SPC Severe Weather Reports & GEFSv12 Reforecast, or asks about evaluating this task. Reports Brier Skill Score (BSS).
Assesses an LLM agent's ability to understand and follow privacy norms while performing real-world tasks. It measures both helpfulness and the rate at which sensitive information is incorrectly exposed. Use when the user wants to benchmark on PrivacyLens, or asks about evaluating this task. Reports privacy leakage rate.
Evaluates the trade-offs between privacy preservation, model utility, and computational/energy costs in hybrid privacy-preserving vision systems. It probes how combining federated learning with differential privacy or secure multi-party computation affects convergence, classification accuracy, and resource consumption across different neural architectures. Use when the user wants to benchmark on Alzheimer MRI Classification, ISIC Skin Lesion Classification, or asks about evaluating this task....
Evaluates large multimodal models' ability to detect, correct, and reason over real-world multimodal inconsistencies in scientific papers. It probes inter-modal mismatch detection, structured reasoning, and robustness to linguistic shortcuts versus genuine visual grounding. Use when the user wants to benchmark on PRISMM-Bench, or asks about evaluating this task. Reports Accuracy (%).
Probes LLM hallucinations across four dimensions (knowledge missing, knowledge errors, reasoning errors, and instruction-following errors) by isolating error sources through controlled, task-specific queries. It evaluates model reliability and guides optimization by measuring error rates across memory, instruction, and reasoning generation stages. Use when the user wants to benchmark on PRISM, or asks about evaluating this task. Reports H-Score.
Evaluates fine-grained, multi-aspect-aware paper-to-paper retrieval by decomposing long-form query papers into aspect-specific views and segmenting candidate papers into section-level representations for targeted retrieval. Use when the user wants to benchmark on SciFullBench, PatentFullBench, or asks about evaluating this task. Reports Recall@K.
Evaluates the ability of sentiment lexicons or models to assign accurate real-valued polarity scores to individual terms, measuring rank correlation with gold standards. Use when the user wants to benchmark on Term test set, or asks about evaluating this task. Reports Kendall's τ coefficient.
Evaluates the effectiveness of various prior-based loss functions (low-level boundary/distance and high-level shape/size constraints) for medical image segmentation across diverse anatomical structures and imaging modalities. Use when the user wants to benchmark on WMH, ISLES, Atrium, Colon, Spleen, Hippocampus, Prostate, ACDC, or asks about evaluating this task. Reports Dice score.
Evaluates large language models' ability to reason about medical ethics using the Principlism framework (autonomy, non-maleficence, beneficence, justice). It probes both theoretical knowledge of ethical principles and their practical application to complex, open-ended clinical dilemmas. Use when the user wants to benchmark on PrinciplismQA, or asks about evaluating this task. Reports Knowledge accuracy, Practice score.
Probes an LLM's ability to align generated responses with a set of natural language constitutional principles without parameter fine-tuning. It measures both overall conformance quality and the reduction of critical principle violations through an inference-time self-correction pipeline. Use when the user wants to benchmark on SafeRLHF, HH-RLHF, or asks about evaluating this task. Reports 5-Point Likert Score Ranking.
Evaluates the quality of Semantic Role Labeling (SRL) systems by measuring precision and recall for predicate senses and argument annotations. It specifically probes a model's step-dependent error propagation by penalizing argument scores when the associated predicate sense is incorrect, while also handling discontinuous and reference arguments. Use when the user has predictions and gold and needs to compute PriMeSRL-Eval.
Evaluates a pre-trained seismic model's multi-task capability on single-station waveforms, specifically phase picking (Pg, Sg, Pn, Sn), P-wave polarization classification, and seismic event type classification. The protocol tests generalization across temporal splits and transfer learning on local data to mitigate dataset imbalance. Use when the user wants to benchmark on CSNCD, or asks about evaluating this task. Reports recall.
Evaluates multi-objective hyperparameter optimization algorithms on deep learning benchmarks, measuring their ability to find high-quality Pareto fronts of validation error and training cost under varying prior conditions and budget constraints. Use when the user wants to benchmark on Yahpo-Gym & PD1 HPO Benchmarks, or asks about evaluating this task. Reports mean dominated hypervolume.
Evaluates a privacy-regulated speech-to-speech translation framework that replaces voice cloning with matching to pre-consented preset voices. Probes classifier robustness, cross-lingual speech naturalness, and inference efficiency across multilingual scenarios. Use when the user wants to benchmark on RAVDESS, CGDD, CAFE, EmoDB, CREMA-D, or asks about evaluating this task. Reports NISQA.
Evaluates a model's ability to classify the sense of a preposition in context. It tests cross-lingual context representation and semi-supervised learning for fine-grained lexical disambiguation. Use when the user wants to benchmark on Web-reviews corpus, SemEval corpus, or asks about evaluating this task. Reports accuracy.
Measures an agent's capability to retain and adhere to user preferences during long, multi-turn conversations. It tests whether the model can maintain consistency without external reminders or with explicit preference cues. Use when the user wants to benchmark on PrefEval, or asks about evaluating this task. Reports preference retention accuracy.
This benchmark evaluates a model's ability to dynamically adapt to evolving user preferences by conditioning on natural language preferences inferred from interaction history. It probes recommendation accuracy, fine- and coarse-grained preference steering, sentiment following, and history consolidation across multiple e-commerce and gaming datasets. Use when the user wants to benchmark on Amazon Beauty, Amazon Sports and Outdoors, Amazon Toys and Games, Steam, or asks about evaluating this ta...
Evaluates a model's ability to identify and rank key moments (shots) in soccer match videos for summarization. It measures how well the model selects representative content when constrained to match the exact duration of a human-curated highlight summary. Use when the user has predictions and gold and needs to compute F1 Score@$T$.