Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,097–9,120 of 21,245 skills
Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.
Evaluates a model's ability to recognize and classify temporal surgical phases in cataract surgery videos, specifically testing classification accuracy and robustness to domain shift across different clinical centers. Use when the user wants to benchmark on Cataract-LMM Phase Recognition Subset, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large language models' ability to detect and mitigate educational safety risks in a student-tailored manner. It probes how well models adapt their responses to individual student profiles across 15 risk domains and 14 psychological and educational attributes. Use when the user wants to benchmark on CASTLE, or asks about evaluating this task. Reports safety score.
Evaluates deep learning models for drug-target interaction prediction across four tasks: scoring (affinity correlation), ranking (pose affinity ordering), docking (native pose retrieval from decoys), and screening (true binder identification among random molecules). Probes both regression accuracy and virtual screening generalization. Use when the user wants to benchmark on CASF-2016, CSAR NRC-HiQ, or asks about evaluating this task. Reports Docking power (Top-1 Success Rate), Screening power...
This evaluation protocol assesses a model's ability to perform session-based next-item recommendation by predicting the subsequent item a user will click based on their recent interaction history. It probes the model's capacity to capture both short-term sequential dependencies and long-term contextual patterns within a session without relying on explicit user profiles. Use when the user wants to benchmark on Yoochoose1/64, Yoochoose1/4, Diginetica, or asks about evaluating this task. Reports...
Evaluates the ability of language models to generate accurate, concise, and legally faithful summaries of long U.S. Supreme Court opinions. It probes the alignment between automatic NLP metrics and expert human judgment in a high-stakes, domain-specific summarization task. Use when the user wants to benchmark on CaseSumm, or asks about evaluating this task. Reports ROUGE.
Evaluates a model's ability to capture sequential user behavior patterns for personalized top-N item recommendation. It probes the model's capacity to model temporal dependencies, skip behaviors, and union-level sequential patterns from historical interactions to predict future items. Use when the user wants to benchmark on MovieLens, Gowalla, Foursquare, Tmall, or asks about evaluating this task. Reports MAP.
This benchmark probes a model's ability to perform fine-grained legal text classification and annotation. It tests whether models can accurately extract specific legal features, such as precedent alteration, issue areas, or ideological valence, from lengthy court opinions using multiple-choice prompts. Use when the user wants to benchmark on CaselawQA, or asks about evaluating this task. Reports accuracy.
Evaluates LLMs and retrieval models on verifying colloquial legal claims against U.S. Supreme Court precedents, measuring both verdict prediction accuracy and the quality of retrieved supporting case evidence. Use when the user wants to benchmark on CaseFacts, or asks about evaluating this task. Reports Verdict Score.
Evaluates a graph-based neural architecture for tabular learning on single and multiple tables, testing its ability to handle mixed numerical and categorical features without requiring schema or entity matching. Use when the user wants to benchmark on TabLLM datasets, Entity Matching datasets, or asks about evaluating this task. Reports performance.
Evaluates the visual classification and vision-language understanding capabilities of large vision-language models (LVLMs) under context-aware ensemble prompting. It probes fine-grained visual recognition, scientific question answering, text-rich VQA, hallucination detection, and multimodal reasoning across diverse benchmarks. Use when the user wants to benchmark on ImageNet, Caltech101, Flower102, Food101, ScienceQA (image subset), TextVQA, POPE, MME, MMBench, CV-Bench, MMVP, or asks about e...
Evaluates the reconstruction quality of Neural Radiance Field (NeRF) models on synthetic vehicle components. It probes both 2D appearance fidelity and 3D geometric accuracy, specifically testing robustness to reflective surfaces, transparent materials, and varying numbers of training viewpoints. Use when the user wants to benchmark on CarPatch, or asks about evaluating this task. Reports PSNR.
Evaluates 3D occupancy prediction models on their ability to reconstruct semantic and instance-level voxel grids from camera inputs. It probes geometric completeness, occlusion reasoning, and instance discrimination in complex autonomous driving scenes. Use when the user wants to benchmark on CarlaOcc, or asks about evaluating this task. Reports mIoU.
Evaluates an autonomous driving agent's ability to navigate urban environments, avoid dynamic and static obstacles, and handle road blockages under varying weather conditions and unseen towns. Use when the user wants to benchmark on CARLA urban driving benchmark, or asks about evaluating this task. Reports Success Rate.
Evaluates the ability of reinforcement learning policies to execute tactical driving maneuvers (e.g., lane changes, roundabout navigation) in a simulated environment mapped from real-world traffic data. It probes how observation modalities and reward structures impact policy generalization and success rates across diverse driving scenarios. Use when the user wants to benchmark on NGSIM, openDD, or asks about evaluating this task. Reports success rate.
Evaluates an autonomous driving model's ability to navigate complex urban and highway environments under varying weather and lighting conditions, focusing on route completion, safety (infraction avoidance), and overall driving performance. Use when the user wants to benchmark on CARLA Longest6, CARLA Town05 Long, or asks about evaluating this task. Reports Driving Score (DS).
Evaluates autonomous driving agents on their ability to navigate complex urban environments while balancing route completion, safety, and rule compliance. It probes how well models handle dynamic traffic interactions, edge-case scenarios, and physical constraints in a closed-loop simulation. Use when the user wants to benchmark on CARLA Leaderboard 2.0 scenarios, CARLA 42 Routes, Town05, or asks about evaluating this task. Reports Driving Score (DS).
Evaluates end-to-end autonomous driving agents on route completion, safety/infractions, and occlusion-aware perception in simulated urban environments. Use when the user wants to benchmark on CARLA Town 05 Long, DOS Benchmark, or asks about evaluating this task. Reports Driving Score (DS).
Evaluates the quality and feasibility of model-agnostic counterfactual explanations generated for tabular data. It probes whether generated counterfactuals successfully flip classifier predictions while maintaining sparsity, proximity, actionability (immutable constraints), and plausibility across multiple binary classification tasks. Use when the user wants to benchmark on adult, COMPAS, Give Me Some Credit, HELOC, Irish, Saheart, Titanic, Wine, or asks about evaluating this task. Reports Su...
Evaluates the end-to-end inference latency of an autonomous driving pipeline that runs parallel reinforcement learning and object detection models in a simulated urban environment. It measures how efficiently the middleware handles communication and computation overhead during real-time sensor processing and action fusion. Use when the user wants to benchmark on CARLA simulator, or asks about evaluating this task. Reports inference_latency.
Evaluates large language models' causal reasoning capabilities across three task categories: causal graph reasoning (adjacency matrix, d-separation, causal direction), knowledge discovery from tabular data, and decision-making under interventions and counterfactuals. Use when the user wants to benchmark on CARL-GT, or asks about evaluating this task. Reports F1 score.
Evaluates the trustworthiness of Medical Large Vision-Language Models across five dimensions: trustfulness (factuality and uncertainty), fairness (demographic disparities), safety (jailbreaking, toxicity, overcautiousness), privacy, and robustness. It probes the models' ability to generate accurate medical information, recognize their own uncertainty, avoid demographic bias, resist adversarial prompts, and handle sensitive data without leakage. Use when the user wants to benchmark on CARES, I...
This benchmark evaluates large language models' ability to perform critical appraisal and methodological reasoning on biomedical scientific articles. It probes whether models can correctly identify study design flaws, statistical limitations, and biases by answering multiple-choice questions derived from authentic French medical exams. The evaluation specifically measures both exact correctness and partial reasoning accuracy under varying context conditions. Use when the user wants to benchma...
Evaluates multimodal fusion of Electronic Health Records (EHR) and Chest X-Rays (CXR) for clinical decision support, specifically testing robustness to missing modalities, temporal imbalance, and subgroup fairness across phenotyping, mortality, and length-of-stay prediction. Use when the user wants to benchmark on CareBench, or asks about evaluating this task. Reports AUROC, AUPRC, F1, Accuracy, Cohen's Kappa.