Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,041–8,064 of 20,861 skills
Evaluates sentiment classification performance across different English dialects (en-US, en-AU, en-UK, en-IN) and tests how label proximity, review length, and sentiment density affect model generalization. Use when the user wants to benchmark on Google Place Reviews (Dialectal Sentiment), or asks about evaluating this task. Reports F1-Score.
Evaluates a robot's ability to perform contact-based manipulation and learn a continuous control policy via reinforcement learning to match a target angle on a potentiometer. Use when the user wants to benchmark on DeltaZ Dial Turning Task, or asks about evaluating this task. Reports reward.
Clinical diagnostic reasoning capability of LLMs, requiring them to generate plausible diagnoses from patient case descriptions and imaging/symptom details. It probes the model's ability to perform complex, multi-step medical deduction and generalize across 28 clinical specialties. Use when the user wants to benchmark on DiagnosisArena, or asks about evaluating this task. Reports accuracy.
Binary classification of lexical semantic change for target words across two diachronic time periods. It probes whether models can reliably detect meaning shifts in Italian using corpus pairs from newspapers and books. Use when the user wants to benchmark on DIACR-Ita, or asks about evaluating this task. Reports classification accuracy.
This benchmark evaluates the ability of machine learning models to perform binary classification on free-text electronic health record (EHR) progress notes related to diabetes. It probes how well different architectures (CNNs, RNNs, SVMs, hybrids) capture local linguistic patterns and generalize across different hospital datasets. Use when the user wants to benchmark on BWH/UTP Clinical Notes, or asks about evaluating this task. Reports AUC.
Evaluates the safety of conversational AI models by measuring their tendency to generate unsafe responses at both the utterance level and within conversational context. It specifically probes context-sensitive unsafety, where responses appear safe in isolation but become harmful when conditioned on prior dialogue history. Use when the user wants to benchmark on DiaSafety, or asks about evaluating this task. Reports proportion.
Evaluates a model's ability to perform multi-dimensional discourse analysis on Bengali climate news articles. It probes capabilities in stance detection, authenticity verification, political influence identification, and various information extraction tasks related to environmental reporting. Use when the user wants to benchmark on Dhoroni, or asks about evaluating this task. Reports F1 Score.
Evaluates the effectiveness of a deep hierarchical ensemble network for large-scale click-through rate (CTR) prediction. It probes the model's ability to capture complex, non-overlapping feature interactions across multiple layers and scale efficiently on industrial-scale data. Use when the user wants to benchmark on Industrial in-house dataset, or asks about evaluating this task. Reports Normalized Entropy (NE) loss.
Evaluates structured OCR extraction fidelity and text degeneration rates on printed, handwritten, and legal documents. Measures how well models adhere to JSON schemas while minimizing pathological generation loops. Use when the user wants to benchmark on DharmaOCR-Benchmark, or asks about evaluating this task. Reports Score.
Evaluates a model's ability to perform sequential next-item recommendation by modeling dynamic collaborative signals and temporal user preferences. It tests how well the system captures high-order item transitions and time-annotated graph structures to predict the next interaction in a user's history. Use when the user wants to benchmark on Amazon-CDs, Amazon-Games, Amazon-Beauty, or asks about evaluating this task. Reports NDCG@10.
Evaluates audio-visual models on their ability to separate target musical instrument sounds from mixed audio using synchronized video cues. It probes cross-modal feature alignment and dynamic fusion of audio and visual signals for source separation in complex environments. Use when the user wants to benchmark on MUSIC, MUSIC-21, or asks about evaluating this task. Reports SDR.
Evaluates automatic dynamic facial micro-expression recognition (MER) models on a large-scale spontaneous micro-expression dataset. It probes the model's ability to classify subtle, high-frame-rate facial movements across seven emotion categories while handling class imbalance and variable video lengths. Use when the user wants to benchmark on DFME, or asks about evaluating this task. Reports Accuracy (ACC).
Evaluates a unified dialogue foundation model across representation, knowledge distillation, and generation capabilities on diverse dialogue-oriented tasks. It probes the model's ability to perform intent detection, slot filling, semantic parsing, dialogue state tracking, text-to-SQL, and end-to-end task-oriented dialogue generation. Use when the user wants to benchmark on DialoGLUE, MULTIWOZ2.0, MULTIWOZ2.2, Spider, CoSQL, CLINC150, BANKING77, HWU64, RESTAURANT8K, DSTC8, TOP, PERSONALCHAT, C...
Evaluates the ability of a quantum annealer (D-Wave) to solve distributed flexible job shop scheduling problems (DFJSP) compared to classical simulated annealing. It probes solver performance in terms of solution quality (energy, makespan, constraint satisfaction) and computational efficiency (runtime scaling) across varying problem sizes. Use when the user wants to benchmark on Custom DFJSP instances (wool textile industry), or asks about evaluating this task. Reports System energy.
Evaluates the robustness of distractor-free novel view synthesis methods against large-scale, diverse distractor scenarios. It measures how well radiance field and 3D Gaussian Splatting models can reconstruct clean 3D scenes from cluttered or dynamically changing inputs without degrading static background quality. Use when the user wants to benchmark on DF3DV-1K, DF3DV-41, or asks about evaluating this task. Reports PSNR.
Evaluates joint perception and manipulation capabilities for hand-object interactions, specifically 2D detection, 6D object pose estimation, and 3D hand pose estimation on real-world RGB-D sequences. Use when the user wants to benchmark on DexYCB, or asks about evaluating this task. Reports precision-coverage.
This evaluation probes a robot policy's ability to successfully reproduce human-demonstrated dexterous manipulation trajectories in physics simulation. It measures robustness by testing performance under nominal conditions and under controlled initial pose perturbations. Use when the user wants to benchmark on DexCanvas, or asks about evaluating this task. Reports success rate.
Evaluates the capability of anomaly detection models to identify rare or deviant data points using only a small set of labeled anomalies as prior knowledge. It probes data efficiency, robustness to varying anomaly contamination levels in unlabeled training data, and the ability to rank anomalies effectively under severe class imbalance. Use when the user wants to benchmark on donors, census, fraud, celeba, backdoor, URL, campaign, news20, thyroid, or asks about evaluating this task. Reports A...
Evaluates the efficacy of covert watermarking techniques embedded in manuscript PDFs to force LLM-generated peer reviews to contain specific hidden markers. It also tests the robustness of these watermarks against common reviewer defenses like paraphrasing, detection prompts, and page cropping, as well as the performance of cryptic prompt injection via gradient-based optimization. Use when the user wants to benchmark on ICLR 2024 submissions, ICLR 2021 submissions, ICLR 2024 submissions (cont...
Evaluates multimodal large language models' ability to pinpoint erroneous content at the token level within long-form image captions. It probes whether models can distinguish between visually grounded facts and hallucinated details by localizing specific tokens that contradict the input image. Use when the user wants to benchmark on DetailVerifyBench, or asks about evaluating this task. Reports token-level F1.
Evaluates a text-to-image model's ability to faithfully render long, descriptive prompts. It probes fine-grained semantic alignment across character presence, attributes, spatial relationships, and scene composition, as well as overall aesthetic and alignment quality using preference models. Use when the user wants to benchmark on DetailMaster, or asks about evaluating this task. Reports CharacterPresence.
Evaluates a model's ability to generate detailed, region-specific descriptions for images and videos, ranging from keywords to multi-sentence captions. It probes fine-grained visual grounding, attribute recognition, and hallucination resistance by comparing generated text against reference captions or using attribute-level positive/negative judgments. Use when the user wants to benchmark on DLC-Bench, LVIS, PACO, Flickr30k Entities, Ref-L4, HC-STVG, VideoRefer-Bench-D, or asks about evaluatin...
Evaluates large language models' ability to perform 3D scene understanding tasks, including single and multi-object visual grounding, 3D scene captioning, and contextual question answering, using object-level text descriptions for relational reasoning. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.25 / Acc@0.5, F1@0.25 / F1@0.5, CIDEr@0.5 / CIDEr, EM / EM-R.
Evaluates the diagnostic accuracy and explainability of various Convolutional Neural Network architectures on dermatological image classification. It specifically probes how well different architectures localize clinically relevant skin characteristics using Grad-CAM heatmaps compared to human dermatologists. Use when the user wants to benchmark on DermXDB, or asks about evaluating this task. Reports image-level Grad-CAM F1 score.