Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,185–8,208 of 20,865 skills
Probes LLMs' cross-cultural emotion understanding by testing their ability to predict emotions and sentiments across six languages, specifically examining how prompt language and explicit country context influence model performance. Use when the user wants to benchmark on CULEMO, or asks about evaluating this task. Reports emotion prediction.
Evaluates lane detection performance in diverse urban and highway scenarios using an F1-measure based on IoU between predicted and ground truth lane lines. Use when the user wants to benchmark on CULane, or asks about evaluating this task. Reports F1-measure.
Evaluates Chinese language understanding and generation capabilities across a hierarchical framework. It probes discourse comprehension, conversational interaction, mathematical reasoning, and multilingual tasks using a multi-level scoring strategy that normalizes model performance against a fixed baseline. Use when the user wants to benchmark on CUGE (lite version), or asks about evaluating this task. Reports normalized capability performance.
This evaluation probes the per-evidence-item utility and trace sensitivity in single-shot retrieval-augmented generation. It measures how removing, replacing, or duplicating retrieved context chunks affects answer correctness, grounding faithfulness, confidence calibration, and reasoning trace stability. Use when the user wants to benchmark on HotpotQA (distractor setting), 2WikiMultihopQA, or asks about evaluating this task. Reports Soft Correctness.
Evaluates contextual code embeddings on six Python source-code understanding tasks, including classification of variable misuse, incorrect binary operators, swapped operands, function-docstring mismatches, and exception types, plus a joint localization and repair task. Use when the user wants to benchmark on ETH Py150 Open Benchmarks, or asks about evaluating this task. Reports classification accuracy.
Evaluates a model's ability to identify and extract relevant text spans from legal contracts corresponding to specific clause categories. It probes domain-specific information extraction and needle-in-a-haystack detection under severe class imbalance. Use when the user wants to benchmark on CUAD, or asks about evaluating this task. Reports Precision@80% Recall.
Evaluates continual test-time adaptation (CTTA) for text understanding across sequential, unobserved domains. It probes a model's ability to adapt to shifting domains using only unlabeled test data while mitigating error accumulation and maintaining cross-domain generalization. Use when the user wants to benchmark on CTTA-Text-Understanding-Benchmark, or asks about evaluating this task. Reports exact match (EM), F1 score.
Evaluates click-through rate (CTR) prediction models for their ability to maximize economic welfare in simulated and real-world ad auction settings, while also measuring standard classification performance. Use when the user wants to benchmark on Synthetic Dataset, Criteo Display Advertising Challenge, or asks about evaluating this task. Reports test-time welfare.
Evaluates the ability of deep learning models to predict click-through rates (CTR) from sparse, high-dimensional categorical features in advertising and recommendation scenarios. It probes how well models capture multi-scale semantic interactions and handle large-scale, imbalanced binary classification tasks typical of real-world ad systems. Use when the user wants to benchmark on Avazu, MovieLens, Weibo, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to predict the next item a user will click or review based on their historical interaction sequence. It probes hierarchical user interest modeling across multiple product dimensions and abstraction levels in recommendation systems. Use when the user wants to benchmark on Amazon Review (Six-Category, Kindle Shop, Electronics), or asks about evaluating this task. Reports AUC.
Evaluates whether neural machine translation models correctly rely on contextual cues when generating target tokens. It compares model-extracted cue-target pairs against human-annotated discourse-level expectations to measure the plausibility of context reliance. Use when the user wants to benchmark on SCAT+, or asks about evaluating this task. Reports Macro F1.
This benchmark evaluates the zero-shot multi-abnormality detection capability of a visual-language foundation model on 3D chest CT volumes. It probes the model's ability to generalize to unseen data distributions and classify multiple pathologies simultaneously without task-specific supervised training. Use when the user wants to benchmark on CT-RATE, RAD-ChestCT, or asks about evaluating this task. Reports AUROC.
Evaluates the ability of segmentation models to accurately delineate brain tissue, cerebrospinal fluid (CSF), and subdural hematomas in post-operative CT scans of hydrocephalic infants. It probes robustness to intensity overlap, anatomical distortion, and limited training data in a real-world clinical setting. Use when the user wants to benchmark on CURE Children's Hospital of Uganda CT Brain Dataset, or asks about evaluating this task. Reports dice-overlap coefficient.
Evaluates compositional symbolic music reasoning by requiring models to chain atomic analyses across multiple musical dimensions (e.g., rhythm, harmony, key, structure) to answer multiple-choice questions derived from expert forums and professional exams. Use when the user wants to benchmark on CSyMR-Bench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the scientific reasoning and domain-grounded visual question answering capabilities of Vision-Language Models (VLMs) in Chinese. It probes the ability to integrate multimodal STEM evidence across physics, chemistry, biology, and mathematics with domain knowledge to solve both multiple-choice and open-ended questions. Use when the user wants to benchmark on CSVQA, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of synthesized speech from single-speaker TTS models trained on the CSS10 datasets across 10 languages. It probes how well models can reproduce natural-sounding audio and accurate pronunciation for held-out test sentences. Use when the user wants to benchmark on CSS10, or asks about evaluating this task. Reports MOS.
This evaluation protocol assesses the ranking performance and robustness of sequential recommendation models trained with confident soft labels. It measures whether predicted item sequences align with actual user interactions and verifies if recommendations correspond to genuinely positive user preferences using explicit rating thresholds. Use when the user wants to benchmark on Last.FM, Yelp, Amazon Electronics, Amazon Movies and TV, or asks about evaluating this task. Reports Recall@n, NDCG@n.
Evaluates the robustness of information retrieval models when processing code-switched queries (English mixed with Mandarin Chinese or Japanese). It probes whether multilingual retrievers and rerankers suffer embedding divergence or performance degradation compared to monolingual English queries across argument, code, biomedical, and instruction-following retrieval tasks. Use when the user wants to benchmark on Touché 2020, HumanEval, TRECCOVID, FollowIR, or asks about evaluating this task. R...
Evaluates a transformer-based framework for Channel State Information (CSI) time-series prediction and wireless sensing classification. It probes the model's ability to recover missing data, predict future CSI sequences, and classify human actions or environmental states from Wi-Fi signals. Use when the user wants to benchmark on WiGesture, WiFall, WiCount, CommPre, or asks about evaluating this task. Reports Accuracy.
Evaluates continual learning capabilities in scene graph generation by measuring how models retain prior object-relationship knowledge while learning new tasks, handle long-tailed data distributions, and generalize to unseen objects and relationships across incremental learning scenarios. Use when the user wants to benchmark on CSEGG, or asks about evaluating this task. Reports Avg. R@20.
Evaluates models on ordinal classification of mammographic masking potential (levels 1–8) and their clinical utility in predicting interval and large invasive cancers. It probes the model's ability to respect ordinal relationships in breast tissue obscuration and correlate these estimates with cancer outcomes. Use when the user wants to benchmark on CSAW-M, or asks about evaluating this task. Reports average mean absolute error (AMAE).
Evaluates a model's ability to detect and localize multiple spoken keywords within continuous, untrimmed audio streams, distinguishing target keywords from background speech and silence. Use when the user wants to benchmark on LibriTop-20, CMAK-7, or asks about evaluating this task. Reports mAP.
This benchmark evaluates automatic speech recognition (ASR) systems on their ability to accurately transcribe spontaneous, full-length dialogues that alternate between Mandarin and English. It probes a model's robustness to language alternation, phonetic mismatches, and contextual dependencies in naturalistic code-switching scenarios. Use when the user wants to benchmark on CS-Dialogue, or asks about evaluating this task. Reports MER.
Evaluates LLMs on end-to-end computer science research workflows by testing their ability to answer scientific questions grounded in academic papers. It probes domain-specific reasoning, factual recall, and methodological understanding across eight research workflow categories. Use when the user wants to benchmark on CS-4k, or asks about evaluating this task. Reports model response score.