Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,465–4,488 of 23,503 skills
Evaluates language model capabilities across knowledge, reasoning, instruction following, and safety using a standard suite of downstream benchmarks. It measures both pretraining quality and post-adaptation performance on established NLP and coding tasks. Use when the user wants to benchmark on MMLU, HellaSwag, ARC-Challenge, ARC-Easy, PIQA, WinoGrande, GSM8k, BBH, HumanEval, AlpacaEval 1.0, XSTest, IFEval, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates a model's ability to classify stance (in favor, against, neutral) toward climate change and related targets, while jointly predicting sentiment (positive, negative, neutral). It probes the synergy between stance classification and sentiment analysis using text and topic features. Use when the user wants to benchmark on Climate Change Tweets Dataset, SemEval-2016 Task 6.A, or asks about evaluating this task. Reports F1 score.
This benchmark evaluates a model's ability to classify the stance (InFavor, Against, or None) of text towards a specific target or query. It specifically probes cross-target generalization by training on multiple targets and testing on a held-out target, while also assessing robustness to sarcastic or figurative language through intermediate sarcasm pre-training. Use when the user wants to benchmark on SemEval 2016 Task 6A Dataset, Multi-Perspective Consumer Health Query Data (MPCHI), or asks...
Evaluates a model's ability to generate fluent, contextually accurate Japanese image captions directly from visual input. It specifically probes whether native-language training data yields better captioning quality compared to a pipeline of English generation followed by machine translation. Use when the user wants to benchmark on STAIR Captions, or asks about evaluating this task. Reports CIDEr.
Tests retrieval-augmented question answering over long screenplay contexts. It probes the model's ability to locate relevant information across chunked text and synthesize accurate answers under different retrieval architectures. Use when the user wants to benchmark on STAGE-QA, or asks about evaluating this task. Reports Question Correctness.
Evaluates an LLM's ability to extract and structure narrative knowledge from full-length movie screenplays into canonical graphs. It probes entity and relation recognition while filtering out event-centric noise to focus on stable world-building elements. Use when the user wants to benchmark on STAGE-KG, or asks about evaluating this task. Reports Entity F1.
Assesses an LLM's capacity to abstract scene-level events into concise, free-form descriptions without schema constraints. It evaluates whether the generated events form a coherent, non-redundant structure and remain factually grounded in the screenplay text. Use when the user wants to benchmark on STAGE-ES, or asks about evaluating this task. Reports Event-Structure Consistency.
Evaluates multimodal models' ability to detect unintended content, structural, and low-level appearance changes in image-to-image transitions. It probes fine-grained visual reasoning and pixel-level alignment capabilities by asking models to assess fidelity across three distinct dimensions. Use when the user wants to benchmark on StableI2I-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the classification performance of SSVEP-based Brain-Computer Interface algorithms using Riemannian geometry on EEG covariance matrices. It probes the robustness of different covariance estimators and online/offline classification pipelines under varying trial lengths, latency delays, and outlier conditions. Use when the user wants to benchmark on SSVEP BCI dataset (12 subjects), or asks about evaluating this task. Reports classification accuracy.
Evaluates the binary sentiment classification capability of hybrid quantum-classical language models on both short and long text sequences. It probes whether adaptive quantum routing and attention mechanisms provide measurable accuracy gains over purely classical or purely quantum baselines on standard NLP benchmarks. Use when the user wants to benchmark on SST-2, IMDB, or asks about evaluating this task. Reports Accuracy.
Evaluates text classification performance on sentence sentiment analysis and news topic categorization. Probes the model's ability to capture non-linear, non-consecutive word interactions (e.g., negation, long-range dependencies) for accurate document/sentence-level prediction. Use when the user wants to benchmark on Stanford Sentiment Treebank (Fine-grained), Stanford Sentiment Treebank (Binary), Sogou Chinese News Corpora, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform sentiment classification on constituent trees, testing both fine-grained (5-class) and binary sentiment prediction at the sentence root and phrase levels. Use when the user wants to benchmark on Stanford Sentiment Treebank, TREC, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify the sentiment of movie review phrases into five fine-grained categories, measuring classification accuracy and error rates. The benchmark probes hierarchical sentiment understanding at the phrase level rather than the full sentence level. Use when the user wants to benchmark on Stanford Sentiment Treebank (SST), or asks about evaluating this task. Reports Error Rate (Fine-Grained).
Evaluates single-step retrosynthesis capability by predicting reactant molecules from a given target product, testing both in-distribution chemical knowledge and out-of-distribution generalization. Use when the user wants to benchmark on USPTO-50K-test, URSA-expert-2026, or asks about evaluating this task. Reports Unique.
Evaluates vision-language models on spatial understanding and general question answering using image-text pairs. It probes capabilities such as object existence, attribute recognition, action identification, counting, positional reasoning, and object identification, specifically testing how well models leverage depth information and spatial reasoning. Use when the user wants to benchmark on SSRBench, or asks about evaluating this task. Reports accuracy.
Evaluates musical reasoning capabilities across rhythm, chords, intervals, and scales using sheet music problems. It tests both textual and visual (staff notation) modalities to measure how well models recognize musical elements and perform logical deductions. Use when the user wants to benchmark on Synthetic Sheet Music Reasoning Benchmark (SSMR-Bench), or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict binding affinity between drug molecules and target proteins. It probes regression accuracy, correlation strength, and ranking consistency across varying data scarcity and generalization settings. Use when the user wants to benchmark on BindingDB, DAVIS, KIBA, or asks about evaluating this task. Reports Concordance Index (CI).
Evaluates the transfer learning capability of self-supervised learning (SSL) pre-trained models on diverse histopathology datasets. It probes domain-specific representation learning by measuring performance on image classification and nuclei instance segmentation tasks under linear probing and fine-tuning protocols. Use when the user wants to benchmark on BACH, CRC, PCam, MHIST, CoNSeP, or asks about evaluating this task. Reports top-1 accuracy.
Probes constrained-manifold spatial reasoning by requiring models to rank structural components based on geometric, topological, and physical constraints in complex 3D engineering scenes. It tests compositional spatial operations like mental rotation, occlusion handling, and force-path reasoning, revealing gaps in structural grounding and 3D constraint consistency. Use when the user wants to benchmark on SSI-Bench, or asks about evaluating this task. Reports Taskwise Accuracy.
Evaluates a model's ability to generate structured, human-centric scene graphs from images by jointly predicting verb predicates and fine-grained semantic role-value pairs for persons and objects. It probes multi-concurrent action understanding, affordance reasoning, and structured visual representation learning. Use when the user wants to benchmark on SSG dataset, Action Genome dataset, or asks about evaluating this task. Reports accuracy.
Evaluates real-time object detection capability by predicting bounding boxes and class scores directly from multi-scale feature maps, eliminating traditional proposal generation steps. It measures how well the model localizes and classifies objects across varying scales and aspect ratios under strict latency constraints. Use when the user wants to benchmark on PASCAL VOC2007, or asks about evaluating this task. Reports mAP.
Evaluates machine translation quality estimation metrics on under-resourced African languages by comparing their predicted scores against human-annotated Direct Assessment (DA) judgments. It probes a model's ability to correlate with human perception of translation adequacy across diverse language pairs, including both reference-based and reference-free settings. Use when the user wants to benchmark on SSA-MTE, or asks about evaluating this task. Reports Spearman correlation.
Evaluates the recommendation accuracy and unlearning effectiveness of session-based recommendation models after deleting a portion of training sessions. It measures how well the model retains predictive performance while successfully preventing the inference of removed items. Use when the user wants to benchmark on Amazon Beauty, Amazon Games, Steam, or asks about evaluating this task. Reports NDCG@K.
Evaluates symbolic regression methods on their ability to recover known physical laws from tabular data, testing both predictive accuracy and structural interpretability while probing robustness against irrelevant dummy variables. Use when the user wants to benchmark on SRSD-Feynman, or asks about evaluating this task. Reports R^2 > 0.999.