Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,009–4,032 of 23,477 skills
Evaluates large language models' ability to understand and generate structured scene graphs from textual narratives. It probes spatial reasoning, action decomposition, and the capacity to map dynamic descriptions to discrete visual or structural elements. Use when the user wants to benchmark on TSG Bench, or asks about evaluating this task. Reports Exact Match (EM) / Accuracy, Precision, Recall, Macro F1.
Evaluates how time series foundation models scale in forecasting accuracy and uncertainty calibration as model size, compute, and training data size increase. It probes both in-distribution generalization and out-of-distribution transfer capabilities across multiple standard time series forecasting benchmarks. Use when the user wants to benchmark on Monash subset, LSF subset, or asks about evaluating this task. Reports NLL.
Evaluates object detection models on traffic surveillance footage under diverse weather conditions and varying degrees of vehicle occlusion. It probes robustness to environmental degradation, scale variation, and dense urban traffic scenarios. Use when the user wants to benchmark on TSBOW, or asks about evaluating this task. Reports mAP50.
Evaluates large language models' ability to perform time series analysis and reasoning across six tasks (anomaly detection, classification, characterization, comparison, data transformation, and temporal relationship) using three question formats (true-or-false, multiple-choice, and puzzling). Use when the user wants to benchmark on TSAQA, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of large multimodal models to generate accurate, domain-agnostic natural language descriptions of time series trends. It probes cross-modal alignment between visual time series plots (or extracted features) and textual trend explanations. Use when the user wants to benchmark on TS-Insights, or asks about evaluating this task. Reports final_score.
Evaluates the forecasting accuracy and computational efficiency of deep learning models on multivariate time series data. It probes how architectural choices, preprocessing steps, and spatial-temporal processing configurations impact performance across varying forecasting horizons. Use when the user wants to benchmark on Weather, Solar-Energy, ECL, Traffic, or asks about evaluating this task. Reports MAE.
Evaluates the factual accuracy and truthfulness of large language models by measuring their ability to select correct answers over common misconceptions. It probes the model's capacity to resist generating plausible but false statements across diverse categories like health, law, and politics. The benchmark specifically tests whether models can identify and output factually correct responses when presented with multiple candidate answers. Use when the user wants to benchmark on TruthfulQA, or...
Evaluates an LLM's factual accuracy and hallucination mitigation across multiple-choice, short-form, and long-form generation tasks. It measures the trade-off between truthfulness and informativeness, and quantifies the exact number of supported versus unsupported facts in generated text. Use when the user wants to benchmark on TruthfulQA, BioGEN, or asks about evaluating this task. Reports Accuracy, True*Info, FActScore.
This benchmark evaluates a model's ability to perform true multimodal in-context learning by requiring it to solve tasks that depend on both visual and textual information from provided demonstrations. It probes whether models can correctly attend to and utilize visual context in few-shot examples rather than relying on superficial textual patterns or prior knowledge. Use when the user wants to benchmark on TrueMICL, or asks about evaluating this task. Reports accuracy.
Binary classification of truck driving risk on specific highway segments based on historical behavior, short-term trip dynamics, and real-time traffic conditions. It probes a model's ability to predict forward collision warning events using a small, highly imbalanced dataset of real-world trajectory data. Use when the user wants to benchmark on Truck Driving Risk Dataset, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).
Evaluates a model's ability to perform video summarization by predicting frame-level importance scores across visual, textual, and audio modalities. It probes the model's capacity for adaptive multimodal fusion and temporal dependency modeling to identify salient segments in long videos. Use when the user wants to benchmark on MoSu, Mr. HiSum, SumMe, TVSum, or asks about evaluating this task. Reports Kendall’s τ (kTau), Spearman’s ρ (sRho).
Evaluates large language models' ability to distinguish factually true statements from factually false and unverifiable ('neither') statements. It probes both prompt-based output probabilities and internal hidden activations to measure veracity classification accuracy and uncertainty quantification. Use when the user wants to benchmark on Trilemma of Truth Datasets, or asks about evaluating this task. Reports MCC.
This benchmark evaluates Turkish natural language understanding across multiple task types, including grammaticality judgment, sentiment analysis, paraphrase detection, and semantic textual similarity. It probes a model's ability to handle agglutinative morphology, culturally adapted expressions, and nuanced semantic equivalence in Turkish. Use when the user wants to benchmark on TrCoLA, TrSST-2, TrMRPC, TrSTS-B, or asks about evaluating this task. Reports accuracy.
Evaluates fact-checking systems across three sub-tasks: retrieving relevant evidence, verifying claim truthfulness, and generating natural language explanations. It specifically probes a model's Hotspot Perception Ability (HPA) by measuring how effectively it allocates reasoning effort and computational resources based on the real-world influence of trending claims. Use when the user wants to benchmark on TrendFact, or asks about evaluating this task. Reports F1-macro.
Evaluates single-object visual tracking performance in first-person vision videos, specifically testing robustness to object manipulation, occlusions, and dynamic interactions under real-time execution constraints. Use when the user wants to benchmark on TREK-150, or asks about evaluating this task. Reports SS, NPS, GSR.
Evaluates a model's ability to perform fine-grained tree species identification using multimodal Earth observation data, specifically leveraging temporal dynamics from optical and radar time series alongside high-resolution imagery. Use when the user wants to benchmark on TreeSatAI-TS, or asks about evaluating this task. Reports weighted F1.
This benchmark evaluates an LLM's ability to perform deep, structured scientific peer review by generating comprehensive reviews and actionable feedback comments. It probes the model's capacity for hierarchical question decomposition, context-aware analysis of long documents, and alignment with human reviewer judgments across multiple quality dimensions. Use when the user wants to benchmark on TreeReview Benchmark (ICLR-2024, NeurIPS-2023, Nature Communications), or asks about evaluating this...
TreeEval probes an LLM's ability to handle complex, adaptive reasoning through dynamically generated hierarchical questions. It evaluates how well a model's relative performance ranking aligns with established leaderboards like AlpacaEval2.0, while testing the framework's efficiency in distinguishing fine-grained capability differences without relying on static datasets. Use when the user wants to benchmark on TreeEval (Dynamic/Benchmark-Free), or asks about evaluating this task. Reports Spea...
Evaluates visual grounded reasoning by requiring models to precisely localize target objects in cluttered scenes and perform second-order reasoning about their interactions. It measures both the correctness of the final answer and the spatial accuracy of the traceable bounding box evidence. Use when the user wants to benchmark on TreeBench, or asks about evaluating this task. Reports Accuracy.
This evaluation protocol compares recursive tree-based neural models against recurrent sequence-based models across multiple NLP tasks. It probes whether syntactic tree structures are necessary for learning representations, particularly for tasks requiring long-distance dependency modeling or hierarchical composition. Use when the user wants to benchmark on Stanford Sentiment Treebank, Pang Sentiment Dataset, UMD-QA, SemEval-2010 Task 8, Discourse Parsing, or asks about evaluating this task. ...
Probes retrieval-augmented generation systems on complex, narrative-driven queries by decomposing information needs into sub-narratives. It evaluates document relevance based on sub-narrative coverage, measures response quality via strict vital recall of fully supported information nuggets, and assesses sentence-level factual grounding against cited documents. Use when the user wants to benchmark on MS MARCO V2.1, or asks about evaluating this task. Reports strict_vital_recall.
Evaluates the factual accuracy and content grounding of RAG-generated answers by checking for the presence of key factual claims (nuggets) extracted from source documents. It also measures answer length to assess the trade-off between conciseness and completeness in system outputs. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports V_strict.
Evaluates neural cross-language information retrieval systems on ad hoc, reranking, and monolingual tasks across Chinese, Persian, and Russian newswire collections. It measures how well models retrieve and rank relevant documents when queries are in English and documents are in other languages, or when queries are human-translated. Use when the user wants to benchmark on TREC 2022 NeuCLIR Collections, or asks about evaluating this task. Reports nDCG@10.