Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,985–4,008 of 23,477 skills
Evaluates named entity recognition (NER) and syntactic NLP capabilities on noisy, informal social media text. It probes a model's ability to handle domain-specific challenges like abbreviations, irregular capitalization, and complex entity structures common in tweets, while establishing baselines for tokenization, lemmatization, POS tagging, and dependency parsing. Use when the user wants to benchmark on Tweebank-NER (TB2), or asks about evaluating this task. Reports entity-level F1.
This evaluation probes a model's ability to correctly route natural language questions to the most appropriate domain-specific QA agent from a large, heterogeneous pool. It measures both sample efficiency (performance with few training examples per agent) and scalability (maintaining accuracy as the number of candidate agents grows to hundreds). Use when the user wants to benchmark on QA-Tasks, Many-Agents, or asks about evaluating this task. Reports Accuracy@1.
Evaluates a model's ability to predict explicit and implicit user negative feedback for video recommendations, including binary judgment of controversial content and classification of dislike reasons. It also tests the model's capability to simulate user viewing behavior (e.g., fast-skip) based on historical interactions and profile data. Use when the user wants to benchmark on TVNF, MovieLens, Steam, or asks about evaluating this task. Reports Recall.
Evaluates the end-to-end performance and optimization capability of a deep learning compiler across diverse hardware back-ends (GPU, CPU, embedded GPU, FPGA) on standard inference workloads. It measures how effectively the compiler automatically generates high-performance kernels compared to hand-tuned vendor libraries and existing frameworks. Use when the user wants to benchmark on DL Inference Workloads (ResNet-18, MobileNet, LSTM, DQN, DCGAN), or asks about evaluating this task. Reports sp...
Evaluates a model's ability to detect lane boundaries in highway driving scenarios using a point-based accuracy metric over predefined row anchors. Use when the user wants to benchmark on TuSimple, or asks about evaluating this task. Reports accuracy.
Evaluates the detection of sentence boundaries in Turkish text across diverse domains (scientific abstracts, news, social media). It tests robustness to formatting variations and punctuation absence. Use when the user wants to benchmark on trseg-41, or asks about evaluating this task. Reports F1-score.
Tests the ability to identify and classify named entities (Person, Location, Organization) in Turkish text. It probes fine-grained token-level classification and boundary detection. Use when the user wants to benchmark on Milliyet-Ner, WikiANN (Turkish subset), or asks about evaluating this task. Reports CoNLL F-1.
Measures the quality of bidirectional translation between Turkish and English. It evaluates how well models capture cross-lingual semantic alignment and syntactic restructuring across different corpus types. Use when the user wants to benchmark on Wmt-16 (Turkish-English subset), MuST-C (Turkish-English subset), or asks about evaluating this task. Reports BLEU Score.
Evaluates a model's ability to capture statistical patterns and long-range dependencies in Turkish text. It tests generative probability estimation at both subword and character levels across news and Wikipedia domains. Use when the user wants to benchmark on trwiki-67, trnews-64, or asks about evaluating this task. Reports Perplexity (Ppl), Bits-per-character (Bpc).
Evaluates automatic speech recognition (ASR) models on Tunisian Arabic dialect audio, measuring overall transcription accuracy and code-switching performance for embedded English and French phrases. Use when the user wants to benchmark on LinTO, TunSwitch, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates the robustness and zero-shot adaptation of vision-language segmentation models under prompt tuning across diverse medical and natural domain datasets. It probes how different prompt tuning strategies and prompt depths handle domain shifts and varying class counts. Use when the user wants to benchmark on Kvasir-SEG, ClinicDB, BKAI, ISIC 2016, DFU 2022, CAMUS, BUSI, CheXlocalize, Cityscapes, PascalVOC, or asks about evaluating this task. Reports Dice score.
Evaluates fine-grained temporal understanding on dense dynamic videos, probing camera motion, scene transitions, action sequences, and multi-subject interactions. It measures how well models capture dynamic visual elements and handle varying video complexities. Use when the user wants to benchmark on TUNA, or asks about evaluating this task. Reports F1 score, Accuracy.
This evaluation protocol probes the language understanding, reasoning, and instruction-following capabilities of Portuguese LLMs across diverse domains including academic exams, natural language inference, physical commonsense, and code generation. It is specifically designed to provide reliable training signals during pretraining and assess post-training alignment. Use when the user wants to benchmark on ARC Challenge, Calame, Global PIQA, HellaSwag, LAMBADA, ENEM, BLUEX, OAB, Belebele, MMLU...
Evaluates the ability of LLMs to improve reasoning performance at test time through self-reflection and targeted variant question synthesis, without external supervision. It probes how well a model can adapt its policy to difficult mathematical and general reasoning problems by diagnosing its own failures and generating corrective training signals. Use when the user wants to benchmark on AMC23, MATH-500, Minerva, OlympiadBench, AIME 2024, AIME 2025, GPQA-Diamond, MMLU-Pro, or asks about evalu...
Evaluates a text-to-speech model's ability to generate audio that matches natural language descriptions of speaker attributes (gender, accent, pitch, speaking rate, recording quality) and overall audio fidelity. It measures both objective acoustic metrics and subjective human ratings of relevance and naturalness. Use when the user wants to benchmark on MLS, LibriTTS-R, or asks about evaluating this task. Reports MOS.
Evaluates the impact of probabilistic versus deterministic duration modeling on the naturalness and intelligibility of non-autoregressive text-to-speech systems. It specifically probes how well stochastic duration predictors handle prosodic variability and disfluencies in spontaneous speech compared to read-aloud speech. Use when the user wants to benchmark on LJ, RS, TSGD2, AptS, or asks about evaluating this task. Reports CMOS.
This protocol evaluates a 3D convolutional auto-encoder for removing reverberation artifacts (clutter) from transthoracic echocardiographic (TTE) sequences. It measures how well the network preserves cardiac structures while suppressing simulated artifacts, using synthetic data with known ground truth for training and validation, and normal in-vivo sequences for testing. Use when the user wants to benchmark on Synthetic TTE sequences, or asks about evaluating this task. Reports reconstruction...
Evaluates the ability of models to detect human body forgeries generated by diffusion models. It probes spatiotemporal motion inconsistencies and generalization across different generation configurations and unseen manipulation models. Use when the user wants to benchmark on TT-DF, or asks about evaluating this task. Reports AUC.
Evaluates lesion detection performance on synthetic and real breast mammography/tomosynthesis images. Probes the model's ability to localize lesions across varying breast densities, lesion sizes, and lesion densities using a free-response receiver operating characteristic (FROC) framework. Use when the user wants to benchmark on T-SYNTH, EMBED, or asks about evaluating this task. Reports FROC (Sensitivity vs. Average False Positives per Image).
This benchmark evaluates an AI system's ability to perform fact verification using time-series evidence. It probes multi-timeframe temporal reasoning, cross-series numerical analysis, and the generation of factually consistent justifications aligned with human annotations. Use when the user wants to benchmark on TSVer, or asks about evaluating this task. Reports Accuracy.
Evaluates models on Time Series Extrinsic Regression (TSER), where the goal is to predict a single continuous scalar value from multivariate time series inputs of varying lengths and dimensions. It probes the model's ability to handle irregular time series, missing values, and diverse domain-specific patterns without imputation. Use when the user wants to benchmark on Monash TSER Archive, or asks about evaluating this task. Reports R2.
Evaluates a model's ability to perform sequential recommendation with a focus on capturing repeat-aware temporal patterns. It measures how well the model balances predicting new items versus recurring items based on user interaction history and time intervals. Use when the user wants to benchmark on RetailRocket, LastFM, Diginetica, or asks about evaluating this task. Reports HR@K.
Evaluates large language models on time series question answering across five tasks: forecasting, imputation, anomaly detection, classification, and open-ended reasoning. It probes the model's ability to integrate textual context with numerical time series data for both precise numerical prediction and natural language explanation. Use when the user wants to benchmark on TSQA, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of discovered motif sets in time series by measuring alignment with ground truth segments. It accounts for variable-length patterns and time warping while penalizing both false discoveries and missed patterns. Use when the user wants to benchmark on TSMD benchmark datasets, or asks about evaluating this task. Reports F1-score.