Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,649–3,672 of 22,955 skills
Evaluates a conversational recommender system's ability to naturally transition topics, recommend relevant items, and generate coherent responses within a dialogue. It probes the model's capacity to leverage historical interactions, user profiles, and topic sequences to maintain semantic flow and recommendation accuracy. Use when the user wants to benchmark on TG-ReDial, or asks about evaluating this task. Reports NDCG@k.
Evaluates the causal reasoning and forecasting accuracy of LLMs and time-series models on multi-domain time-series data. It probes whether step-by-step reasoning and external event context improve numerical predictions or introduce narrative bias, particularly in stochastic versus pattern-rich domains. Use when the user wants to benchmark on TFRBench, or asks about evaluating this task. Reports MASE.
Evaluates a model's ability to understand visual metaphors and image implications by verifying multiple factual and inferential propositions per image. It probes fine-grained visual perception, multi-hop reasoning, and theory of mind through structured true-false questioning. Use when the user wants to benchmark on TFQ-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify short DNA sequences as transcription factor binding sites or not, capturing its capacity to learn regulatory sequence patterns from genomic data. Use when the user wants to benchmark on TFBS classification, or asks about evaluating this task. Reports AUC.
Evaluates cross-modal video retrieval models that must jointly process visual context and scene text (OCR tokens) to match sentence queries with relevant videos. Probes the model's ability to read, comprehend, and align fine-grained text semantics with visual frames in real-world scenarios. Use when the user wants to benchmark on TextVR, or asks about evaluating this task. Reports R@K (Recall@K).
Evaluates a model's ability to perform texture-aware segmentation by measuring how well it segments regions based on repeating texture patterns rather than semantic shape cues. It tests generalization on both synthetic texture-only images and natural images, while also checking for catastrophic forgetting on standard semantic benchmarks. Use when the user wants to benchmark on RWTD, STMD, ADE20K, or asks about evaluating this task. Reports mIoU.
Evaluates foundation models on tabular prediction tasks that require leveraging mixed categorical, numerical, and free-text features across diverse real-world domains. The benchmark tests whether models can maintain predictive performance when text features contain semantic ambiguity, synonym variation, or noise, while preserving structural tabular signals. Use when the user wants to benchmark on fraud, kick, osha, cards, complaints, spotify, airbnb, beer, houses, laptops, mercari, permits, w...
Evaluates text-to-image generation models on their ability to render long, dense, and structurally complex text accurately within images. It probes both semantic alignment between the prompt and the generated image, and precise character/word-level OCR fidelity across diverse layouts, styles, and real-world scenes. Use when the user wants to benchmark on TextAtlasEval, or asks about evaluating this task. Reports OCR Accuracy (Acc.).
This evaluation protocol assesses how well text-to-video generation models align generated content with textual prompts across fine-grained attributes like object counts, colors, actions, and spatial relationships. It measures both semantic alignment and visual/motion quality to determine if refinement techniques successfully correct misalignments without degrading fidelity. Use when the user wants to benchmark on EvalCrafter, T2V-CompBench, or asks about evaluating this task. Reports Text-Vi...
Evaluates a model's ability to translate natural language questions into correct SQL queries for a complex, real-world industrial database. It also probes schema-linking precision by measuring how accurately the model identifies the required tables from the schema. Use when the user wants to benchmark on Industrial Energy Database Benchmark, or asks about evaluating this task. Reports Accuracy.
Evaluates the reliability of text-to-SQL benchmarks by quantifying annotation error rates and measuring how these errors distort agent execution accuracy and leaderboard rankings. Use when the user wants to benchmark on BIRD, Spider 2.0-Snow, or asks about evaluating this task. Reports annotation error rate.
Evaluates the vulnerability of differential privacy-based text sanitization methods by measuring how accurately an attacker can reconstruct original sensitive or personally identifiable information (PII) tokens from their sanitized counterparts. It probes the effectiveness of Bayesian inference-based reconstruction attacks against state-of-the-art sanitization defenses. Use when the user wants to benchmark on SST-2, AGNEWS, QNLI, Yelp, or asks about evaluating this task. Reports ASR.
Evaluates a model's ability to generate images with accurate, legible, and layout-controlled text based on text prompts or masked regions. It probes text coherence, character-level rendering fidelity, and alignment between generated text and background imagery. Use when the user wants to benchmark on MARIO-10M, DrawBenchText, or asks about evaluating this task. Reports OCR(F-measure).
This benchmark evaluates a model's ability to perform text-queried audio source separation, specifically its capacity to isolate target sound events from mixed audio based on natural language instructions. It probes both acoustic fidelity (spectral and signal-level accuracy) and semantic alignment (how well the separated audio matches the textual description). Use when the user wants to benchmark on AudioCaps, Clotho v2, FSD50K, 3 Sets, MUSIC, or asks about evaluating this task. Reports LSD.
Evaluates the robustness of finetuned transformer models (BERT, GPT-2, T5) to various text perturbations (e.g., dropping nouns/verbs, character changes, adding text) across classification and generation tasks. It measures how much model performance degrades when inputs are syntactically or semantically altered. Use when the user wants to benchmark on GLUE, XSum, CommonGen, SQuAD, or asks about evaluating this task. Reports Accuracy, Robustness Score.
Evaluates the model's ability to align facial images with their textual descriptions by retrieving the correct image given a text query, and vice versa. It measures how well the model learns cross-modal semantic correspondence for face-centric data. Use when the user wants to benchmark on CelebA-Caption, MM-CelebA, or asks about evaluating this task. Reports R@5, R@10.
Evaluates the ability of centroid-based clustering algorithms to group unlabeled text documents into semantically coherent clusters. It measures clustering accuracy, label alignment with ground truth, and how closely learned centroids match true cluster centers. Use when the user wants to benchmark on Bank77, CLINC, GoEmo, MASSIVE, StackExchange, or asks about evaluating this task. Reports ACC, NMI.
Evaluates text classification performance across multiple sentiment, subjectivity, question classification, and topic categorization tasks. It probes the model's ability to capture contextual and syntactic features from sequential text using 2D matrix representations and spatial pooling. Use when the user wants to benchmark on MR, SST-1, SST-2, Subj, TREC, 20Newsgroups, or asks about evaluating this task. Reports accuracy.
Systematic comparison of generative (AR, MLM, Diffusion) and discriminative (encoder) transformer models on text classification tasks, focusing on sample efficiency, robustness to input noise, and output calibration/ordinality. Use when the user wants to benchmark on AG News, Emotion, SST2, SST5, Multiclass Sentiment Analysis, Twitter Financial News Sentiment, IMDb, Hate Speech Offensive, or asks about evaluating this task. Reports weighted-F1 score.
Evaluates a text-to-video model's ability to accurately render and animate specific text within a video scene while maintaining visual quality and temporal consistency. It probes character-level text fidelity, resistance to text collapse during motion, and overall video generation quality. Use when the user wants to benchmark on LAION subset, or asks about evaluating this task. Reports Sen. Acc.
Evaluates the impact of test-time scaling (TTS) inference strategies on Vision-Language Models across multimodal reasoning and perception tasks. It measures how techniques like Chain-of-Thought, Best-of-N, Self-Consistency, and Self-Refinement improve or degrade performance on open-source versus closed-source models. Use when the user wants to benchmark on MathVista, MMMU, MMBench, or asks about evaluating this task. Reports accuracy.
Evaluates whether a zero-shot prompting method (OOC) improves stratified invariance and counterfactual invariance in LLM text classification predictions across real-world and synthetic datasets, while measuring retention of predictive accuracy. Use when the user wants to benchmark on civilcomments (Toxic Comments), Bios (Occupation), Amazon Fashion Reviews, Discrimination (Synthetic), MIMIC-III/SBDH (Clinical), Semantic Leakage Tasks, or asks about evaluating this task. Reports SI-bias.
Evaluates the convergence speed and final test performance of distributed synchronous versus asynchronous stochastic gradient descent algorithms. It probes whether backup workers in synchronous training can mitigate stragglers without degrading accuracy due to gradient staleness. Use when the user has predictions and gold and needs to compute test accuracy.
Evaluates the robustness of Android malware classifiers against spatio-temporal experimental bias. It measures how model performance degrades when trained on past application data and tested on future data, while accounting for realistic malware-to-goodware class distributions. Use when the user wants to benchmark on Android malware dataset (2014-2016), or asks about evaluating this task. Reports F1-Score.