Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,081–4,104 of 23,477 skills
Probes the capability of real-time tumor and surrogate localization in MRI-guided radiotherapy using 2D sagittal cine MRI sequences. It evaluates how well algorithms can track anatomical motion across varying frame rates and multi-vendor MRI-linac hardware under clinically relevant conditions. Use when the user wants to benchmark on TrackRAD2025, or asks about evaluating this task. Reports tracking performance.
Evaluates the ability of deep learning models to detect and track high-speed, tiny objects (tennis and badminton balls) in broadcast sports videos. It probes robustness to motion blur, occlusion, and domain shifts by comparing single-frame vs. multi-frame tracking and transfer learning across different sports. Use when the user wants to benchmark on Tennis, Badminton, or asks about evaluating this task. Reports F1-measure.
Probes a GNN's capability to perform edge scoring on highly sparse, irregular scientific graphs by predicting the probability that a directional connection between two 3D space-point measurements originates from the same particle. Use when the user wants to benchmark on TrackML, or asks about evaluating this task. Reports edge probability.
Evaluates a model's ability to track objects through appearance-changing state transformations and explicitly model those transformations as a state graph. It probes spatiotemporal continuity, zero-shot object recovery, and semantic reasoning about object interactions. Use when the user wants to benchmark on VOST, VSCOS, M3-VOS, DAVIS 2017, VOST-TAS, or asks about evaluating this task. Reports Jaccard (J).
This benchmark evaluates an LLM's ability to detect and classify reward hacking behaviors in multi-turn code generation trajectories. It specifically probes contrastive anomaly detection capabilities by presenting clusters of mixed benign and malicious trajectories, testing whether models can disentangle subtle semantic and syntactic exploit patterns without prior taxonomy exposure. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports Detection Rate.
This benchmark evaluates a model's ability to predict post-click Gross Merchandise Volume (GMV) under delayed feedback conditions. It specifically probes how well models adapt to rapidly evolving label distributions through online streaming training and whether they can effectively handle the distinct statistical properties of single-purchase versus repurchase transactions. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports AUC.
This protocol evaluates training-free partial audio deepfake detection by analyzing the temporal continuity of frozen speech foundation model embeddings. It probes a model's ability to detect splice boundaries and synthetic insertions in speech without requiring labeled training data or architectural modifications. Use when the user wants to benchmark on PartialSpoof, HalfTruth Audio Deepfake (HAD), ADD 2023 Track 2, LlamaPartialSpoof, or asks about evaluating this task. Reports EER.
Evaluates the quality and efficiency of trace encoding methods for process mining event logs. It probes how well encodings preserve trace similarities (expressivity), their computational cost as data scales (scalability), and their suitability for downstream process mining tasks. Use when the user wants to benchmark on Process Mining Event Log Scenarios (1-5), or asks about evaluating this task. Reports T4.
Evaluates a model's ability to classify OpenTelemetry workflow traces as benign, suspicious, or malicious, and assesses its knowledge of cybersecurity frameworks via multiple-choice questions. Use when the user wants to benchmark on OpenTelemetry Workflow Traces, or asks about evaluating this task. Reports Overall Accuracy.
Evaluates large language models' ability to perform multi-table question answering across varying context lengths (8K–64K tokens) and complex reasoning tasks. It probes cross-table inference, symbolic reasoning, and handling of real-world relational data without Wikipedia bias. Use when the user wants to benchmark on TQA-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the performance and energy efficiency of a Tensor Processing Unit (TPU) and alternative hardware designs across six specific neural network workloads. It probes how architectural parameters like memory bandwidth, clock rate, and matrix multiply unit size impact throughput and power consumption. Use when the user wants to benchmark on TPU Benchmark Workloads (MLP0, MLP1, LSTM0, LSTM1, CNN0, CNN1), or asks about evaluating this task. Reports Watt/die.
Evaluates large language models' ability to perform analytical calculations in hypersonic thermal protection system engineering using closed-form formulas and thermodynamic relations, without relying on external simulation tools. Use when the user wants to benchmark on TPS-CalcBench, or asks about evaluating this task. Reports relative_error.
Evaluates the discriminative capability of a speaker verification model by measuring the true positive rate at fixed false positive rate thresholds. It probes how well the model's embedding space separates same-speaker pairs from different-speaker pairs under controlled error constraints. Use when the user has predictions and gold and needs to compute TPR@FPR.
Evaluates an LLM's ability to generate structurally complex SQL queries for real-world decision-making workloads. It probes the model's capacity to handle deep nesting, multiple joins, diverse column references, and complex filtering conditions compared to simpler benchmarks. Use when the user wants to benchmark on TPC-DS, or asks about evaluating this task. Reports structural_similarity.
Evaluates unsupervised anomalous sound detection systems on miniature machine operating sounds. It probes the ability of models to learn normal acoustic patterns and identify deviations caused by mechanical faults or environmental variations. Use when the user wants to benchmark on ToyADMOS, or asks about evaluating this task. Reports AUC-ROC.
Evaluates whether Multimodal Large Language Models (MLLMs) can generate structurally valid, low-toxicity alternative molecules from toxic inputs while adhering to drug-likeness, synthetic feasibility, and structural similarity constraints. It probes the model's ability to perform structure-aware molecular editing and cross-modal scientific reasoning. Use when the user wants to benchmark on ToxiMol, or asks about evaluating this task. Reports Toxicity Repair Success Rate.
Evaluates how well automated toxicity classifiers align with diverse human perceptions of harmful content, specifically measuring how demographic background and personal harassment experiences influence toxicity judgments. Use when the user wants to benchmark on Toxicity Perspectives Dataset, or asks about evaluating this task. Reports interrater agreement (Cohen's kappa).
Evaluates the toxicity of text sequences (prompts and model continuations) by scoring them with a black-box API. It probes how well models generate non-toxic text and how sensitive toxicity metrics are to API updates and score drift over time. Use when the user wants to benchmark on REALTOXICITYPROMPTS, or asks about evaluating this task. Reports Toxic Fraction.
Probes the ability of text generation models to produce non-toxic content by measuring average toxicity scores and comparing them via statistical significance testing. It specifically evaluates how accounting for classifier uncertainty affects the reliability of these comparisons. Use when the user wants to benchmark on BOLD, RealToxicityPrompts, or asks about evaluating this task. Reports Confidence Interval.
Evaluates binary toxic language classification under extreme data scarcity and severe class imbalance. It probes how well classifiers can detect the minority 'threat' class when trained on a very small labeled dataset, and measures the effectiveness of various data augmentation techniques in improving recall and macro-F1. Use when the user wants to benchmark on Seed, or asks about evaluating this task. Reports macro-averaged F1-score.
This benchmark evaluates the individual and group fairness of toxicity classifiers on online comments. It probes whether model predictions remain stable when sensitive identity tokens are swapped (individual fairness) and whether prediction accuracy is equitable across different demographic groups (group fairness). Use when the user wants to benchmark on Toxic Comment Classification Challenge, or asks about evaluating this task. Reports Balanced Accuracy (BA).
Evaluates transformer and RNN models on their ability to classify toxic comments while measuring classification accuracy and inference speed. It specifically probes identity-based bias by measuring how well models distinguish between toxic and normal comments across demographic subgroups. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Macro AUROC.
Evaluates graph neural networks for few-shot toxic molecule classification. It probes the model's ability to generalize from very limited labeled examples (shots) and adapt to new query sets using meta-learning and graph augmentation techniques. Use when the user wants to benchmark on Tox21, or asks about evaluating this task. Reports ROC-AUC Score.
Evaluates molecular toxicity prediction capabilities across diverse AI architectures (descriptor-based models, neural networks, tabular transformers, and zero-shot LLMs) on a standardized chemical safety benchmark. Use when the user wants to benchmark on Tox21 Challenge dataset, or asks about evaluating this task. Reports performance.