Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,553–6,576 of 20,836 skills
Probes the resource-latency trade-off between FPGA programmable logic (hls4ml) and AMD AI Engines for dense neural network layers. It quantifies the minimum PL hardware resources required to match AIE inference latency, identifying architectural crossover points for different layer shapes and reuse factors. Use when the user has predictions and gold and needs to compute LARE (Latency-Adjusted Resource Equivalence).
Probes the ability to extract aspect-based sentiment quadruples (target, aspect, opinion, sentiment) from text in low-resource agglutinative languages. It evaluates exact-match performance across entity detection, relation linking, and full quadruple composition. Use when the user wants to benchmark on LASQ, or asks about evaluating this task. Reports F1.
Evaluates deep learning models on video-based laparoscopic surgical training tasks. It probes the model's ability to recognize task-specific procedural errors and predict structured global skill ratings from synchronized stereo video streams. Use when the user wants to benchmark on LASANA, or asks about evaluating this task. Reports error_recognition.
This protocol evaluates the cross-lingual safety alignment of LLMs by measuring how frequently they comply with jailbreak prompts across multiple languages and resource levels. It simultaneously verifies that safety alignment does not degrade general capabilities such as multilingual knowledge, reasoning, and instruction following. Use when the user wants to benchmark on MultiJail, HarmBench (translated), M-MMLU, MT-Bench, MGSM, or asks about evaluating this task. Reports Attack Success Rate ...
Evaluates how well vision models and latent action representations capture semantic action categories and map visual features to low-level robotic control trajectories. It probes both high-level action understanding and physical grounding for generalizable vision-to-action alignment across diverse robotic and human motion datasets. Use when the user wants to benchmark on VLABench, CALVIN, RoboCOIN, AgiBotWorld-Beta, or asks about evaluating this task. Reports Top-1 Accuracy.
Evaluates an LLM's ability to perform legal argument reasoning by predicting the correct continuation of a court's argument chain. Given case facts and preceding arguments, the model must select the most plausible next argument from multiple options, testing its understanding of legal logic and precedent application. Use when the user wants to benchmark on LAR-ECHR, or asks about evaluating this task. Reports accuracy.
Evaluates large language models on their proficiency in Lao, a low-resource Southeast Asian language. It probes factual knowledge, K12 curriculum alignment, culturally grounded reasoning, bilingual translation fidelity, and open-ended generation quality through multiple-choice, translation, and pairwise arena tasks. Use when the user wants to benchmark on LaoBench, or asks about evaluating this task. Reports Accuracy.
Evaluates multilingual language transfer by measuring how well models adapt to German and Bulgarian while preserving source language (English) capabilities. Probes catastrophic forgetting and cross-lingual generalization across reasoning, math, reading comprehension, and commonsense tasks. Use when the user wants to benchmark on Multilingual Language Transfer Benchmarks (EN/DE/BG), or asks about evaluating this task. Reports normalized accuracy.
This metric probes an LLM's internal representation quality and cross-lingual alignment by measuring how closely the embedding space of a target language clusters around an English baseline. It quantifies multilingual capability and pre-training data imbalance by computing similarity scores across specific transformer layers. Use when the user has predictions and gold and needs to compute Language Ranker.
Evaluates language identification models on noisy web crawl data to measure their ability to accurately filter in-language sentences for low-resource languages. It probes domain mismatch and class imbalance effects on real-world LangID deployment. Use when the user wants to benchmark on Web Crawl & Held-out Eval Set, or asks about evaluating this task. Reports precision.
Evaluates models' ability to predict lane-level traffic speed and flow by modeling spatio-temporal dependencies on graph-structured lane networks. It tests performance across both regular and irregular lane configurations, emphasizing both predictive accuracy and training efficiency. Use when the user wants to benchmark on PeMS, PeMSF, HuaNan, or asks about evaluating this task. Reports MAE.
Evaluates the internal reasoning dynamics of LLMs on multi-choice tasks by tracking intermediate thought states and measuring convergence behavior. It quantifies consistency, uncertainty, and perplexity across different model scales, reasoning tasks, and decoding methods to visualize how reasoning trajectories evolve toward correct or incorrect answers. Use when the user wants to benchmark on AQuA, MMLU, StrategyQA, CommonSenseQA, or asks about evaluating this task. Reports reasoning accuracy.
Evaluates a transformer variant's ability to retrieve relevant past context blocks and maintain language modeling performance over extended sequence lengths. It probes long-range dependency retention and random-access memory retrieval capabilities compared to standard and recurrent transformer baselines. Use when the user wants to benchmark on PG-19, arXiv math papers, RedPajama (subset), or asks about evaluating this task. Reports perplexity.
Evaluates semantic segmentation models on high-resolution aerial imagery for mapping four land cover classes: buildings, woodlands, water, and roads. It probes the model's ability to accurately segment fine-grained, small, and narrow objects in rural environments from RGB imagery. Use when the user wants to benchmark on LandCover.ai, or asks about evaluating this task. Reports mIoU.
Evaluates land cover segmentation models on Sentinel-2 imagery using sparse annotations to predict fuel maps. It tests the model's ability to generalize across European regions affected by wildfires and compare against dense ground truth datasets (LUCAS, Urban Atlas). Use when the user wants to benchmark on Sentinel-2 Land Cover Dataset, or asks about evaluating this task. Reports F1 score.
This benchmark evaluates a model's ability to generate long-form, personalized question-answering responses by aligning outputs with fine-grained, user-specific information needs extracted from community Q&A narratives. It probes aspect-based response quality rather than binary correctness, measuring how well generated answers address individual criteria tailored to a specific user's profile. Use when the user wants to benchmark on LaMP-QA, or asks about evaluating this task. Reports aspect-b...
Evaluates open-vocabulary multi-object tracking by requiring models to follow multiple targets in video sequences guided by natural language descriptions. It probes the model's ability to jointly perform text-grounded detection and long-term identity association across diverse, real-world scenarios. Use when the user wants to benchmark on LaMOT, or asks about evaluating this task. Reports HOTA.
Evaluates the morphosyntactic well-formedness and grammaticality of generated natural language text. It measures how closely a generated sentence or corpus adheres to language-specific dependency rules extracted from treebanks. Use when the user has predictions and gold and needs to compute L'AMBRE.
Evaluates Chinese legal LLM capabilities across three hierarchical levels: basic legal NLP, basic legal application, and complex legal application. Probes tasks including named entity recognition, judicial summarization, case recognition, judgment prediction, legal question answering, and legal reasoning generation to measure domain-specific text processing, analysis, and reasoning skills. Use when the user wants to benchmark on LAiW Legal Evaluation Dataset (LED), or asks about evaluating th...
Evaluates the probabilistic forecasting skill and ensemble calibration of AI weather models by comparing them against a parameter-free lagged ensemble baseline. It probes whether models trained with long-lead-time objectives suffer from under-dispersion and poor variance calibration despite strong deterministic accuracy. Use when the user wants to benchmark on Atmospheric reanalysis / IFS HRES, or asks about evaluating this task. Reports CRPS.
Evaluates a language-assisted feature transformation framework for anomaly detection. It probes the model's ability to use textual prompts to define normality boundaries and selectively suppress or emphasize specific image attributes without retraining, across both semantic and industrial anomaly detection benchmarks. Use when the user wants to benchmark on Colored MNIST, Waterbirds, CelebA, MVTec AD, VisA, or asks about evaluating this task. Reports AUROC.
Evaluates open-vocabulary and closed-set object detection capabilities on remote sensing imagery. It probes a model's ability to detect novel Earth-based objects without prior training on them, as well as its efficiency when fine-tuned with limited labeled data. Use when the user wants to benchmark on LAE-1M, DIOR, DOTAv2.0, LAE-80C, or asks about evaluating this task. Reports mAP.
This benchmark evaluates the robustness of histopathology image classification models to both uniform and asymmetric label noise. It compares the performance of contrastive deep embeddings against non-contrastive backbones and image-based noise-robust loss functions. The protocol measures how well classifiers maintain accuracy when training labels are corrupted. Use when the user wants to benchmark on NCT-CRC-HE-100K, PatchCamelyon, BACH, MHIST, LC25000, GasHisSDB, or asks about evaluating th...
Evaluates GPT-3's ability to predict ground-truth labels for given instances, and analyzes whether explanation quality correlates with prediction correctness across different datasets. Use when the user wants to benchmark on CommonsenseQA, SNLI, or asks about evaluating this task. Reports accuracy.