Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,393–4,416 of 23,503 skills
Evaluates models' ability to generate calibrated probabilistic forecasts of supply chain disruptions from raw news text. It probes temporal generalization, uncertainty quantification, and the prioritization of high-risk signals for decision-making. Use when the user wants to benchmark on Supply Chain Disruption Forecasting Dataset, or asks about evaluating this task. Reports Brier score.
Evaluates the ability of a gradient-free, locally-updated spiking neural network to classify images using population-level spike agreement metrics. It probes whether replacing backpropagation with supervised Spike Agreement-Dependent Plasticity (SADP) and Cohen’s κ can achieve competitive vision and biomedical classification performance while maintaining biological plausibility and hardware compatibility. Use when the user wants to benchmark on MNIST, Fashion-MNIST, CIFAR-10, LC25000, Brain M...
Evaluates the ability of a predictor model to estimate the performance of instruction-following language models on unseen tasks, using only the task instruction as input. It probes the fundamental challenge of third-party model transparency and controllability at the task level. Use when the user wants to benchmark on SuperNI, or asks about evaluating this task. Reports RMSE.
Evaluates a language model's ability to follow diverse natural language instructions across various NLP tasks in a zero-shot setting. It measures how well the model generalizes to unseen tasks without in-context examples. Use when the user wants to benchmark on SUPER-NATURALINSTRUCTIONS, or asks about evaluating this task. Reports ROUGE-L.
Evaluates general-purpose language understanding across eight diverse tasks including coreference resolution, question answering, and natural language inference. It probes a model's ability to handle complex reasoning, sample-efficient learning, and transfer learning beyond standard GLUE capabilities. Use when the user wants to benchmark on SuperGLUE, or asks about evaluating this task. Reports SuperGLUE score.
Evaluates deep chemical reasoning capabilities of LLMs using expert-curated, entity-masked multiple-choice problems. It probes both final-answer accuracy and the fidelity of the reasoning process against expert-annotated solution paths, while also assessing the impact of multimodal inputs on complex chemical problem-solving. Use when the user wants to benchmark on SUPERChem-A11, SUPERChem-release, SUPERChem-holdout, SUPERChem-100, Multimodal-Essential Subset, or asks about evaluating this tas...
Evaluates the effectiveness of a proactive validation system for cloud AI infrastructure in detecting hardware defects, selecting optimal benchmark subsets, and balancing validation cost against system reliability. It probes the ability of automated criteria and selection algorithms to distinguish healthy nodes from degraded ones while minimizing downtime and maximizing GPU utilization. Use when the user wants to benchmark on Cluster Benchmark Dataset, or asks about evaluating this task. Repo...
Evaluates speech models on three downstream tasks: Intent Classification, Keyword Spotting, and Automatic Speech Recognition. It specifically probes noise robustness by comparing performance on clean speech versus speech corrupted with out-of-distribution CHiMe3 background noise. Use when the user wants to benchmark on SUPERB, or asks about evaluating this task. Reports Accuracy (Acc%).
Evaluates pre-trained speech models on downstream spoken language understanding tasks. It probes the model's ability to classify spoken intents, fill semantic slots in transcriptions, and detect specific keywords in audio. Use when the user wants to benchmark on SUPERB, or asks about evaluating this task. Reports test accuracy.
Evaluates a model's ability to perform spatial super-resolution on global weather forecast data, specifically upscaling temperature and cloud coverage maps from 1° to 0.5° resolution. It measures pixel-wise reconstruction accuracy against high-resolution ground truth. Use when the user wants to benchmark on GraphCast-ERA5 Paired Dataset, or asks about evaluating this task. Reports MSE.
Evaluates instruction-following and cross-task generalization capabilities of language models on a massive, diverse benchmark of 1,616 NLP tasks spanning 76 task types and 55 languages. It measures how well models trained on a mix of tasks perform on unseen tasks when given natural language instructions. Use when the user wants to benchmark on Super-NaturalInstructions, or asks about evaluating this task. Reports human evaluation metric.
Evaluates a decentralized multi-agent reinforcement learning algorithm that selectively shares high-temporal-difference-error experiences. It probes cooperative and competitive multi-agent coordination, credit assignment, and communication efficiency in anonymous environments with separate per-agent reward signals. Use when the user wants to benchmark on PettingZoo (Pursuit, Battle, Adversarial-Pursuit), or asks about evaluating this task. Reports total mean episode reward.
This benchmark evaluates how well automatic summarization metrics align with human judgments across multiple quality dimensions. It probes whether standard n-gram, embedding-based, and reference-less metrics reliably predict human-perceived coherence, consistency, fluency, and relevance of generated summaries. Use when the user wants to benchmark on SummEval, or asks about evaluating this task. Reports Kendall’s tau.
Evaluates the factual consistency of human reference summaries across popular abstractive summarization benchmark datasets. It probes whether widely used datasets contain systematic factual errors or low-abstraction artifacts that compromise their validity as training and evaluation standards. Use when the user wants to benchmark on CNN/DM, XSUM, XL-Sum (English), or asks about evaluating this task. Reports Factuality Score.
This evaluation probes a model's ability to compress long documents into concise summaries while preserving topical coverage, cross-sentence coherence, and factual consistency under strict token budgets. It measures how well extractive or generative methods balance semantic relevance with structural discourse cues across diverse domains. Use when the user wants to benchmark on CNN/DailyMail, GovReport, arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-2.
Evaluates the sum spectral efficiency of a hybrid centralized-distributed precoding scheme in fronthaul-constrained cell-free massive MIMO networks, comparing it against fully centralized and fully distributed baselines under varying fronthaul capacities and antenna configurations. Use when the user has predictions and gold and needs to compute sum spectral efficiency (sum SE).
Evaluates the robustness and generalization of 3D policy learning models for robotic manipulation across varying environmental conditions, temporal horizons, and real-world interference. It probes spatial understanding, fine-grained pose control, and resilience to domain randomization and lighting changes. Use when the user wants to benchmark on RoboTwin 2.0, ManiSkill2, Real-World Manipulation, or asks about evaluating this task. Reports Success Rate (%).
Evaluates the efficiency and compactness of subword tokenizers on Hindi, English, and code-mixed text by measuring how many tokens are generated per word and how frequently words are split into multiple tokens. Use when the user has predictions and gold and needs to compute Subword Fertility (SF).
Evaluates the faithfulness of image attribution methods by measuring how prediction confidence changes as important regions are removed or added. It also probes the ability to identify specific image regions that cause model misclassifications. Use when the user wants to benchmark on Celeb-A, VGG-Face2, CUB-200-2011, or asks about evaluating this task. Reports Deletion AUC.
Evaluates models' ability to classify six subjective linguistic features (Assertive, Cautious, Optimistic, Specific, Clear, Relevant) in financial earnings call question-and-answer transcripts. It probes how well models capture nuanced, tone-based, and domain-specific communication cues beyond factual content. Use when the user wants to benchmark on SubjECTive-QA, or asks about evaluating this task. Reports weighted F1 score.
This benchmark evaluates text anonymization methods by measuring both span-level masking accuracy and subject-level privacy leakage. It probes whether anonymized text successfully prevents adversarial LLMs from inferring personal identifiable information (PII) and sensitive attributes, while maintaining text utility. Use when the user wants to benchmark on PANORAMA, TAB, or asks about evaluating this task. Reports CPR.
Evaluates the quality of distributed subgraph representations learned by subgraph2vec on graph classification and clustering tasks. It probes the model's ability to capture structural and semantic similarities in chemical/biological graphs and Android application control-flow graphs. Use when the user wants to benchmark on MUTAG, PTC, PROTEINS, NCI1, NCI109, CLONE260, TRAIN10K, TEST10K, or asks about evaluating this task. Reports Accuracy.
Evaluates zero-shot segmentation accuracy on cellular and subcellular microscopy imagery, and assesses the downstream reliability of extracted morphological features for drug hit validation in high-content screening assays. Use when the user wants to benchmark on Cell segmentation datasets, Hit validation datasets, or asks about evaluating this task. Reports Dice Score (DSC), Z'-factor.
Evaluates a visual-to-audio dubbing model's ability to generate emotionally consistent, speaker-identical speech that aligns temporally with video lip movements. It probes multi-scale style learning across unseen speakers, reference audio variations, and phoneme-level lip-sync accuracy. Use when the user wants to benchmark on V2C-Animation, GRID, or asks about evaluating this task. Reports WER.