Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,857–8,880 of 21,231 skills
Evaluates few-shot learning capabilities of pre-trained language models across sentence classification, question answering, and named entity recognition tasks. It measures how well models adapt with limited labeled examples (10, 20, 30 shots) compared to fully supervised settings and human performance. Use when the user wants to benchmark on SST-2, MNLI, NER, MRC, or asks about evaluating this task. Reports macro-averaged results.
Evaluates Chinese language understanding across nine diverse tasks, including text classification, natural language inference, semantic similarity, and machine reading comprehension. It probes a model's ability to handle Chinese-specific linguistic phenomena, whole-word masking, and token-level vs. global understanding through a standardized fine-tuning pipeline. Use when the user wants to benchmark on CLUE, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to perform generative error correction (GER) for automatic speech recognition by reformulating the task as a cloze test. The model must select the correct hypothesis from a 5-best N-best list to minimize word error rate while maintaining source speech fidelity. Use when the user wants to benchmark on HyPoradise (GER benchmark), or asks about evaluating this task. Reports WER (%).
Evaluates continual learning capabilities on Visual Question Answering (CLVQA) by measuring how well a model retains knowledge from previous tasks while learning new ones across scene-incremental and function-incremental settings. Use when the user wants to benchmark on CLOVE, or asks about evaluating this task. Reports average accuracy (%).
Evaluates time series forecasting models, particularly pre-trained Transformers, on cloud operations data. It probes zero-shot generalization, architectural efficiency, and scaling behavior against classical and deep learning baselines. Use when the user wants to benchmark on azure2017, borg2011, ali2018, or asks about evaluating this task. Reports sMAPE.
Evaluates the ability of systems to detect context-aware anomalies in cloud environments by jointly analyzing system logs and performance metrics, and to classify the specific anomaly scenario. It also tests generalization to point-level anomalies using log-only or metric-only data. Use when the user wants to benchmark on CloudAnoBench, BGL, Thunderbird, HDFS_v1, or asks about evaluating this task. Reports F1-score.
This benchmark evaluates how effectively different Virtual Machines (VMs) can execute specific application workloads by ranking them according to weighted hardware attributes. It probes the capability to map domain-specific application requirements to underlying infrastructure performance characteristics. Use when the user wants to benchmark on Cloud VM Benchmarking Suite, or asks about evaluating this task. Reports S_i.
Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture. Use when the user wants to benchmark on Clotho v2, or asks about evaluating this task. Reports SDRi.
Evaluates the robustness of image classification models when trained on datasets with inherent label noise and class imbalance. It probes how effectively learning algorithms can filter mislabeled samples and adapt to skewed class distributions without manual curation. Use when the user wants to benchmark on Clothing1mPP, or asks about evaluating this task. Reports Noise Rate.
Evaluates cross-modal retrieval and zero-shot classification capabilities for remote sensing imagery (SAR and multispectral optical) paired with text descriptions. It probes how well unified semantic embeddings align heterogeneous geospatial data with natural language for crisis event and land cover analysis. Use when the user wants to benchmark on CrisisLandMark, or asks about evaluating this task. Reports nDCG@1000.
Evaluates the ability of voice cloning models to preserve speaker identity and acoustic characteristics across different speech conditions, including neutral and emotional speech. It measures how closely generated audio matches the reference speaker's embedding and signal properties without human intervention. Use when the user wants to benchmark on LS test-clean, TESS, or asks about evaluating this task. Reports cosine similarity (WavLM).
Evaluates a universal adversarial perturbation framework designed to defend against zero-shot voice cloning TTS models. It measures how well the perturbation degrades cloned audio quality and speaker similarity while preserving the perceptual fidelity of the original protected speech. Use when the user wants to benchmark on VCTK, LibriSpeech ASR, LibriTTS-R, LJSpeech, Common Voice, or asks about evaluating this task. Reports DSR.
Evaluates the ability of audio classifiers to distinguish between real human speech and AI-generated cloned voices across single and multi-speaker scenarios. It also probes robustness against adversarial audio laundering, including additive noise and AAC transcoding, to assess how well different feature representations (learned, spectral, perceptual) generalize and resist degradation. Use when the user wants to benchmark on ElevenLabs (EL), Uberduck (UD), WaveFake (WF), TIMIT-ElevenLabs, or a...
Evaluates a model's ability to determine whether two code snippets share the same semantics or to retrieve relevant code snippets from a repository. It probes semantic code similarity and code retrieval capabilities. Use when the user wants to benchmark on BigCloneBench, POJ-104, or asks about evaluating this task. Reports Overall.
Evaluates the naturalness and quality of synthesized speech across single-speaker, multi-speaker, and multi-emotion TTS models. It measures how closely generated audio matches human ground truth in terms of overall speech quality and emotional similarity. Use when the user wants to benchmark on Baker, AISHELL3, LJSpeech, LibriTTS, Emotional Speech Dataset (ESD), or asks about evaluating this task. Reports MOS.
Evaluates the zero-shot robustness of CLIP models to natural distribution shifts by measuring classification accuracy on four ImageNet-derived datasets. It probes how pre-training data composition and quality affect generalization to out-of-distribution images like sketches, renditions, and novel viewpoints. Use when the user wants to benchmark on ImageNet-V2, ImageNet-R, ImageNet-Sketch, ObjectNet, or asks about evaluating this task. Reports accuracy.
Evaluates zero-shot classification performance of CLIP-based vision-language models on chest X-rays, assessing fairness across demographic subgroups (age, sex, race) and robustness to spurious correlations (presence of chest drains in pneumothorax cases). Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Reports AUPRCadj.
This evaluation protocol probes a model's ability to perform continual learning (CIL) using vision-language models (CLIP) without catastrophic forgetting. It measures how well the model retains knowledge from previous tasks while adapting to new ones, specifically testing stability-plasticity trade-offs under varying task splits and replay settings. Use when the user wants to benchmark on ImageNetR, ImageNetA, CIFAR-100, or asks about evaluating this task. Reports Last.
Evaluates clinical text-to-SQL capabilities by requiring models to generate executable BigQuery queries that perform multi-table joins, temporal reasoning, and patient-similarity cohort analysis on electronic health record data. Use when the user wants to benchmark on CLINSQL, or asks about evaluating this task. Reports Execution Score.
Evaluates the physiological realism and clinical fidelity of synthetic 12-lead ECGs by measuring how often expert clinicians can correctly distinguish them from real clinical recordings, and how accurately they can diagnose specific pathologies in both synthetic and real signals. Use when the user wants to benchmark on MedalCare-XL, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness of clinical NLP models against real-world input noise across four standard tasks. It probes whether models maintain performance when text contains character- or word-level perturbations (e.g., misspellings, deletions, negations) that remain human-readable. Use when the user wants to benchmark on i2b2, MedSTS, MedNLI, or asks about evaluating this task. Reports evaluation scores (accuracy/F1).
Evaluates transformer-based models on clinical relation extraction tasks, measuring their ability to identify and classify relationships between medical entities in text. It compares general vs. clinical-pretrained architectures and binary vs. multi-class classification strategies. Use when the user wants to benchmark on MADE1.0, n2c2, or asks about evaluating this task. Reports strict micro-averaged F1-score.
Evaluates multimodal clinical reasoning and medical knowledge by testing models on standardized medical exams, text-based QA benchmarks, and medical imaging visual question-answering tasks. Use when the user wants to benchmark on USMLE, MedQA, MMLU, MedXpertQA, VQA-RAD, BraTS, PathVQA, Blood Cell VQA, BreaKHis, EMBED, InBreast, CMMD, CBIS-DDS, or asks about evaluating this task. Reports percentage of correct answers.
Evaluates the safety, accuracy, and interaction quality of a clinical AI voice agent using real-world production call data and clinician-validated simulations. It probes system-level reliability across clinical tasks, conversational dynamics, and operational performance to determine if the model handles noisy, multi-turn healthcare conversations safely. Use when the user wants to benchmark on Live Patient Calls, Clinician-Validated Simulations, HEART, or asks about evaluating this task. Repor...