Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,513–7,536 of 20,843 skills
This benchmark evaluates federated time-series anomaly detection systems by measuring how well models maintain detection accuracy when trained across decentralized clients compared to centralized baselines. It probes the robustness of anomaly detection architectures under federated learning protocols and varying degrees of non-IID data partitioning. Use when the user wants to benchmark on Time-series anomaly detection datasets (specific names not provided in excerpt), or asks about evaluating...
Evaluates the ability of convolutional neural networks to classify facial expressions into six basic emotion categories (happiness, surprise, sadness, anger, disgust, fear) from static images and video frames. The benchmark tests both classification accuracy and the model's capacity to generalize across in-the-wild and posed facial expression datasets. Use when the user wants to benchmark on RAF-DB, FER2013, FER GiMeFive, or asks about evaluating this task. Reports Accuracy (%).
Evaluates the faithfulness of abstractive summaries by generating questions from summary sentences and verifying if the answers can be extracted from the source document. It probes a model's ability to avoid hallucination and maintain factual grounding relative to the source text. Use when the user has predictions and gold and needs to compute FEQA.
Evaluates global medium-range weather forecasting accuracy over 10-day lead times across multiple atmospheric variables. It measures point-wise prediction error and anomaly correlation against observed climatology, with explicit latitude weighting to account for spherical grid distortion. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports ACC.
Evaluates the generalization and transferability of emotion recognition models across heterogeneous physiological signal datasets. It probes how well different modeling paradigms (handcrafted features, raw signal deep learning, and contrastive pretraining) perform under subject-independent, cross-dataset, and low-data regimes. Use when the user wants to benchmark on WESAD, NURSE, EMOGNITION, UBFC_PHYS, PhyMER, EmoWear, MAUS, CLAS, CASE, Unobtrusive, CEAP-360VR, ScientISST MOVE, Dapper, ForDig...
This benchmark evaluates how well LLM-generated feedback aligns with human preferences for text summarization, and tests whether preference learning (DPO) using multi-dimensional feedback improves summary quality over supervised fine-tuning. Use when the user wants to benchmark on FeedSum, or asks about evaluating this task. Reports Spearman correlation.
This benchmark evaluates retrieval-based question answering systems and their ability to incorporate post-deployment user feedback. It probes a model's capacity to retrieve relevant answer passages, generate human-like explanations for answer quality, and rerank candidate answers using interactive feedback signals. Use when the user wants to benchmark on FEEDBACKQA, or asks about evaluating this task. Reports accuracy.
Evaluates personalized federated learning (PFL) methods on heterogeneous multi-modal and text-only clients. It probes a model's ability to personalize to its own data distribution ('Self') while maintaining generalization to unseen or other clients' tasks ('Others') under both static and dynamic distribution shifts. Use when the user wants to benchmark on DRAKE, HFLB, Fed-Scope, Fed-Aya, Fed-LLM-Large, or asks about evaluating this task. Reports A_last.
Evaluates the effectiveness of federated hyperparameter optimization (FedHPO) methods across diverse FL tasks. It measures how well optimizers can find high-performing hyperparameter configurations under distributed, communication-constrained, and heterogeneous data settings. Use when the user wants to benchmark on FedHPO-B, or asks about evaluating this task. Reports mean rank.
Evaluates the computational efficiency and security overhead of federated graph neural network training. Measures how system-level metrics like training time, FLOPs, and parameter counts scale across diverse graph datasets under non-IID data partitioning and secure aggregation protocols. Use when the user wants to benchmark on SIDER, BACE, Clintox, BBBP, Tox21, FreeSolv, ESOL, Lipo, hERG, QM9, Ciao, Epinions, CORA, Citeseer, DBLP, PubMed, or asks about evaluating this task. Reports Wall-clock...
Evaluates the effectiveness and efficiency of federated fine-tuning large language models using parameter-efficient fine-tuning (PEFT) algorithms across code generation, general language, and mathematical reasoning tasks under different data heterogeneity and privacy constraints. Use when the user wants to benchmark on Fed-CodeAlpaca, Fed-Dolly, Fed-GSM8K-3, HumanEval, HELM, GSM8K-test, or asks about evaluating this task. Reports Evaluation Scores(%).
Evaluates federated learning algorithms on authentic IoT data modalities under realistic constraints like non-IID data partitioning, label noise, and quantized training. It probes how data heterogeneity, client sampling ratios, and hardware limitations affect model convergence and final performance across diverse sensing tasks. Use when the user wants to benchmark on WISDM-W, WISDM-P, UT-HAR, Widar, VisDrone, CASAS, AEP, EPIC-SOUNDS, or asks about evaluating this task. Reports Accuracy (%).
Evaluates the fairness of federated learning models across heterogeneous client distributions. It probes whether server-level aggregation masks persistent unfairness at the individual client level by measuring demographic disparity and equal opportunity difference on bias-heterogeneous tabular datasets. Use when the user wants to benchmark on attribute-silo, value-silo, attribute-device, value-device, or asks about evaluating this task. Reports DD.
Evaluates the performance of heterogeneous federated fine-tuning methods across multiple NLP domains (instruction following, NLU, finance) under both IID and non-IID data partitioning settings. It probes the ability of LoRA-based federated learning algorithms to maintain model accuracy while accommodating resource-constrained clients with varying rank allocations. Use when the user wants to benchmark on Natural Instructions, GLUE benchmark, FPB, FIQA, TFNS, or asks about evaluating this task....
Evaluates federated learning algorithms on echocardiogram video segmentation tasks under label incompleteness and high data heterogeneity across multiple medical institutions. It probes how FL methods handle missing annotations and conflicting labels from different clinical sites. Use when the user wants to benchmark on Fed-ECHO, or asks about evaluating this task. Reports DICE.
Evaluates federated learning algorithms on ECG classification tasks under non-IID and long-tailed label distribution challenges across multiple medical institutions. It probes how well FL methods generalize across heterogeneous clinical data and handle class imbalance without centralizing all data. Use when the user wants to benchmark on Fed-ECG, or asks about evaluating this task. Reports Micro F1-Score (Mi-F1).
Evaluates a derivative-free camera control policy's ability to align vision-language models in 3D multi-object scenes. It probes robustness to viewpoint changes and object occlusions using minimal demonstration data. Use when the user wants to benchmark on FeatureID-3DS, PartialView-3DS, or asks about evaluating this task. Reports prediction error.
Evaluates vision-language models on financial credit document understanding, covering perception tasks like document type recognition and key information extraction, as well as reasoning tasks such as validity checking and numerical calculation under strict low-latency constraints. Use when the user wants to benchmark on FCMBench, or asks about evaluating this task. Reports F1 score.
Evaluates the accuracy and stability of numerical schemes for solving the functionalized Cahn-Hilliard gradient flow under varying morphological complexity regimes. It tests how well different time-stepping methods (IMEX, SAV, PSD, ETD) capture defect formation, pearling, and interface dynamics in stiff, nonlinear PDE regimes. Use when the user wants to benchmark on FCH Gradient Flow Benchmarks, or asks about evaluating this task. Reports L2 relative error.
Evaluates a model's ability to verify factual claims against retrieved evidence, specifically focusing on claims derived from ambiguous information-seeking questions. It also measures the effectiveness of transfer learning from crowdsourced fact-checking data to professional fact-checking benchmarks. Use when the user wants to benchmark on FaVIQ, Snopes, SciFACT, or asks about evaluating this task. Reports accuracy.
Evaluates a full-duplex turn detection model's ability to classify speech turn states (complete, incomplete, backchannel, wait) in real-time streaming conditions, measuring robustness to acoustic ambiguity, noise, and speech overlap. Use when the user wants to benchmark on FastTurn test set, Easy Turn, Smart Turn, or asks about evaluating this task. Reports Accuracy.
Evaluates multi-object tracking performance in highly crowded, complex urban traffic environments. It probes a tracker's ability to maintain identity consistency under severe occlusions, varying lighting conditions, and across diverse object classes using motion and structural cues rather than appearance models. Use when the user wants to benchmark on FastTrack, or asks about evaluating this task. Reports HOTA.
This evaluation protocol assesses the ability of Large Speech-Language Models to process and understand both short and long-form audio inputs across multiple tasks. It specifically probes speech comprehension, spoken question answering, dialogue understanding, emotion recognition, automatic speech recognition, and long-speech information retrieval under varying compression ratios. Use when the user wants to benchmark on LongSpeech-Eval, speech_QA_iemocap (AIR-Bench), LibriSQA, LibriTTS (OpenA...
Evaluates the long-form factuality of LLM-generated responses by extracting claims, verifying them against evidence, and comparing the system's factuality scores against human annotations. Use when the user wants to benchmark on FaStfact-Bench, or asks about evaluating this task. Reports F₁@K′.