Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,393–7,416 of 20,843 skills
Evaluates the ability of large protein language models to predict protein fitness under constrained, low-data scenarios. It probes mutation-level generalization, overfitting risks, and the impact of model depth and structural information on predictive accuracy across diverse protein families. Use when the user wants to benchmark on FLIP benchmark, or asks about evaluating this task. Reports MSE.
This evaluation probes the inference throughput and generation quality of LLMs across diverse hardware and software configurations. It also tests a predictive modeling framework designed to optimize system co-design by forecasting performance metrics based on model and hardware features. Use when the user wants to benchmark on OpenOrca, Open MLPerf Dataset, or asks about evaluating this task. Reports Tokens/s.
Evaluates multilingual spoken language understanding (SLU) across 102 languages for topical classification and 92 languages for spoken multiple-choice QA, testing cross-lingual transfer, speech-to-text translation, and robustness to audio quality variations. Use when the user wants to benchmark on SIB-Fleurs, Belebele-Fleurs, or asks about evaluating this task. Reports accuracy.
Evaluates universal speech representations across 102 languages using few-shot learning on parallel speech data. Probes capabilities in automatic speech recognition (ASR), speech language identification, and retrieval tasks. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports character level error rate.
This evaluation probes a model's ability to perform automatic speech recognition in low-resource and zero-supervised settings by leveraging joint speech-text representation learning. It specifically measures how well the model can transcribe unseen languages using only untranscribed audio and graphemic text, without relying on manually labeled speech data. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports CER.
Evaluates a multimodal LLM's ability to perform visual reasoning across heterogeneous medical modalities (2D images, 3D volumes, videos) and generate clinical reports. It probes diagnostic accuracy, cross-modal generalization, temporal understanding, and structured medical knowledge integration. Use when the user wants to benchmark on OmniMedVQA, PMC-VQA, VQA-RAD, PathVQA, SLAKE, MIMIC-CXR, IU-Xray, M3D-VQA, MedVideoBench, or asks about evaluating this task. Reports accuracy, ROUGE-L, CIDEr.
Evaluates a unified vision-language foundation model across 35 downstream tasks spanning vision classification, natural language understanding, and multimodal reasoning/retrieval. It probes the model's ability to generalize from joint unimodal and multimodal pretraining to zero-shot and fine-tuned downstream settings. Use when the user wants to benchmark on GLUE (MNLI, CoLA, MRPC, QQP, SST-2, QNLI, RTE, STS-B), 22 Vision Datasets (ImageNet, Food101, CIFAR10, CIFAR100, Cars, Aircraft, DTD, Pet...
Evaluates the ability of quantum and classical models to perform full-waveform inversion (FWI) by predicting subsurface velocity maps from scaled seismic waveform data. It probes the effectiveness of physics-guided data scaling and layer-wise variational quantum circuit designs in geophysical imaging tasks. Use when the user wants to benchmark on FlatVelA, or asks about evaluating this task. Reports SSIM.
Evaluates multi-agent coordination and path planning for train rescheduling in a dynamic grid-world railway simulation. It probes the ability of agents to adapt to partial observability, handle congestion, and coordinate under switching constraints and dynamic disruptions. Use when the user wants to benchmark on Flatland Competition 2020, or asks about evaluating this task. Reports overall score.
Evaluates LLMs on fine-grained alignment capabilities by decomposing instruction-following performance into 12 sub-skills across four domains (Logical Thinking, Background Knowledge, Problem Handling, User Alignment). It measures how well models adhere to specific quality criteria like factuality, logical robustness, and harmlessness on a per-instance basis. Use when the user wants to benchmark on FLASK, or asks about evaluating this task. Reports FLASK skill score.
Evaluates the robustness and efficiency of text-guided visual token pruning in large multimodal models. It probes whether aggressive token compression (retaining 32–128 tokens for images, 114–455 for video) degrades performance on image and video question-answering tasks, and measures cross-modal grounding quality via spatial alignment and semantic overlap metrics. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MMBench-CN, MM Vet, TGIF-...
Evaluates the effectiveness of various Retrieval-Augmented Generation (RAG) methods across text and multimodal question-answering tasks. It probes how different retrieval strategies, context compression techniques, and generator optimizations impact answer accuracy and faithfulness on single-hop and multi-hop datasets. Use when the user wants to benchmark on NQ, TriviaQA, HotpotQA, 2WikiMultihopQA, Gaokao-MM, MultimodalQA, MathVista, or asks about evaluating this task. Reports Acc.
Evaluates the effectiveness of a frequency-domain-guided KV cache compression method (FlashCache) on multimodal long-context understanding tasks. It measures how well the model preserves accuracy under varying KV cache retention ratios and quantifies the computational overhead and decoding latency compared to baseline eviction methods. Use when the user wants to benchmark on MileBench, MUIRBench, MMMU, V*, HR-Bench, FAVOR-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates bilingual (Spanish-English) financial understanding, prediction, and generation capabilities of LLMs. Probes cross-lingual transfer, domain-specific instruction following, and performance disparity between high-resource and low-resource financial tasks. Use when the user wants to benchmark on FLARE-ES, or asks about evaluating this task. Reports Acc, F1.
Evaluates the instruction-tuning effectiveness of models trained on the Flan 2022 collection across held-in, chain-of-thought, and held-out benchmarks. It probes zero-shot and few-shot generalization capabilities on reasoning, knowledge, and natural language understanding tasks. Use when the user wants to benchmark on MMLU, BBH, or asks about evaluating this task. Reports zero-shot/few-shot accuracy.
Assesses practical financial application capabilities across a hierarchical framework of 10 primary and 21 secondary scenarios. It probes models' ability to perform real-world financial tasks such as compliance checking, document generation, risk control, data extraction, and client analysis. Use when the user wants to benchmark on FLAME-Sce, or asks about evaluating this task. Reports usability rate.
Evaluates federated learning algorithms for decentralized robotic manipulation across heterogeneous environments. It probes a model's ability to generalize from distributed, non-IID demonstrations under visual and physical perturbations, measuring both action prediction fidelity and task completion success. Use when the user wants to benchmark on FLAME, or asks about evaluating this task. Reports RMSE.
Evaluates foundation and reasoning-reinforced language models across 20 core financial NLP tasks. It probes capabilities in numeric reasoning, entity classification, information retrieval, question answering, and summarization, while also measuring inference efficiency and cost. Use when the user wants to benchmark on FLaME, or asks about evaluating this task. Reports F1.
Evaluates financial domain knowledge and certification exam readiness in Chinese and English. It probes models' ability to answer multiple-choice questions across 14 professional financial certifications with varying difficulty levels. Use when the user wants to benchmark on FLAME-Cer, or asks about evaluating this task. Reports accuracy rate.
Evaluates semantic segmentation models for fine-grained land cover classification and crop type mapping using multi-sensor remote sensing imagery. It probes the model's ability to fuse spatial, spectral, and temporal modalities (aerial RGBI, SPOT, Sentinel-1/2, DEM) for pixel-level prediction at 20 cm resolution. Use when the user wants to benchmark on FLAIR-HUB, or asks about evaluating this task. Reports mIoU.
Evaluates the ability of models to perform high-resolution land-cover semantic segmentation on aerial imagery. It probes robustness to spatial, temporal, and multi-sensor domain shifts, as well as handling radiometric inconsistencies and phenological variations across diverse landscapes. Use when the user wants to benchmark on FLAIR-one, or asks about evaluating this task. Reports mIoU.
Evaluates federated learning models on real-world, non-IID image data with user-level heterogeneity and long-tailed label distributions. It probes how model convergence and multi-label classification performance degrade under privacy constraints (differential privacy) and distributed training compared to centralized baselines. Use when the user wants to benchmark on FLAIR, or asks about evaluating this task. Reports averaged precision (AP).
Evaluates large reasoning models on automatically verifiable textual problem-solving tasks, including academic coursework, word puzzles, cipher deciphering, and algorithmic coding. It probes the models' ability to follow instructions, perform logical deduction, and produce correctly formatted final answers under varying reasoning effort settings. Use when the user wants to benchmark on FlagEval Textual, or asks about evaluating this task. Reports accuracy.
Evaluates federated learning methods for medical image segmentation under non-IID data distributions, measuring segmentation accuracy and robustness across multiple clinical tasks and imaging modalities. It compares generic and personalized FL approaches against local training baselines to assess client drift, fairness, and generalization. Use when the user wants to benchmark on Fed-Vessel, Fed-Prostate, Fed-COSAS, Fed-BUS, Fed-MG, Fed-Polyp, Fed-Pancreas, Fed-M&Ms, FeTS2022, or asks about ev...