Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,313–9,336 of 21,343 skills
Evaluates the vulnerability of Federated Graph Neural Networks (FedGNN) to classification backdoor attacks across node-level and graph-level tasks. It systematically measures how global factors (data distribution, attacker count, attack timing, overlap) and local factors (trigger size, type, position, poisoning rate) influence attack success and transferability to clean clients. Use when the user wants to benchmark on Unspecified (13 datasets across 6 domains), or asks about evaluating this t...
This benchmark evaluates vision-language and large language models on clinical reasoning for musculoskeletal disorders. It probes capabilities ranging from medical knowledge recall and unimodal interpretation to open-ended multimodal diagnosis, treatment planning, and text-image inconsistency detection. The protocol highlights the performance gap between structured multiple-choice questions and complex, free-form clinical reasoning tasks. Use when the user wants to benchmark on B&J benchmark,...
Evaluates LLMs on real-world financial reasoning tasks, including numerical calculation, temporal reasoning, information extraction, prediction recognition, and knowledge-based QA in Chinese. It probes the models' ability to handle noisy, context-dependent financial data and produce structured, reasoned outputs. Use when the user wants to benchmark on BizFinBench, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to perform quantitative reasoning in business and finance, specifically focusing on synthesizing executable code from financial questions, extracting numeric values from structured and unstructured data, and applying domain-specific financial knowledge. Use when the user wants to benchmark on BizBench, or asks about evaluating this task. Reports accuracy.
Evaluates monocular head pose estimation accuracy by predicting 6DoF rotation (yaw, pitch, roll) from RGB images, comparing absolute single-image regression against relative two-view transformation prediction. Use when the user wants to benchmark on BIWI Kinect Head Pose Database, or asks about evaluating this task. Reports MAE.
Evaluates the robustness of neural network architectures (MLPs, CNNs, and Differentiable Weightless Networks) to parameter bit-flips under varying corruption rates. It measures how task accuracy degrades as a function of bit error rate (BER) and isolates the impact of architectural hyperparameters like precision, width, depth, activation functions, and sparsity. Use when the user wants to benchmark on MLPerf Tiny, or asks about evaluating this task. Reports accuracy.
Evaluates the performance of the BiSHop model on tabular classification and regression tasks, probing its ability to handle mixed feature types, bi-directional cellular learning, and generalized sparse modern Hopfield layers. Use when the user wants to benchmark on Tabular Benchmarks (Adult, Bank, Blastchar, Income, SeismicBump, Shrutime, Spambase, Qsar, Jannis, CR), or asks about evaluating this task. Reports AUC (%).
Evaluates the trade-off between safety and efficiency of energy-function-based safe control algorithms in human-robot interaction and robot co-working scenarios. It tests how well different controllers navigate toward goals while avoiding collisions with human or robot agents across varying dynamic models. Use when the user wants to benchmark on BIS (Benchmark of Interactive Safety), or asks about evaluating this task. Reports efficiency_score.
Evaluates deep learning models on multi-label audio classification for avian bioacoustics, specifically probing robustness to covariate shift, class imbalance, and noisy labels in passive acoustic monitoring scenarios. Use when the user wants to benchmark on BirdSet, or asks about evaluating this task. Reports cmAP.
This benchmark evaluates a model's ability to detect fraudulent user accounts in e-commerce and app review platforms by analyzing their rating patterns and temporal behavior. It probes whether the model can identify users who exhibit extreme rating biases or bursty posting times indicative of coordinated spam or defamation campaigns. Use when the user wants to benchmark on Flipkart, SWM, or asks about evaluating this task. Reports precision@k.
Evaluates an LLM's ability to generate executable Python code for file-based data retrieval tasks from natural language questions. It probes the model's capacity to handle explicit procedural logic, resolve ambiguous user intent, and correctly apply domain knowledge without relying on implicit database semantics. Use when the user wants to benchmark on BIRD-Python, or asks about evaluating this task. Reports LLM-based Execution Accuracy (EX).
Evaluates an LLM's ability to generate syntactically correct and semantically accurate SQL queries from natural language questions over large, real-world databases. It probes database schema understanding, value matching, external knowledge incorporation, and query execution efficiency. Use when the user wants to benchmark on BIRD, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates the capability of large language models to generate correct SQL queries from natural language questions. It specifically probes how annotation noise and errors in benchmark datasets affect model performance and reliability. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of models to perform clinical named entity recognition in Urdu, specifically identifying and classifying biomedical entities like diseases, genes, and proteins within clinical text sequences. It probes sequence labeling capabilities in a low-resource, domain-specific language setting. Use when the user wants to benchmark on BioUNER, or asks about evaluating this task. Reports F1 score.
Evaluates models on insect biodiversity monitoring by testing closed-world species identification, open-world genus-level grouping for novel species, and zero-shot clustering of multimodal embeddings against taxonomic ground truth. Use when the user wants to benchmark on BIOSCAN-5M, or asks about evaluating this task. Reports Fine-tuned accuracy.
Evaluates a compound AI architecture's ability to retrieve, synthesize, and reason across cross-disciplinary scientific knowledge. It probes performance on established single-domain science benchmarks and a novel benchmark specifically designed for bio-AI cross-domain synthesis and reasoning. Use when the user wants to benchmark on LitQA2, GPQA, WMDP, HLE-Bio, BioSage Cross-Disciplinary Benchmark, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large language models on biomedical question-answering, specifically probing their factuality, robustness to linguistic variations (paraphrasing and typos), and susceptibility to demographic bias (age and gender). It distinguishes between extractive and abstractive reasoning capabilities using expert-verified QA pairs from clinical documents. Use when the user wants to benchmark on BioPulse-QA, or asks about evaluating this task. Reports F1.
Evaluates a Persian biomedical large language model's ability to generate accurate, domain-specific long-form answers and summaries. It probes subject-specific knowledge acquisition, knowledge synthesis, and evidence-based reasoning by comparing model outputs against human-written biomedical references. Use when the user wants to benchmark on BioPars-BENCH, or asks about evaluating this task. Reports BERTScore.
This evaluation probes a model's ability to verify scientific claims against provided or retrieved evidence in a binary classification setting. It measures performance on Supported vs. Refuted labels, testing factual grounding, uncertainty calibration, and the impact of atomic decomposition and web corroboration. Use when the user wants to benchmark on BIONLI-300, or asks about evaluating this task. Reports Balanced Accuracy.
Probes an AI agent's ability to answer pharmacology questions by querying federated biomedical knowledge graphs. It evaluates three access methods—direct MCP tools, text-to-Cypher generation, and standalone LLM reasoning—to measure factual accuracy and query efficiency. Use when the user wants to benchmark on BiomedQA, or asks about evaluating this task. Reports Accuracy.
Evaluates the robustness and classification accuracy of deep learning models on biomedical time-series signals (ECG and EEG). It probes the model's ability to handle class imbalance, signal noise, and diverse diagnostic categories without relying on traditional oversampling techniques. Use when the user wants to benchmark on PTB Diagnostic ECG Database, MIT-BIH Arrhythmia Database, UCI Seizure EEG Dataset, or asks about evaluating this task. Reports Accuracy, F1 Score.
This benchmark evaluates large language models on four core biomedical natural language processing tasks: event extraction, relation extraction, named entity recognition, and text classification. It probes the models' ability to identify complex biomedical entities, relationships, and events, as well as classify medical texts, highlighting precision-recall trade-offs in domain-specific applications. Use when the user wants to benchmark on PHEE, Genia2013, Genia2011, DDI, GIT, BioRED, BC5CDR, ...
Probes an LLM's ability to generate syntactically and semantically correct Cypher queries for a biomedical knowledge graph, and execute them to answer domain-specific questions without hallucination. Use when the user wants to benchmark on Custom Biomedical QA Benchmark, or asks about evaluating this task. Reports accuracy.
Evaluates the domain-adaptive post-training of multimodal large language models on biomedical visual question answering tasks, measuring how well models generalize to specialized medical domains using both open and closed evaluation splits. Use when the user wants to benchmark on SLAKE, PathVQA, VQA-RAD, PMC-VQA, or asks about evaluating this task. Reports accuracy.