All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 22,716
- 947
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 10,705–10,728 of 22,716 skills
- Bird EvalEvaluates an LLM's ability to generate syntactically correct and semantically accurate SQL queries from natural language questions over large, real-world databases. It probes database schema understanding, value matching, external knowledge incorporation, and query execution efficiency. Use when the user wants to benchmark on BIRD, or asks about evaluating this task. Reports Execution Accuracy (EX).Votes: 0GitHub stars: 3
- Bird Bench EvalEvaluates the capability of large language models to generate correct SQL queries from natural language questions. It specifically probes how annotation noise and errors in benchmark datasets affect model performance and reliability. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Biouner EvalEvaluates the ability of models to perform clinical named entity recognition in Urdu, specifically identifying and classifying biomedical entities like diseases, genes, and proteins within clinical text sequences. It probes sequence labeling capabilities in a low-resource, domain-specific language setting. Use when the user wants to benchmark on BioUNER, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Bioscan 5m EvalEvaluates models on insect biodiversity monitoring by testing closed-world species identification, open-world genus-level grouping for novel species, and zero-shot clustering of multimodal embeddings against taxonomic ground truth. Use when the user wants to benchmark on BIOSCAN-5M, or asks about evaluating this task. Reports Fine-tuned accuracy.Votes: 0GitHub stars: 3
- Biosage Scientific EvalEvaluates a compound AI architecture's ability to retrieve, synthesize, and reason across cross-disciplinary scientific knowledge. It probes performance on established single-domain science benchmarks and a novel benchmark specifically designed for bio-AI cross-domain synthesis and reasoning. Use when the user wants to benchmark on LitQA2, GPQA, WMDP, HLE-Bio, BioSage Cross-Disciplinary Benchmark, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Biopulse Qa EvalThis benchmark evaluates large language models on biomedical question-answering, specifically probing their factuality, robustness to linguistic variations (paraphrasing and typos), and susceptibility to demographic bias (age and gender). It distinguishes between extractive and abstractive reasoning capabilities using expert-verified QA pairs from clinical documents. Use when the user wants to benchmark on BioPulse-QA, or asks about evaluating this task. Reports F1.Votes: 0GitHub stars: 3
- Bionpars Bench EvalEvaluates a Persian biomedical large language model's ability to generate accurate, domain-specific long-form answers and summaries. It probes subject-specific knowledge acquisition, knowledge synthesis, and evidence-based reasoning by comparing model outputs against human-written biomedical references. Use when the user wants to benchmark on BioPars-BENCH, or asks about evaluating this task. Reports BERTScore.Votes: 0GitHub stars: 3
- Bionli 300 EvalThis evaluation probes a model's ability to verify scientific claims against provided or retrieved evidence in a binary classification setting. It measures performance on Supported vs. Refuted labels, testing factual grounding, uncertainty calibration, and the impact of atomic decomposition and web corroboration. Use when the user wants to benchmark on BIONLI-300, or asks about evaluating this task. Reports Balanced Accuracy.Votes: 0GitHub stars: 3
- Biomedqa EvalProbes an AI agent's ability to answer pharmacology questions by querying federated biomedical knowledge graphs. It evaluates three access methods—direct MCP tools, text-to-Cypher generation, and standalone LLM reasoning—to measure factual accuracy and query efficiency. Use when the user wants to benchmark on BiomedQA, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Biomedical Timeseries Classification EvalEvaluates the robustness and classification accuracy of deep learning models on biomedical time-series signals (ECG and EEG). It probes the model's ability to handle class imbalance, signal noise, and diverse diagnostic categories without relying on traditional oversampling techniques. Use when the user wants to benchmark on PTB Diagnostic ECG Database, MIT-BIH Arrhythmia Database, UCI Seizure EEG Dataset, or asks about evaluating this task. Reports Accuracy, F1 Score.Votes: 0GitHub stars: 3
- Biomedical Nlp EvalThis benchmark evaluates large language models on four core biomedical natural language processing tasks: event extraction, relation extraction, named entity recognition, and text classification. It probes the models' ability to identify complex biomedical entities, relationships, and events, as well as classify medical texts, highlighting precision-recall trade-offs in domain-specific applications. Use when the user wants to benchmark on PHEE, Genia2013, Genia2011, DDI, GIT, BioRED, BC5CDR, ...Votes: 0GitHub stars: 3
- Biomedical Cypher EvalProbes an LLM's ability to generate syntactically and semantically correct Cypher queries for a biomedical knowledge graph, and execute them to answer domain-specific questions without hallucination. Use when the user wants to benchmark on Custom Biomedical QA Benchmark, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Biomed Vqa EvalEvaluates the domain-adaptive post-training of multimodal large language models on biomedical visual question answering tasks, measuring how well models generalize to specialized medical domains using both open and closed evaluation splits. Use when the user wants to benchmark on SLAKE, PathVQA, VQA-RAD, PMC-VQA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Biomed Enriched EvalEvaluates the biomedical knowledge, clinical reasoning, and domain-specific comprehension of language models using multiple-choice question-answering benchmarks across anatomy, clinical medicine, genetics, and multilingual medical QA. Use when the user wants to benchmark on MMLU Professional Medicine, MedQA, MedMCQA, PubMedQA, FrenchMedMCQA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Biological Visual EvalEvaluates frozen visual embedding extractors on ecological trait alignment, fine-grained intra-species variation preservation, and zero-shot/few-shot transfer learning across diverse biological domains. Use when the user wants to benchmark on FishNet, NeWT, AwA2, Herb, PlantDoc, Life stage-Diff/Align, Sex-Diff/Align, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Biological Mllm Merging EvalEvaluates the ability of merged multimodal large language models to perform cross-modal biological reasoning tasks, specifically predicting molecular interactions with proteins/cells and predicting enzyme functionality. It probes whether embedding-space-aware merging preserves modality-specific expertise better than parameter-space heuristics or fine-tuning. Use when the user wants to benchmark on Biological MLLM Interaction & Functionality Benchmarks, or asks about evaluating this task. Repo...Votes: 0GitHub stars: 3
- Biodenoising EvalEvaluates the ability of audio denoising models to remove background noise from animal vocalization recordings without access to clean reference data during training. It measures how well models generalize across diverse species and environments using synthetic mixtures and a held-out benchmark set. Use when the user wants to benchmark on Biodenoising benchmark set, or asks about evaluating this task. Reports SI-SDR.Votes: 0GitHub stars: 3
- Bioclinical Modernbert EvalEvaluates long-context clinical NLP encoders on biomedical entity recognition, clinical text classification, and demographic information extraction. Probes the model's ability to process full-length clinical notes (up to 8,192 tokens) and retain domain-specific knowledge without truncation. Use when the user wants to benchmark on ChemProt, Phenotype, Social History, DEID, COS, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Biocap EvalEvaluates zero-shot species classification and fine-grained text-image retrieval capabilities in biological domains. Probes the model's ability to align visual features with taxonomic labels and descriptive natural language without task-specific fine-tuning. Use when the user wants to benchmark on NABirds, Meta-Album (Plankton, Insects, Insects 2), IDLE-OO Camera Traps, Rare Species, PlantNet, Fungi, PlantVillage, Med. Leaf, INQUIRE-Rerank, Cornell Bird, PlantID, or asks about evaluating this...Votes: 0GitHub stars: 3
- Biobert Re EvalTests a model's capability to extract biomedical relations (gene-disease, gene-chemical) from text using minimal task-specific modifications. It probes whether domain-specific pre-training improves relation classification on small-scale biomedical corpora. Use when the user wants to benchmark on GAD, EU-ADR, CHEMPROT, or asks about evaluating this task. Reports entity-level F1.Votes: 0GitHub stars: 3
- Biobert Qa EvalEvaluates factoid question answering performance on small biomedical datasets, testing transfer learning effectiveness from domain-specific pre-training. It measures how well the model retrieves exact or lenient answers to biomedical queries. Use when the user wants to benchmark on BioASQ 4b, BioASQ 5b, BioASQ 6b, or asks about evaluating this task. Reports Mean Reciprocal Rank (MRR).Votes: 0GitHub stars: 3
- Biobert Ner EvalEvaluates a model's ability to identify and classify biomedical entities (diseases, drugs/chemicals, genes/proteins, species) in text using transfer learning from domain-specific pre-training. It tests whether contextualized representations improve entity boundary detection and classification on small biomedical corpora. Use when the user wants to benchmark on NCBI disease, 2010 i2b2/VA, BC5CDR, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, or asks about evaluating this task. Reports entity...Votes: 0GitHub stars: 3
- Bing Chat Open Domain EvalEvaluates an LLM's ability to perform zero-shot dialogue segmentation and joint state tracking on real-world, open-domain human-LLM conversations. It probes the model's capacity to identify topic shifts, assign intent and domain labels, and maintain context over long multi-turn interactions without hallucination. Use when the user wants to benchmark on Bing Chat (Internal Human-LLM Dialogue Dataset), or asks about evaluating this task. Reports JGA (I/D).Votes: 0GitHub stars: 3
- Binding Pose Prediction EvalEvaluates a model's ability to predict the 3D spatial arrangement (binding pose) of a small molecule ligand when docked to a target protein structure. It probes geometric reasoning, conformational sampling, and the capacity to generate physically plausible protein-ligand complexes. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports L-RMSD.Votes: 0GitHub stars: 3