All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,376
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,385–9,408 of 21,376 skills
- Biobert Qa EvalEvaluates factoid question answering performance on small biomedical datasets, testing transfer learning effectiveness from domain-specific pre-training. It measures how well the model retrieves exact or lenient answers to biomedical queries. Use when the user wants to benchmark on BioASQ 4b, BioASQ 5b, BioASQ 6b, or asks about evaluating this task. Reports Mean Reciprocal Rank (MRR).Votes: 0GitHub stars: 3
- Biobert Ner EvalEvaluates a model's ability to identify and classify biomedical entities (diseases, drugs/chemicals, genes/proteins, species) in text using transfer learning from domain-specific pre-training. It tests whether contextualized representations improve entity boundary detection and classification on small biomedical corpora. Use when the user wants to benchmark on NCBI disease, 2010 i2b2/VA, BC5CDR, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, or asks about evaluating this task. Reports entity...Votes: 0GitHub stars: 3
- Bing Chat Open Domain EvalEvaluates an LLM's ability to perform zero-shot dialogue segmentation and joint state tracking on real-world, open-domain human-LLM conversations. It probes the model's capacity to identify topic shifts, assign intent and domain labels, and maintain context over long multi-turn interactions without hallucination. Use when the user wants to benchmark on Bing Chat (Internal Human-LLM Dialogue Dataset), or asks about evaluating this task. Reports JGA (I/D).Votes: 0GitHub stars: 3
- Binding Pose Prediction EvalEvaluates a model's ability to predict the 3D spatial arrangement (binding pose) of a small molecule ligand when docked to a target protein structure. It probes geometric reasoning, conformational sampling, and the capacity to generate physically plausible protein-ligand complexes. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports L-RMSD.Votes: 0GitHub stars: 3
- Binary RewardThis section outlines the reward design used during reinforcement learning training, which functions as the primary evaluation metric. It probes the model's ability to perform multimodal logical reasoning and produce correctly formatted final answers. The protocol relies on a strict binary correctness check rather than partial credit for reasoning steps. Use when the user has predictions and gold and needs to compute binary_reward.Votes: 0GitHub stars: 3
- Bimr Interpretability EvalThis benchmark evaluates the reasoning interpretability of knowledge graph completion models by measuring how well their generated multi-hop paths or rules can be understood and validated. It probes whether models produce semantically reasonable explanations rather than just statistically valid paths, highlighting the gap between link prediction accuracy and actual explainability. Use when the user wants to benchmark on WD15K, FB15K-237, or asks about evaluating this task. Reports GI (Global ...Votes: 0GitHub stars: 3
- Bimedx2 Medical EvalEvaluates a bilingual (Arabic-English) large multimodal model's ability to understand diverse medical imaging modalities, answer visual questions, and generate or summarize clinical reports. It probes factual accuracy, clinical relevance, and linguistic quality across text-only, visual-question-answering, and report-generation tasks. Use when the user wants to benchmark on BiMed-MBench, Rad-VQA, SLAKE, Path-VQA, MIMIC-CXR, MIMIC-III, or asks about evaluating this task. Reports accuracy, F1, F...Votes: 0GitHub stars: 3
- Billsum EvalThis benchmark evaluates the ability of models to automatically generate concise and accurate summaries of complex, nested legislative texts. It probes extractive and abstractive summarization capabilities in a highly technical, domain-specific legal context. Use when the user wants to benchmark on BillSum, or asks about evaluating this task. Reports ROUGE F-Score.Votes: 0GitHub stars: 3
- Bigpatent EvalEvaluates abstractive summarization models on patent documents, probing their ability to capture global discourse structure, maintain entity coherence, and generate novel content without excessive repetition or fabrication. Use when the user wants to benchmark on BIGPATENT, or asks about evaluating this task. Reports ROUGE-1 F1.Votes: 0GitHub stars: 3
- Bigobench EvalEvaluates whether LLMs can accurately predict the time and space complexity of given code snippets and generate new code that satisfies explicit complexity constraints. It probes algorithmic reasoning and scalability awareness beyond mere syntactic or functional correctness. Use when the user wants to benchmark on BigO(Bench), or asks about evaluating this task. Reports Pass@k.Votes: 0GitHub stars: 3
- Bigearthnet Txt EvalEvaluates vision-language models on remote sensing tasks including image captioning, binary visual question answering, multiple-choice questions, and referring expression/point detection. It probes the models' ability to understand multi-sensor (SAR + multispectral) and RGB Earth observation imagery, follow complex spatial instructions, and generate accurate land-use/land-cover descriptions or localized bounding boxes. Use when the user wants to benchmark on BigEarthNet.txt, or asks about eva...Votes: 0GitHub stars: 3
- Bigcodebench EvalEvaluates large language models' ability to generate correct, executable code for complex programming tasks requiring diverse function calls and compositional reasoning. It also probes instruction-following capabilities by comparing performance on verbose prompts versus condensed natural-language instructions. Use when the user wants to benchmark on BigCodeBench, or asks about evaluating this task. Reports Pass@1.Votes: 0GitHub stars: 3
- Bigbench Arithmetic EvalEvaluates large language models on arithmetic reasoning across addition, subtraction, multiplication, and division tasks with varying digit lengths. It probes the model's ability to handle large-number computation, number tokenization consistency, and stepwise reasoning without relying on external tools. Use when the user wants to benchmark on BIG-bench arithmetic, Extra arithmetic tasks, or asks about evaluating this task. Reports exact string match.Votes: 0GitHub stars: 3
- Big2015 Malware Detection EvalEvaluates the effectiveness of multimodal visual feature fusion (grayscale, entropy graph, SimHash) using VGG16 for binary malware classification and family detection. It probes the model's ability to handle imbalanced malware datasets and detect obfuscated binaries. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F1-score.Votes: 0GitHub stars: 3
- Big2015 EvalEvaluates the ability of deep learning models to classify malware binaries into their respective family types using image-based representations, specifically probing performance on imbalanced class distributions. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F-Score.Votes: 0GitHub stars: 3
- Big Bench Predictability EvalEvaluates the predictability of LLM performance across diverse BIG-bench tasks using an MLP-based predictor. It probes how well model scale, task type, and in-context examples correlate with actual benchmark scores, and tests robustness under different holdout strategies. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports R².Votes: 0GitHub stars: 3
- Big Bench EvalEvaluates large language models on probabilistic reasoning over sets, including conceptual combination, semantic outlier detection, and logical deduction. It probes the model's ability to decouple knowledge retrieval from final inference and aggregate uncertain facts without relying solely on self-attention. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Big Ann Competition EvalEvaluates approximate nearest neighbor (ANN) indexing methods across four realistic, constrained workloads: filtered search, out-of-distribution data, sparse vectors, and streaming updates. It probes the trade-off between search accuracy (recall) and query throughput under strict memory and time constraints. Use when the user wants to benchmark on Big ANN Challenge Datasets, or asks about evaluating this task. Reports 10-recall@10.Votes: 0GitHub stars: 3
- Big 15 Malware Detection EvalEvaluates a graph-based adversarial domain adaptation model's ability to classify Windows malware and benign binaries under concept drift using minimal labeled samples. It probes the model's robustness to real-world malware evolution by comparing performance against baselines on the Big-15 dataset. Use when the user wants to benchmark on Big-15, or asks about evaluating this task. Reports performance metrics.Votes: 0GitHub stars: 3
- Bibldr Drug Repositioning EvalEvaluates a model's ability to predict novel drug-disease associations by modeling them as a recommendation task using bidirectional behavioral sequences and prototype spaces. It probes cold-start generalization and robustness to highly sparse interaction data. Use when the user wants to benchmark on Gdataset, Cdataset, LRSSL, or asks about evaluating this task. Reports AUPRC.Votes: 0GitHub stars: 3
- Bib Ref Parser EvalEvaluates open-source bibliographic reference and citation parsers on their ability to extract structured metadata fields (author, source, year, volume, issue, page, organization) from raw citation text. Compares machine learning-based versus rule-based approaches, and assesses the impact of domain-specific retraining. Use when the user wants to benchmark on Unspecified bibliographic dataset, or asks about evaluating this task. Reports F1.Votes: 0GitHub stars: 3
- Biasinear EvalEvaluates the robustness and sensitivity of multimodal large language models (MLLMs) to perturbations in spoken multiple-choice questions. It probes how models handle variations in language, accent, speaker gender, and answer option ordering, measuring both absolute correctness and prediction stability across conditions. Use when the user wants to benchmark on BiasInEar, or asks about evaluating this task. Reports Question Entropy.Votes: 0GitHub stars: 3
- Biasig EvalEvaluates text-to-image models for multi-dimensional social biases across demographic attributes (sex, race, age) by measuring implicit distributional divergence, explicit instruction-following accuracy, and whether bias manifests as ignorance or discrimination. Use when the user wants to benchmark on BiasIG, or asks about evaluating this task. Reports Implicit Bias Score ($S_{sum}$).Votes: 0GitHub stars: 3
- Bias Quantization EvalEvaluates how weight-activation quantization affects model capabilities, stereotypes, fairness, toxicity, and sentiment across demographic subgroups. It probes whether aggressive compression amplifies historical bias, disparate outcomes, and inter-subgroup disparities in generated text. Use when the user wants to benchmark on MMLU, RedditBias, WinoBias, DiscrimEval, DT-Fairness, BOLD, StereoSet, or asks about evaluating this task. Reports MMLU accuracy.Votes: 0GitHub stars: 3