All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,376
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,433–9,456 of 21,376 skills
- Benchie Fl EvalEvaluates Open Information Extraction (OIE) systems on their ability to extract fact-based triples from text. It uses a conservative exact-matching function with synset-based clustering to penalize non-informative copies and reward precise fact extraction, while also measuring correlation with downstream QA and knowledge base tasks. Use when the user wants to benchmark on BenchIE^FL, or asks about evaluating this task. Reports exact-match.Votes: 0GitHub stars: 3
- Benchecg EvalEvaluates ECG foundation models on diverse clinical tasks including classification, regression, detection, and survival analysis across multiple populations and signal lengths. It probes the model's ability to generalize across datasets, modalities (ECG vs PPG), and long-context temporal dependencies. Use when the user wants to benchmark on CODE-15%, Sleep-Apnea-ECG, MIT-BIH Arrhythmia, PTB-XL, CPSC2018, MIMIC-IV-ECG, Exercise-ECG, or asks about evaluating this task. Reports BenchECG score.Votes: 0GitHub stars: 3
- Bench2drive Vl EvalEvaluates vision-language models in closed-loop autonomous driving by assessing their perception, prediction, planning, and behavioral reasoning capabilities within a CARLA simulator. It probes the model's ability to process raw sensor inputs, generate causally consistent natural language reasoning, and produce valid control actions across diverse and out-of-distribution driving scenarios. Use when the user wants to benchmark on Bench2Drive-VL, or asks about evaluating this task. Reports LLM-...Votes: 0GitHub stars: 3
- Bench2drive Speed EvalEvaluates autonomous driving policies' ability to follow explicit user commands for target speed and overtake/follow behaviors in closed-loop simulations. It measures how well models track desired speeds and execute passing maneuvers while maintaining safety, comfort, and traffic compliance. Use when the user wants to benchmark on Bench2Drive-Speed, or asks about evaluating this task. Reports Speed-Adherence Score.Votes: 0GitHub stars: 3
- Bench Push EvalEvaluates the transferability and performance of reinforcement learning policies for mobile robot navigation and pushing-based manipulation tasks. It probes how well policies trained in simulation handle real-world sim-to-real gaps, clutter, and sparse rewards across varying obstacle densities. Use when the user wants to benchmark on Bench-Push (Maze & Box-Delivery), or asks about evaluating this task. Reports $S_{\text{manip}}$.Votes: 0GitHub stars: 3
- Bells EvalEvaluates LLM supervision systems and frontier models on their ability to detect harmful content across varying harm severities (benign, borderline, harmful) and adversarial sophistication levels (direct prompts vs. jailbreaks). It measures detection capability, robustness to adversarial transformations, and metacognitive coherence between harm classification and response behavior. Use when the user wants to benchmark on BELLS benchmark, or asks about evaluating this task. Reports BELLS Score.Votes: 0GitHub stars: 3
- Belief Tracking Policy EvalEvaluates neural belief tracking models on their accuracy, calibration, and runtime efficiency, and measures how different uncertainty estimates (confidence, total uncertainty, knowledge uncertainty) affect downstream dialogue policy performance in both simulated and human-user environments. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).Votes: 0GitHub stars: 3
- Beir Nl EvalEvaluates zero-shot information retrieval capabilities of lexical, dense, and reranking models on Dutch-language queries and documents. It probes how well models generalize to a machine-translated benchmark without fine-tuning, measuring both ranking quality and recall performance. Use when the user wants to benchmark on MSMARCO, TREC-COVID, NFCorpus, NQ, HotpotQA, FiQA-2018, ArguAna, Touche-2020, CQADupstack, Quora, DBPedia, SciDocs, SciFact, FEVER, Climate-FEVER, or asks about evaluating th...Votes: 0GitHub stars: 3
- Beir EvalEvaluates zero-shot information retrieval capabilities across 18 diverse domains and query types. It probes a model's ability to retrieve relevant documents without domain-specific fine-tuning, highlighting performance variations due to domain shifts, query length, and lexical versus semantic matching. Use when the user wants to benchmark on BEIR, or asks about evaluating this task. Reports nDCG@10.Votes: 0GitHub stars: 3
- Being H05 Robot EvalEvaluates cross-embodiment generalization and manipulation capabilities of Vision-Language-Action models across heterogeneous real robots and simulation benchmarks. Probes spatial reasoning, long-horizon planning, bimanual coordination, and zero-shot transfer to unseen task-embodiment pairs. Use when the user wants to benchmark on Real-robot task suite, LIBERO, RoboCasa, or asks about evaluating this task. Reports success rate (%).Votes: 0GitHub stars: 3
- Behavioral Prediction EvalPredicts individual strategic decisions by conditioning on structured psychometric trait profiles (e.g., Big Five personality traits). It probes a model's ability to map high-dimensional psychological embeddings to discrete behavioral outcomes in unseen situational contexts. Use when the user wants to benchmark on Strategic Scenario Dataset, or asks about evaluating this task. Reports balanced accuracy, macro-F1.Votes: 0GitHub stars: 3
- Behavioral Fraud Patterns EvalThis benchmark evaluates whether synthetic tabular data generators preserve complex behavioral fraud patterns beyond simple statistical fidelity. It specifically probes temporal inter-event timing, burst-like transaction sequences, shared-infrastructure graph motifs across accounts, and velocity-rule trigger rates. Use when the user wants to benchmark on IEEE-CIS Fraud Detection, Amazon Fraud Dataset, or asks about evaluating this task. Reports P1: IET W1.Votes: 0GitHub stars: 3
- Behavior1k EvalEvaluates embodied AI agents on long-horizon, human-centered manipulation tasks in a realistic physics-based simulation. It probes the agent's ability to plan and execute complex sequences of action primitives (pick, place, navigate, etc.) while handling rigid, articulated, and deformable objects. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports task success rate.Votes: 0GitHub stars: 3
- Behavior 1k EvalThis evaluation probes a robot's ability to execute long-horizon, zero-shot rearrangement tasks in unexplored indoor-outdoor environments using grounded language reasoning. It measures how well the system interprets natural language instructions, reasons over 3D scene graphs, and coordinates sequential manipulation actions to satisfy multiple goal conditions. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate (SR).Votes: 0GitHub stars: 3
- Beexai EvalEvaluates post-hoc explainable AI (XAI) attribution methods on tabular data across binary classification, multi-class classification, and regression tasks. It measures how well feature importance scores align with core XAI desiderata—faithfulness, plausibility, robustness, and complexity—using ground-truth-aligned quantitative metrics. Use when the user wants to benchmark on inria-soda/tabular-benchmark, OpenML-CC18 Curated Classification, or asks about evaluating this task. Reports Infidelity.Votes: 0GitHub stars: 3
- Beerqa EvalEvaluates open-domain question answering systems on their ability to retrieve and synthesize information across varying numbers of reasoning steps (single-hop to three-hop) without relying on structured metadata or predefined hop counts. Use when the user wants to benchmark on SQuAD Open, HotpotQA, BeerQA, or asks about evaluating this task. Reports exact match (EM).Votes: 0GitHub stars: 3
- Beep Korean Toxic Speech EvalThis benchmark evaluates models' ability to detect social bias (gender and other types) and hate speech in Korean online news comments. It probes whether models can distinguish between hate speech, offensive language, and neutral comments, and whether incorporating bias labels improves hate speech detection. Use when the user wants to benchmark on BEEP!, or asks about evaluating this task. Reports F1.Votes: 0GitHub stars: 3
- Bee 8b Mllm EvalEvaluates the visual reasoning, factual accuracy, OCR, chart understanding, and mathematical capabilities of fully open multimodal large language models (MLLMs) against a comprehensive suite of established benchmarks. The protocol tests the model's ability to process images and text prompts, generate responses in a thinking mode, and achieve high scores across general VQA, document/chart analysis, and complex math/reasoning tasks. Use when the user wants to benchmark on Bee-8B Evaluation Benc...Votes: 0GitHub stars: 3
- Beavertails Moderation EvalThis evaluation probes the safety moderation and context-comprehension capabilities of external text moderation APIs. It measures how well automated systems align with human and expert-preference labels when assessing harmfulness across specific risk categories in QA pairs. Use when the user wants to benchmark on BeaverTails Evaluation Dataset, or asks about evaluating this task. Reports agreement.Votes: 0GitHub stars: 3
- Beaver EvalProbes LLMs' ability to generate correct SQL queries from natural language questions over complex, enterprise-scale databases. It specifically evaluates handling of high schema complexity, multi-table joins, aggregations, and column-to-table mapping in real-world business contexts. Use when the user wants to benchmark on BEAVER, or asks about evaluating this task. Reports execution accuracy.Votes: 0GitHub stars: 3
- Beatv2 EvalEvaluates cross-dataset generalization of co-speech gesture generation on a standard English benchmark, measuring gesture quality, beat consistency, and diversity. Use when the user wants to benchmark on BEATv2, or asks about evaluating this task. Reports FGD.Votes: 0GitHub stars: 3
- Beat It EvalEvaluates a model's ability to generate 3D dance motions that are temporally synchronized with musical beats and controllable via sparse keyframes, while maintaining kinematic plausibility and motion diversity. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports BAS.Votes: 0GitHub stars: 3
- Beat Backdoor Detection EvalEvaluates the ability of a black-box defense mechanism to detect backdoor-unaligned samples in LLMs by measuring changes in the model's refusal behavior when a malicious probe is concatenated to the input. Use when the user wants to benchmark on MaliciousInstruct + Advbench + UltraChat-200k, or asks about evaluating this task. Reports AUROC.Votes: 0GitHub stars: 3
- Bearllm Fault Diagnosis EvalEvaluates a multimodal LLM framework's ability to perform bearing fault diagnosis, anomaly detection, and maintenance recommendation using vibration signals and textual prompts. It probes cross-condition generalization and zero-shot transfer across diverse industrial bearing datasets. Use when the user wants to benchmark on MBHM, JUST, IMS, CWRU, XJTU, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3