All categories
Business & Operations
Operations, strategy, finance, sales, support, management, and planning
- 31,233
- 1,302
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse business & operations skills
Showing 10,801–10,824 of 31,233 skills
- Tpc H Tpc C Energy EvalEvaluates the energy efficiency and performance trade-offs of a distributed database cluster versus a single high-end server under OLAP and OLTP workloads. It probes the system's ability to maintain energy proportionality through dynamic node scaling and measures the overhead incurred during data migration and cluster reconfiguration. Use when the user wants to benchmark on TPC-H, TPC-C, or asks about evaluating this task. Reports energy consumption per query.Votes: 0GitHub stars: 3
- Tcmsd Sd EvalEvaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification. Use when the user wants to benchmark on TCM-SD, or asks about evaluating this task. Reports Macro-F1.Votes: 0GitHub stars: 3
- Swim EvalEvaluates real-time instance segmentation performance of lightweight models under strict onboard hardware constraints, measuring inference speed, memory usage, and segmentation accuracy for spacecraft boundary localization. Use when the user wants to benchmark on SWiM, or asks about evaluating this task. Reports RAM_footprint.Votes: 0GitHub stars: 3
- Robointer Vqa EvalEvaluates embodied reasoning capabilities of vision-language models on robotic manipulation tasks. It probes spatial understanding and generation (e.g., object grounding, grasp pose prediction) and temporal understanding and generation (e.g., motion trace reconstruction, multi-step planning) across diverse indoor and tabletop scenarios. Use when the user wants to benchmark on RoboInter-VQA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Representation Benchmark EvalEvaluates how input representation choices—quantization granularity, value encoding, temporal encoding, and vocabulary remapping—affect downstream predictive performance on clinical outcomes. It probes the model's ability to extract and utilize structured medical event sequences for binary classification and regression tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.Votes: 0GitHub stars: 3
- Redundancy Detection EvalEvaluates a model's ability to determine whether pairs of P, I, or O spans in an abstract refer to the same underlying information. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.Votes: 0GitHub stars: 3
- Phycritic EvalEvaluates multimodal models' ability to perform physical reasoning, spatial cognition, and egocentric task planning, as well as their capacity to act as reliable critics/judges for physical AI tasks. Use when the user wants to benchmark on PhyCritic-Bench, VL-RewardBench, Multimodal-RewardBench, CosmosReason1-Bench, CV-Bench, EgoPlanBench2, or asks about evaluating this task. Reports accuracy (overall/macro).Votes: 0GitHub stars: 3
- Pediatric Brain Tumor Wsi EvalThis benchmark evaluates the ability of weakly supervised multiple instance learning models to classify pediatric brain tumors from whole-slide histopathology images. It probes fine-grained diagnostic discrimination across varying class granularities (2 to 7 classes) under conditions of class imbalance and limited data. Use when the user wants to benchmark on Pediatric brain tumor WSI dataset, or asks about evaluating this task. Reports Macro F1.Votes: 0GitHub stars: 3
- Orak EvalEvaluates LLM agents' long-horizon decision-making and gameplay capabilities across 12 diverse video games spanning six genres. It probes the effectiveness of different agentic strategies (zero-shot, reflection, planning, skill-management) and the impact of multimodal inputs (text vs. visual) on action inference. The benchmark also assesses generalization to unseen in-game scenarios, out-of-distribution games, and non-game tasks like math and web navigation. Use when the user wants to benchma...Votes: 0GitHub stars: 3
- Nlt Few Shot Classification EvalEvaluates whether few-shot classification methods can adapt to new tasks without using support set labels at test time. It probes the model's ability to cluster or classify query images based solely on support set images and learned representations. Use when the user wants to benchmark on Omniglot, miniImageNet, tieredImageNet, CUB, Meta-Dataset, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Nim4 Asr EvalEvaluates automatic speech recognition performance across diverse acoustic and linguistic domains, including English, Mandarin, dialects, code-switching, and in-car conversational scenarios. It measures transcription accuracy and hallucination rates to assess model robustness, latency, and customization capabilities. Use when the user wants to benchmark on LibriSpeech, VoxPopuli, MLS-English, AISHELL-1, AISHELL-2, AISHELL-2021-Eval, WeNetSpeech, SpeechIO, WeNetSpeech-Chuan, WeNetSpeech-Yue, K...Votes: 0GitHub stars: 3
- Morehopqa EvalEvaluates multi-step reasoning capabilities beyond simple information extraction in question answering. It probes models' ability to perform arithmetic, commonsense, and symbolic reasoning by extending standard multi-hop questions with additional reasoning layers. Use when the user wants to benchmark on MoreHopQA, or asks about evaluating this task. Reports EM.Votes: 0GitHub stars: 3
- Mobile Gpu Inference EvalThis benchmark evaluates the inference performance of a mobile GPU-accelerated neural network engine. It measures execution latency and peak memory consumption across different hardware platforms, precision formats (FP32 vs FP16), and batch sizes. It probes the engine's ability to leverage GPU parallelism and half-precision arithmetic for efficient edge deployment. Use when the user wants to benchmark on MNIST, Cifar-10, Style Transfer Dataset, or asks about evaluating this task. Reports exec...Votes: 0GitHub stars: 3
- Mobile EvalEvaluates an autonomous mobile device agent's ability to execute multi-step UI operations using visual perception. It probes task planning, self-reflection, and cross-application interaction under varying instruction complexity. Use when the user wants to benchmark on Mobile-Eval, or asks about evaluating this task. Reports Success (Su).Votes: 0GitHub stars: 3
- Minisuperb EvalEvaluates self-supervised speech models by measuring computational efficiency (forward MACs) and downstream task performance using a lightweight, offline feature extraction protocol. It probes the trade-off between model complexity and representation quality across speech tasks while enabling rapid early-stage model screening. Use when the user wants to benchmark on MiniSUPERB, or asks about evaluating this task. Reports forward MACs.Votes: 0GitHub stars: 3
- Mem Rec Ctr EvalEvaluates the click-through rate prediction capability of deep learning recommendation models under varying memory budgets and compression techniques. It probes how well alternative representation schemes (like Bloom filter encoding) maintain recommendation accuracy while drastically reducing embedding table sizes. Use when the user wants to benchmark on Avazu, Criteo-Kaggle, Criteo-Terabyte, or asks about evaluating this task. Reports ROC-AUC.Votes: 0GitHub stars: 3
- Masakhanews EvalThis benchmark evaluates the ability of language models and classical ML algorithms to classify news articles into predefined topics across 16 typologically diverse African languages. It probes multilingual representation quality, script handling, and few-shot/fine-tuning performance in low-resource settings. Use when the user wants to benchmark on MasakhaNEWS, or asks about evaluating this task. Reports weighted F1-score.Votes: 0GitHub stars: 3
- Magic Apt Detection EvalEvaluates the ability of a self-supervised graph representation learning model to detect Advanced Persistent Threats (APTs) in system audit logs. It probes multi-granularity anomaly detection (batched log-level and system entity-level) under a strict unsupervised setting where only benign data is available for training. Use when the user wants to benchmark on StreamSpot, Unicorn Wget, DARPA Engagement 3, or asks about evaluating this task. Reports Precision.Votes: 0GitHub stars: 3
- Longspeech EvalEvaluates long-form speech processing capabilities across transcription, translation, summarization, and higher-level reasoning tasks. It probes models' ability to maintain semantic consistency, track temporal progression, and extract structured information from ~10-minute audio segments. Use when the user wants to benchmark on LongSpeech, or asks about evaluating this task. Reports WER, BLEU-4, Numeric Accuracy, Strict Accuracy.Votes: 0GitHub stars: 3
- Llmstructbench EvalEvaluates large language models' ability to extract structured data from natural-language emails into valid JSON objects that adhere to a provided schema. It measures both syntactic validity (structural correctness) and semantic accuracy (correct value extraction) across varying levels of JSON nesting complexity. Use when the user wants to benchmark on LLMStructBench, or asks about evaluating this task. Reports DOC.Votes: 0GitHub stars: 3
- Lithology Classification EvalEvaluates a model's ability to perform multi-class sequence labeling on multi-channel well-log time series data for lithology classification. It tests the model's capacity to handle diverse geological settings, manage distribution shifts, and produce stratigraphically plausible predictions. Use when the user wants to benchmark on SEAM, Facies, FORCE, GeoLink, or asks about evaluating this task. Reports Weighted F1.Votes: 0GitHub stars: 3
- Ieee Cis Fraud Detection EvalThis evaluation probes a model's ability to detect financial fraud in a federated, privacy-preserving setting using quantum-enhanced neural networks. It measures classification performance on imbalanced transaction data while assessing robustness against simulated quantum hardware noise. Use when the user wants to benchmark on IEEE-CIS Fraud Detection, or asks about evaluating this task. Reports binary classification accuracy.Votes: 0GitHub stars: 3
- Hm Eqa EvalThis benchmark evaluates a robot's ability to autonomously explore an unseen indoor environment and answer multiple-choice questions requiring object identification, counting, spatial reasoning, and multi-goal navigation. It probes the agent's multimodal perception, iterative reasoning, and navigation efficiency in a dynamic, tool-invoking workflow. Use when the user wants to benchmark on HM-EQA, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Hermes Video Understanding EvalEvaluates a training-free KV cache management framework for real-time streaming and offline video understanding. It probes the model's ability to maintain temporal coherence and answer questions accurately under strict token/memory budgets, while measuring inference efficiency. Use when the user wants to benchmark on StreamingBench, OVO-Bench, RVS (Ego & Movie), MVBench, VideoMME, Egoschema, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3