Business & Operations
Operations, strategy, finance, sales, support, management, and planning
Browse business & operations skills
Showing 8,953–8,976 of 29,751 skills
Evaluates self-supervised speech models by measuring computational efficiency (forward MACs) and downstream task performance using a lightweight, offline feature extraction protocol. It probes the trade-off between model complexity and representation quality across speech tasks while enabling rapid early-stage model screening. Use when the user wants to benchmark on MiniSUPERB, or asks about evaluating this task. Reports forward MACs.
Evaluates the click-through rate prediction capability of deep learning recommendation models under varying memory budgets and compression techniques. It probes how well alternative representation schemes (like Bloom filter encoding) maintain recommendation accuracy while drastically reducing embedding table sizes. Use when the user wants to benchmark on Avazu, Criteo-Kaggle, Criteo-Terabyte, or asks about evaluating this task. Reports ROC-AUC.
This benchmark evaluates the ability of language models and classical ML algorithms to classify news articles into predefined topics across 16 typologically diverse African languages. It probes multilingual representation quality, script handling, and few-shot/fine-tuning performance in low-resource settings. Use when the user wants to benchmark on MasakhaNEWS, or asks about evaluating this task. Reports weighted F1-score.
Evaluates the ability of a self-supervised graph representation learning model to detect Advanced Persistent Threats (APTs) in system audit logs. It probes multi-granularity anomaly detection (batched log-level and system entity-level) under a strict unsupervised setting where only benign data is available for training. Use when the user wants to benchmark on StreamSpot, Unicorn Wget, DARPA Engagement 3, or asks about evaluating this task. Reports Precision.
Evaluates long-form speech processing capabilities across transcription, translation, summarization, and higher-level reasoning tasks. It probes models' ability to maintain semantic consistency, track temporal progression, and extract structured information from ~10-minute audio segments. Use when the user wants to benchmark on LongSpeech, or asks about evaluating this task. Reports WER, BLEU-4, Numeric Accuracy, Strict Accuracy.
Evaluates large language models' ability to extract structured data from natural-language emails into valid JSON objects that adhere to a provided schema. It measures both syntactic validity (structural correctness) and semantic accuracy (correct value extraction) across varying levels of JSON nesting complexity. Use when the user wants to benchmark on LLMStructBench, or asks about evaluating this task. Reports DOC.
Evaluates a model's ability to perform multi-class sequence labeling on multi-channel well-log time series data for lithology classification. It tests the model's capacity to handle diverse geological settings, manage distribution shifts, and produce stratigraphically plausible predictions. Use when the user wants to benchmark on SEAM, Facies, FORCE, GeoLink, or asks about evaluating this task. Reports Weighted F1.
This evaluation probes a model's ability to detect financial fraud in a federated, privacy-preserving setting using quantum-enhanced neural networks. It measures classification performance on imbalanced transaction data while assessing robustness against simulated quantum hardware noise. Use when the user wants to benchmark on IEEE-CIS Fraud Detection, or asks about evaluating this task. Reports binary classification accuracy.
This benchmark evaluates a robot's ability to autonomously explore an unseen indoor environment and answer multiple-choice questions requiring object identification, counting, spatial reasoning, and multi-goal navigation. It probes the agent's multimodal perception, iterative reasoning, and navigation efficiency in a dynamic, tool-invoking workflow. Use when the user wants to benchmark on HM-EQA, or asks about evaluating this task. Reports Accuracy.
Evaluates a training-free KV cache management framework for real-time streaming and offline video understanding. It probes the model's ability to maintain temporal coherence and answer questions accurately under strict token/memory budgets, while measuring inference efficiency. Use when the user wants to benchmark on StreamingBench, OVO-Bench, RVS (Ego & Movie), MVBench, VideoMME, Egoschema, or asks about evaluating this task. Reports accuracy.
Evaluates a graph neural network classifier's ability to determine whether credit card transactions require customer contact, aiming to optimize fraud detection workflows and reduce false positives. Use when the user wants to benchmark on Custom credit card transaction dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the scalability, heterogeneity handling, realism, and privacy overhead of the Flower federated learning framework across various datasets and device configurations. Use when the user wants to benchmark on Amazon Book Reviews, FEMNIST, RealWorld, CIFAR-10, FashionMNIST, or asks about evaluating this task. Reports training time.
Evaluates computational workflow anomaly detection by benchmarking models on detecting injected CPU and HDD performance anomalies in distributed workflow execution logs. It tests the ability of tabular, graph, and text-based methods to identify anomalous nodes within Directed Acyclic Graph (DAG) workflow executions. Use when the user wants to benchmark on Flow-Bench, or asks about evaluating this task. Reports ROC-AUC.
Evaluates the ability of a federated learning model to perform multi-step time-series forecasting of node-level traffic in optical networks, measuring how prediction accuracy degrades with longer history windows and forecasting horizons. Use when the user wants to benchmark on Optical network traffic dataset, or asks about evaluating this task. Reports R² score.
Evaluates the performance, efficiency, and robustness of Spiking Neural Networks (SNNs) deployed on edge hardware or simulated on conventional processors. It probes hardware-independent algorithmic complexity and system-level execution metrics under resource-constrained, latency-sensitive conditions. Use when the user has predictions and gold and needs to compute accuracy / mAP / MSE.
Evaluates the inference performance of heterogeneous edge AI platforms (CPU, GPU, NPU) across fundamental linear algebra primitives and diverse neural network models. It probes hardware efficiency in compute-bound versus memory-bound workloads, batch processing scalability, and quantization support. Use when the user wants to benchmark on Matrix Multiplication, Matrix-Vector Multiplication, Dot Product, MobileNetV2, LSTM, TinyLlama, or asks about evaluating this task. Reports Latency (ms).
Evaluates the economic value of hydropower reservoir operations based on sub-seasonal to seasonal precipitation forecasts under varying energy price differentials. It probes how forecast horizon and reservoir storage capacity interact with market prices to determine operational profitability. Use when the user wants to benchmark on 10-year reservoir timeseries, or asks about evaluating this task. Reports overall value of water.
Evaluates hardware-aware neural architecture search (NAS) methods on edge IoT tasks, measuring classification accuracy alongside hardware constraints like memory footprint, latency, and computational complexity (OPs) on a RISC-V IoT SoC. It benchmarks both mask-based and path-based differentiable NAS approaches across image classification, visual wake words, keyword spotting, and anomaly detection tasks. Use when the user wants to benchmark on CIFAR-10, MSCOCO 2014, Speech Commands v2, DCASE2...
This benchmark evaluates discrete audio tokenizers across speech, music, and general audio domains. It probes their ability to preserve acoustic fidelity during compression and decompression, as well as their effectiveness when used as inputs for downstream discriminative and generative audio tasks. Use when the user wants to benchmark on LibriSpeech test-clean, MUSDB, Audioset test-set, DASB Benchmark, or asks about evaluating this task. Reports SDR.
Evaluates the efficiency of display advertising bidding strategies by measuring how well they align advertiser payments with actual conversion attribution. It probes the model's ability to predict conversion probability and adjust bids dynamically to avoid overbidding after early clicks, ultimately maximizing advertiser utility under budget constraints. Use when the user wants to benchmark on Criteo Attribution Dataset, or asks about evaluating this task. Reports U_A.
Evaluates a graph convolutional neural network's ability to predict optimal next-node preferences for constrained shortest path problems with mandatory waypoints, to accelerate constraint programming solvers. Use when the user wants to benchmark on Maneuver benchmark, Exploration benchmark, or asks about evaluating this task. Reports Number of instances resolved with proof of optimality.
This benchmark evaluates multimodal complaint analysis by measuring a model's ability to jointly process multi-turn textual dialogues and accompanying images to classify fine-grained aspects and severity levels of customer grievances. It probes cross-modal alignment, multi-label classification, and robustness to class imbalance and subjective tone variations. Use when the user wants to benchmark on CIViL, or asks about evaluating this task. Reports macro F1-score.
This benchmark evaluates the inference accuracy and hardware performance of binary neural networks (BNNs) deployed on non-volatile memory crossbar architectures. It probes how hardware constraints like ADC resolution and first-layer input precision affect model accuracy, latency, energy efficiency, and chip area. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports accuracy.
Evaluates big data processing systems under diverse workload patterns (record insertion, statistics computation, iterative graph computation) to measure throughput, latency, and execution time across relational, text, and graph data types. Use when the user wants to benchmark on BigOP Log Monitoring & PageRank Workloads, or asks about evaluating this task. Reports throughput (ops/sec).