All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,385
- 892
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,841–9,864 of 21,385 skills
- Adiabatic Quantum Benchmark EvalEvaluates the performance of adiabatic quantum optimization on complex network analysis tasks. It benchmarks quantum annealing against classical methods on Chimera Ising spin glass instances, independent set problems, planted-solution instances, and community detection. Use when the user wants to benchmark on Chimera Ising spin glass instances, Independent set problems, Planted-solution instances, Community detection, or asks about evaluating this task. Reports D-Wave run-time estimation.Votes: 0GitHub stars: 3
- Ader Sr EvalEvaluates continual learning performance for session-based recommendation by measuring how well a model maintains prediction accuracy on historical items while adapting to new sessions over time. It probes stability-plasticity trade-offs by averaging recommendation quality across multiple sequential update cycles. Use when the user wants to benchmark on DIGINETICA, YOOCHOOSE, or asks about evaluating this task. Reports Recall@k.Votes: 0GitHub stars: 3
- Adept Prosody Clone EvalEvaluates a zero-shot multispeaker TTS model's ability to clone both speaker voice and fine-grained prosody from untranscribed reference audio. It measures intelligibility, spectral/prosodic fidelity, and perceptual similarity against human references. Use when the user wants to benchmark on ADEPT, or asks about evaluating this task. Reports Phone Error Rate (PER).Votes: 0GitHub stars: 3
- Ade20k Scene Parse EvalEvaluates a model's ability to perform dense pixel-wise semantic segmentation across 150 common scene categories, including both discrete objects and amorphous 'stuff' classes. It probes fine-grained scene understanding and the model's capacity to handle class imbalance and varying object scales. Use when the user wants to benchmark on SceneParse150, or asks about evaluating this task. Reports Mean IoU.Votes: 0GitHub stars: 3
- Adcraft EvalEvaluates reinforcement learning agents' ability to optimize bidding strategies and budget allocation in a non-stationary, stochastic Search Engine Marketing (SEM) simulation. It probes how well policies handle sparse feedback, shifting reward landscapes, and long-term profitability constraints over a simulated campaign. Use when the user wants to benchmark on AdCraft Environment, or asks about evaluating this task. Reports NCP.Votes: 0GitHub stars: 3
- Adbench EvalEvaluates tabular anomaly detection models on their ability to identify outliers in medium- and high-dimensional datasets by measuring ranking quality and precision-recall trade-offs under a standardized semi-supervised protocol. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports ROC-AUC.Votes: 0GitHub stars: 3
- Adasum Scaling EvalEvaluates the algorithmic and system efficiency of the Adasum distributed gradient combiner compared to naive gradient averaging across different hardware interconnects and model scales. It probes the ability of synchronous SGD to scale to large effective batch sizes while maintaining convergence accuracy and reducing time-to-accuracy. Use when the user wants to benchmark on ImageNet, SQuAD 1.1, MNIST, or asks about evaluating this task. Reports epochs_to_target_accuracy.Votes: 0GitHub stars: 3
- Adaptmmbench EvalEvaluates Vision-Language Models' ability to dynamically select between text-only and tool-augmented reasoning modes, and assesses the quality, efficiency, and final accuracy of their reasoning processes across multimodal domains. Use when the user wants to benchmark on AdaptMMBench, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Adaptive Thinking EvalEvaluates a model's ability to dynamically adjust its reasoning depth based on problem difficulty, balancing computational efficiency against task accuracy. It also measures safety alignment by assessing the model's harmless response rate on adversarial or harmful prompts. Use when the user wants to benchmark on MATH500, AIME2024, AMC2023, Olympiad Bench, GSM8K, BeaverTails, HarmfulQA, or asks about evaluating this task. Reports pass@1 accuracy.Votes: 0GitHub stars: 3
- Adaptive Sgd EvalEvaluates the training efficiency and convergence accuracy of sparse deep learning optimizers on large-scale, high-dimensional multi-class classification tasks with extreme label sparsity. Use when the user wants to benchmark on Amazon-670k, Delicious-200k, or asks about evaluating this task. Reports time-to-accuracy.Votes: 0GitHub stars: 3
- Adaptive Query Routing EvalEvaluates retrieval and answer generation methods across structured financial, legal, and medical documents. It probes how well different architectures handle varying query complexities, cross-references, and domain-specific structural requirements. Use when the user wants to benchmark on Controlled Multi-Domain Corpus, FinanceBench, or asks about evaluating this task. Reports Quality.Votes: 0GitHub stars: 3
- Adaptive Llm Testing EvalThis evaluation probes the effectiveness of diversity-based adaptive test selection strategies for black-box LLM applications. It measures how quickly and reliably different prioritization methods detect failures in prompt templates compared to random baselines, while also assessing the diversity of generated outputs. Use when the user wants to benchmark on BBH & P3 Prompt Templates, or asks about evaluating this task. Reports APFD.Votes: 0GitHub stars: 3
- Adapter Fedllm Privacy EvalEvaluates the privacy vulnerability of adapter-based federated large language models against gradient inversion attacks. It measures how accurately an adversary can reconstruct private training text from shared adapter gradients under varying batch sizes, model architectures, and defensive mechanisms. Use when the user wants to benchmark on CoLA, SST, Rotten Tomatoes, or asks about evaluating this task. Reports ROUGE-1.Votes: 0GitHub stars: 3
- Adamerging EvalEvaluates multi-task model merging methods on image classification tasks by measuring average accuracy across multiple datasets, generalization to unseen tasks, and robustness to image corruptions. Use when the user wants to benchmark on Image Classification Bench (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD), or asks about evaluating this task. Reports Avg Acc.Votes: 0GitHub stars: 3
- Adadecode EvalEvaluates the inference speedup and output consistency of adaptive layer parallelism for LLM decoding compared to standard autoregressive generation. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports speedup.Votes: 0GitHub stars: 3
- Adacompress EvalEvaluates a reinforcement learning-based adaptive JPEG compression framework for cloud computer vision services. It measures how effectively the system balances image file size reduction against the accuracy degradation of downstream black-box vision models, while accounting for end-to-end latency overhead compared to standard JPEG baselines. Use when the user wants to benchmark on ImageNet, DNIM, or asks about evaluating this task. Reports relative top-5 accuracy.Votes: 0GitHub stars: 3
- Ad4ad EvalEvaluates visual anomaly detection models for autonomous driving by measuring their ability to detect and precisely localize defects or hazards in road scenes. It probes the trade-off between detection accuracy, pixel-level localization precision, and computational efficiency for onboard deployment. Use when the user wants to benchmark on AD4AD (AnoVox), or asks about evaluating this task. Reports P-AP.Votes: 0GitHub stars: 3
- Ad2 Bench EvalEvaluates multimodal large language models on autonomous driving tasks under adverse weather and complex scenes. It probes base and advanced visual perception, relational understanding, event reasoning, and the coherence of hierarchical chain-of-thought reasoning. Use when the user wants to benchmark on AD^2-Bench, or asks about evaluating this task. Reports Avg-S.Votes: 0GitHub stars: 3
- Ad Personalization Ips EvalEvaluates the predictive accuracy and decision-making value of ad targeting policies using different information sets (contextual, geographical, behavioral). It specifically tests whether geographical and behavioral data act as complements or substitutes in improving click-through rates, accounting for user exposure history. Use when the user wants to benchmark on Real-world Ad Impression Dataset, or asks about evaluating this task. Reports IPS policy value.Votes: 0GitHub stars: 3
- Actormind EvalThis benchmark evaluates a model's ability to perform speech role-playing by generating persona-consistent, emotionally grounded audio responses. It specifically probes the model's capacity for accurate voice impersonation, precise content delivery, and alignment with target emotional prosody in a conversational context. Use when the user wants to benchmark on ActorMindBench, or asks about evaluating this task. Reports RP-MOS.Votes: 0GitHub stars: 3
- Activitynet Comp EvalEvaluates fine-grained temporal and compositional alignment in video-text models by testing their ability to distinguish between videos and captions that contain subtle structural disruptions. It probes sensitivity to temporal reordering, action word replacement, and segment-level misalignment, as well as the model's robustness to combined disruptions. Use when the user wants to benchmark on ActivityNet-Comp, YouCook2-Comp, or asks about evaluating this task. Reports binary classification acc...Votes: 0GitHub stars: 3
- Activity Recognition EvalEvaluates a model's ability to recognize human activities in real-time by simultaneously learning from skeletal pose data and object attributes. It probes the integration of multi-modal cues (color, shape, distance, or object probabilities) for accurate and efficient activity classification in robotics scenarios. Use when the user wants to benchmark on Cornell Activity Dataset (CAD-60), MSR Daily Activity 3D Dataset, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Activity Chain Synthesis EvalEvaluates a generative model's ability to synthesize realistic human mobility patterns by producing activity chains and location trajectories conditioned on socio-demographic and household attributes. It probes the model's capacity to capture temporal dynamics, activity type distributions, transition probabilities, and household interdependencies at both the sequence and system levels. Use when the user wants to benchmark on Household Travel Survey (HTS) / NHTS, or asks about evaluating this ...Votes: 0GitHub stars: 3
- Active Nerf EvalEvaluates the accuracy of 3D geometry reconstruction from multi-view images using active pattern projection. It measures how closely the predicted point cloud matches the ground truth geometry and assesses robustness under varying view counts and camera-projector baselines. Use when the user wants to benchmark on NeRF derivative (synthetic), Real-world capture (RealSense D415), or asks about evaluating this task. Reports Chamfer Distance (mm).Votes: 0GitHub stars: 3