Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,649–6,672 of 20,836 skills
Evaluates whether a reduced subset of benchmark tests preserves the relative performance ranking of software variants compared to a full test suite. It probes the fidelity of black-box performance comparisons when benchmark execution costs are minimized. Use when the user has predictions and gold and needs to compute Kendall target.
Evaluates machine reading comprehension and question answering on low-resource Kiswahili. It tests a model's ability to read short stories and accurately extract or generate answers to posed questions. The protocol uses a held-out test split to measure token-level overlap and exact string matching against ground truth answers. Use when the user wants to benchmark on KenSwQuAD, or asks about evaluating this task. Reports F1 score.
Evaluates supervised machine learning classifiers for anomaly detection in IoT network traffic, specifically probing their ability to identify intrusion attack categories under severe class imbalance. Use when the user wants to benchmark on KDD Cup 1999, or asks about evaluating this task. Reports Accuracy.
Evaluates network intrusion detection models by measuring per-class detection rates and false positive rates across normal and attack traffic categories. Probes the classifier's ability to balance sensitivity to rare attack types while minimizing misclassification of benign traffic. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate.
Evaluates an intrusion detection system's ability to classify network traffic into normal and specific attack categories (probe, dos, u2r, r2l) using genetic algorithm-optimized feature selection and rule generation. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate (DR).
Evaluates a network intrusion detection system's ability to classify TCP/IP connections as either normal or one of several attack types based on 41 network features. It measures how well the model discriminates between benign traffic and specific intrusion categories such as DoS, Probe, R2L, and U2R. Use when the user wants to benchmark on KDD-Cup 99, or asks about evaluating this task. Reports accuracy.
Evaluates the capability of an unsupervised cellular automata-based framework to detect network intrusions and anomalous traffic patterns. It probes the model's ability to learn spatial-temporal dependencies from raw network connection records and distinguish between normal and malicious activities. Use when the user wants to benchmark on KDD Cup, or asks about evaluating this task. Reports Intrusion Detection Vs Positive Rate.
Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on KDD CUP 99, or asks about evaluating this task. Reports accuracy.
Evaluates instruction-tuned LLMs on e-commerce tasks across two Amazon KDD Cup'24 tracks. It measures model performance on development and official test sets to assess retrieval, ranking, and generation capabilities in a commercial setting. Use when the user wants to benchmark on Amazon KDD Cup'24, or asks about evaluating this task. Reports scores.
Evaluates a model's ability to answer complex questions over knowledge graphs by retrieving relevant subgraphs and generating correct answer entities. It probes multi-hop reasoning capabilities and robustness to missing edges in incomplete knowledge bases. Use when the user wants to benchmark on ComplexWebQuestions, WebQuestionsSP, WebQuestions, GrailQA, or asks about evaluating this task. Reports Hit@1.
Evaluates the ability of distantly supervised relation extraction and knowledge base validation systems to correctly predict and rank triples in web-scale knowledge graphs. It probes how well global graph structure and confidence scoring can refine noisy extractions and reduce logical inconsistencies. Use when the user wants to benchmark on NYT-FB, CC-DBP, NELL-165, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to learn first-order logic rules for knowledge base completion and object classification. It probes rule generation efficiency, scalability to longer rules, and few-shot generalization on relational data. Use when the user wants to benchmark on Even-and-Successor (ES), FB15K-237, WN18, Visual Genome (via GQA), or asks about evaluating this task. Reports MRR.
Evaluates multilingual sentiment classification models on Kazakh customer reviews, probing their ability to handle code-switching, mixed scripts, and imbalanced class distributions across polarity and numerical score prediction tasks. Use when the user wants to benchmark on KazSAnDRA, or asks about evaluating this task. Reports macro-F1.
Evaluates the acoustic quality, intelligibility, and diacritic sensitivity of a Kashmiri text-to-speech system. It probes the model's ability to accurately map Perso-Arabic script with explicit diacritics to natural-sounding speech and maintain spectral fidelity under low-resource conditions. Use when the user wants to benchmark on Curated Kashmiri Corpus, or asks about evaluating this task. Reports MCD, MOS.
Evaluates a model's ability to manipulate specific semantic attributes (e.g., technique, skill level) in human motion data while preserving untargeted attributes and anatomical accuracy. It probes latent space disentanglement and the capacity for controlled, attribute-level motion editing. Use when the user wants to benchmark on Kyokushin karate dataset, or asks about evaluating this task. Reports linear separability.
Evaluates multilingual vision-language reasoning by testing models on multiple-choice questions about images entirely in their native language. It probes cultural and linguistic authenticity, assessing how well models handle complex multimodal reasoning without relying on English translations. Use when the user wants to benchmark on Kaleidoscope, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to manipulate and apply stored knowledge across logical reasoning, reading comprehension, and natural language understanding tasks. It specifically probes the 'known & incorrect' phenomenon where models possess relevant facts but fail to apply them correctly during inference. Use when the user wants to benchmark on AbsR, Commonsense (Common), Big Bench Hard (BBH), RACE-H, RACE-M, MMLU, ARC-c, ARC-e, or asks about evaluating this task. Reports accuracy.
This benchmark probes an LLM's ability to understand and generate culturally appropriate responses for Filipino contexts. It evaluates whether models can align with the lived experiences, values, and preferred strategies of action of average native Filipino speakers across nuanced socio-cultural scenarios. Use when the user wants to benchmark on Kalahi, or asks about evaluating this task. Reports MC1.
Evaluates a provenance-based intrusion detection system's ability to identify anomalous system behavior and reconstruct attack footprints from whole-system kernel-level logs. It probes the model's capacity to distinguish between benign and malicious activity in temporal windows without relying on attack signatures. Use when the user wants to benchmark on Manzoor et al., DARPA-E3-THEIA, DARPA-E3-CADETS, DARPA-E3-ClearScope, DARPA-E5-THEIA, DARPA-E5-CADETS, DARPA-E5-ClearScope, DARPA-OpTC, or a...
Evaluates the zero-shot and few-shot generalization capability of text-to-SQL parsers on realistic, industrial-style database schemas with obscure column names and unrestricted natural language questions. It probes the model's ability to perform schema linking, constraint parsing, and SQL generation without extensive domain-specific training data. Use when the user wants to benchmark on KaggleDBQA, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates models' capabilities in knowledge abstraction, concretization, and completion within entity-concept knowledge graphs. It specifically probes multi-hop reasoning through hierarchical relations and cross-view knowledge transfer between entities and concepts. Use when the user wants to benchmark on KACC, or asks about evaluating this task. Reports Hits@10.
Evaluates automated exoplanet vetting pipelines on K2 transit lightcurves by comparing their planet candidate versus false positive dispositions against established ground truth from the NASA Exoplanet Archive. It probes the ability of tools to correctly identify true transiting planets while filtering out astrophysical false positives like eclipsing binaries and blended stars. Use when the user wants to benchmark on K2 Planet Candidate Catalog (DAVE Benchmark), or asks about evaluating this ...
Evaluates a model's ability to predict the number of k-barriers for intrusion detection in wireless sensor networks deployed over circular regions, based on geometric and deployment parameters. Use when the user wants to benchmark on WSN k-barrier simulation dataset, or asks about evaluating this task. Reports RMSE.
Probes the factual and logical reliability of LLM-based judges and reward models by testing their ability to distinguish between objectively correct responses and subtly flawed ones across knowledge, reasoning, math, and coding domains. Use when the user wants to benchmark on JudgeBench, or asks about evaluating this task. Reports accuracy.