Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,836
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,625–6,648 of 20,836 skills

Kl ReductionA

Probes the statistical alignment between a candidate pretraining dataset and a target reference distribution (e.g., The Pile or Wikipedia/books). It quantifies how well the dataset's hashed n-gram frequencies match the desired language model pretraining distribution, serving as a proxy for downstream pretraining performance. Use when the user has predictions and gold and needs to compute KL reduction.

researchpythongo
0
3
Kitti Subt Vo Depth EvalA

Evaluates unsupervised monocular visual odometry and depth estimation methods on challenging driving and subterranean environments. Probes the model's ability to predict consistent 6-DoF ego-motion and recover accurate depth maps without ground-truth supervision. Use when the user wants to benchmark on KITTI, DARPA Subterranean Challenge, or asks about evaluating this task. Reports relative translation error ($t_{err}$), relative rotation error ($r_{err}$).

researchpythongo
0
3
Kitti Speed Estimation EvalA

Ego-vehicle longitudinal speed estimation from monocular video sequences. It probes a model's ability to infer real-world velocity by combining optical flow magnitude and monocular depth/disparity cues over time. Use when the user wants to benchmark on KITTI, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Kitti Optical Flow EvalA

Evaluates the accuracy of unsupervised optical flow estimation methods on standard driving scenes. It measures the pixel-wise displacement error between predicted and ground-truth flow fields to quantify estimation quality. Use when the user wants to benchmark on KITTI2012, or asks about evaluating this task. Reports EPE.

researchpython
0
3
Kitti Lidar Flow EvalA

This evaluation probes a model's ability to estimate dense optical flow directly from sparse, noisy LiDAR range scans without using RGB images. It measures prediction accuracy against real-world ground truth flow maps and evaluates robustness to occlusions and foreground/background motion. Use when the user wants to benchmark on KITTI Tracking & Flow 2015, or asks about evaluating this task. Reports EPE (End-Point-Error).

researchpython
0
3
Kitti Fc EvalA

Evaluates the robustness of optical flow estimation models when subjected to various digital, illumination, weather, noise, and blur corruptions. It measures both absolute performance degradation and relative robustness compared to clean data across in-domain and out-of-domain training settings. Use when the user wants to benchmark on KITTI-FC, or asks about evaluating this task. Reports EPE.

researchpythongit
0
3
Kitti Eigen Depth EvalA

Evaluates the accuracy of self-supervised monocular depth estimation models on urban driving scenes. It probes the model's ability to predict per-pixel depth from a single image or video sequence, handling occlusions, moving objects, and scale ambiguity. Use when the user wants to benchmark on KITTI 2015 (Eigen split), Make3D, or asks about evaluating this task. Reports Abs Rel, δ < 1.25.

researchpythongo
0
3
Kitti Depth Flow Pose EvalA

Evaluates a model's ability to jointly estimate monocular depth, optical flow, and camera ego-motion from consecutive video frames in driving scenes. It probes geometric consistency, motion handling, and self-supervised learning robustness on standard autonomous driving benchmarks. Use when the user wants to benchmark on KITTI Raw, KITTI Flow 2012, KITTI Flow 2015, KITTI Odometry, KITTI Eigen Split, or asks about evaluating this task. Reports EPE.

researchpythongo
0
3
Kitti Depth EvalA

Evaluates self-supervised monocular depth estimation models on outdoor driving scenes, measuring both the geometric accuracy of predicted depth maps and the reliability of associated uncertainty estimates. Use when the user wants to benchmark on KITTI, or asks about evaluating this task. Reports Abs Rel.

researchpython
0
3
Kitsune Iot Nids EvalA

Evaluates online unsupervised anomaly detection systems for network intrusion detection on IoT surveillance and network traffic. It measures how well models distinguish between normal traffic and various attack types (e.g., DoS, MITM, malware) using streaming packet features. Use when the user wants to benchmark on Kitsune IoT Network Datasets, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Kits21 Segmentation EvalA

This benchmark evaluates the ability of deep learning models to perform multi-organ and multi-lesion semantic segmentation on 3D medical imaging data. It specifically probes a model's capacity to accurately delineate kidneys, renal tumors, and renal cysts from corticomedullary-phase CT scans, testing both volumetric overlap and boundary precision. Use when the user wants to benchmark on KiTS21, or asks about evaluating this task. Reports dice.

researchpythontesting
0
3
Kitchen Sink Anomaly Detection EvalA

Evaluates the robustness and sensitivity of various jet substructure feature sets (Energy Flow Polynomials, subjettiness, and their combination) for model-agnostic resonant anomaly detection in high-energy physics dijet events. It compares performance across Ideal Anomaly Detection (IAD) and CWoLa hunting setups using multiple Beyond Standard Model signal topologies. Use when the user wants to benchmark on LHCO & BSM dijet signals, or asks about evaluating this task. Reports max(SIC), sigma_0...

researchpythonperformance
0
3
Kinect Action Recognition EvalA

Evaluates the robustness of Kinect-based action recognition algorithms across single-view and cross-view scenarios. Probes how well models handle viewpoint variation, motion variability, and different sensor modalities (depth, skeleton, RGB-D) on standardized benchmarks. Use when the user wants to benchmark on MSRAction3D Dataset, 3D Action Pairs Dataset, Cornell Activity Dataset (CAD-60), UWA3D Single View Dataset, UWA3D Multiview Dataset, or asks about evaluating this task. Reports average ...

researchpythongo
0
3
Kimi K1.5 Benchmark EvalA

Evaluates multimodal reasoning, coding, and instruction-following capabilities across text, code, and vision tasks using a standardized suite of academic benchmarks. Use when the user wants to benchmark on MMLU, IF-Eval, CLUEWSC, C-EVAL, HumanEval-Mul, LiveCodeBench, Codeforces, AIME 2024, MATH-500, MMMU, MATH-Vision, MathVista, or asks about evaluating this task. Reports exact-match accuracy (EM).

researchpythongo
0
3
Kimi Audio EvalA

Evaluates an audio foundation model's capabilities across automatic speech recognition, general audio understanding, audio-to-text conversational reasoning, and end-to-end speech conversation. Use when the user wants to benchmark on LibriSpeech, FLEURS, AISHELL-1, AISHELL-2, WenetSpeech, Kimi-ASR Internal Testset, MMAU, ClothoAQA, VocalSound, Nonspeech7k, MELD, TUT2017, CochlScene, OpenAudioBench, VoiceBench, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythongo
0
3
Kilt EvalA

Evaluates a model's ability to perform knowledge-intensive language tasks by jointly assessing output generation accuracy and evidence retrieval from a fixed Wikipedia snapshot. It measures how well models can produce correct answers while providing verifiable text-span provenance to justify predictions. Use when the user wants to benchmark on KILT, or asks about evaluating this task. Reports KILT scores.

researchpythongo
0
3
Kidney Histopathology EvalA

Evaluates histopathology foundation models on kidney-specific downstream tasks, including tile-level morphological classification, molecular information estimation, and slide-level diagnostic/prognostic inference across diverse staining protocols (H&E, PAS, PASM, IHC). Use when the user wants to benchmark on Kidney Digital Pathology Benchmark, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).

researchpythongit
0
3
Kgquiz EvalA

Evaluates large language models' ability to store, retrieve, and reason over factual knowledge encoded in parametric memory across five progressively complex tasks. It probes basic fact verification, multiple-choice discrimination, open-ended entity generation, multi-hop factual editing, and comprehensive entity description generation. It measures how well models generalize encoded knowledge across commonsense, encyclopedic, and biomedical domains under increasing reasoning complexity. Use wh...

researchpythongo
0
3
Kgqa4mat EvalA

Evaluates a model's ability to translate natural language questions into correct graph database queries (Cypher or SPARQL) for knowledge graph question answering. It probes the model's capacity for logical reasoning, schema understanding, and formal language generation in both a domain-specific materials science setting and a general-domain multilingual benchmark. Use when the user wants to benchmark on KGQA4MAT, QALD-9, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Kgqa EvalA

This benchmark evaluates the ability of conversational AI models and traditional knowledge graph question-answering systems to accurately answer natural language questions over structured knowledge graphs. It probes factual grounding, recall on exhaustive lists, robustness to linguistic variations, and determinism across general and academic domains. Use when the user wants to benchmark on QALD-9, YAGO, DBLP, MAG, or asks about evaluating this task. Reports Micro F1 score.

researchpythongo
0
3
Kg Benchmark EvalA

Evaluates the query execution performance and scalability of various knowledge graph systems across e-commerce, academic, and digital twin domains. It measures how efficiently triple stores and property graphs handle subsumption, recursive queries, and large-scale RDF datasets under realistic workload conditions. Use when the user wants to benchmark on BSBM, LUBM, DTBM, or asks about evaluating this task. Reports average response time (seconds).

researchpythongit
0
3
Kfineval Pilot EvalA

Evaluates Korean financial language models across three core capabilities: factual knowledge recall, multi-step legal/financial reasoning, and safety alignment against adversarial toxic prompts. It probes domain-specific understanding, procedural reasoning, and robustness to financial fraud or privacy-violating queries. Use when the user wants to benchmark on KFinEval-Pilot, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Keyword Spotting Efficiency EvalA

Evaluates the energy efficiency and inference speed of various hardware platforms (CPU, GPU, neuromorphic chips) running a keyword spotting neural network on audio data. It measures how power consumption and latency scale with network size and batch configuration. Use when the user wants to benchmark on Keyword Spotting Dataset, or asks about evaluating this task. Reports energy cost per inference (J).

researchpythongit
0
3
Keyinst EvalA

This evaluation probes a model's ability to formulate correct SQL queries from natural language questions, specifically focusing on capturing structural semantics like GROUP BY, HAVING, ORDER BY, and set operations. It measures how well prompt engineering techniques or fine-tuning improve SQL generation accuracy across different database schemas and question complexities. Use when the user wants to benchmark on StrucQL, Spider, Bird, or asks about evaluating this task. Reports execution accur...

researchpythongo
0
3