Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,769–6,792 of 20,827 skills
Evaluates an AI system's ability to assess scientific research ideas across classification, selection, ranking, and comparison tasks. It probes knowledge-grounded reasoning and multi-perspective evaluation aligned with human expert judgments and conference acceptance standards. Use when the user wants to benchmark on D_point, D_group, D_pair, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to predict traffic accident injury severity (fatal, serious, slight) from tabular accident data. It specifically probes performance on highly imbalanced multi-class classification, focusing on minority-class accuracy and the impact of data imputation and resampling techniques. Use when the user wants to benchmark on UK DfT Traffic Accident Dataset (2005–2019), or asks about evaluating this task. Reports Overall classification accuracy.
Evaluates LLMs on complex, multi-step reasoning and agentic search tasks, including single-hop and multi-hop question answering as well as deep research benchmarks requiring web search and synthesis. Use when the user wants to benchmark on NQ, TQA, PopQA, HQA, 2Wiki, MSQ, Bamb, BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.
Evaluates large language models' ability to follow complex, multi-constraint instructions by decomposing them into granular criteria (Content, Linguistic, Style, Format, Number) and measuring adherence. It probes fine-grained instruction following rather than holistic response quality. Use when the user wants to benchmark on InFoBench, or asks about evaluating this task. Reports DRFR.
This protocol evaluates the accuracy and efficiency of influence function approximation methods (DataInf, LiSSA, Hessian-free) in matching exact influence values, detecting mislabeled training data, and identifying training points that most impact a test instance's loss across text and image generation tasks. Use when the user wants to benchmark on GLUE (binary classification subsets), Custom Text Generation Datasets, Custom Image Generation Datasets, or asks about evaluating this task. Repor...
Evaluates the conversational and foundational capabilities of LLMs fine-tuned on instruction datasets, comparing them against proprietary and open-source baselines across multiple standard benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard, MT-Bench, MATH, GSM-8K, HumanEval, MBPP, MMLU, LUC-EVAL, or asks about evaluating this task. Reports Overall*.
Evaluates continual learning methods on a procedurally generated benchmark of 500 shape classification tasks. It probes a model's ability to learn incrementally over a long horizon without catastrophic forgetting, while maintaining open-set recognition and one-shot generalization capabilities on unseen shapes. Use when the user wants to benchmark on Infinite dSprites (idSprites), or asks about evaluating this task. Reports average test accuracy.
Probes how different CPU microarchitectures (Haswell, Broadwell, Skylake) and cache hierarchies affect the inference latency and throughput of production-scale DNN recommendation models under varying batch sizes and co-location scenarios. Use when the user has predictions and gold and needs to compute inference latency.
Evaluates the inference performance of four deep learning frameworks (TensorRT, ONNX Runtime, OpenVINO, TensorFlow XLA) across four CNN architectures on GPU hardware. It probes how configuration settings, graph optimizations, and batch sizes impact inference speed and resource utilization, including co-localized model ensembles. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports speed.
This benchmark evaluates large language models' ability to perform informal mathematical reasoning on Olympiad-level inequality problems. It probes step-wise deductive chain integrity by decomposing proofs into bound estimation and relation prediction subtasks, requiring models to generate logically sound derivations rather than just final answers. Use when the user wants to benchmark on IneqMath, or asks about evaluating this task. Reports LLM-as-judge accuracy.
This benchmark evaluates 6D object pose estimation, detection, and segmentation capabilities in realistic industrial environments. It specifically probes a model's ability to handle challenging conditions such as heavy occlusion, background clutter, reflective surfaces, textureless materials, and object symmetry. Use when the user wants to benchmark on IndustryShapes Classic, IndustryShapes Extended, or asks about evaluating this task. Reports Average Recall (AR).
This benchmark evaluates an AI system's ability to detect and localize industrial safety hazards (e.g., oil spills, chemical stains) in complex factory environments. It probes the model's spatial grounding precision and its capacity to generalize from synthetic or limited real-world data to unseen operational sites. Use when the user wants to benchmark on Public Spill Data, Proprietary Factory Data, Synthetic Spill (SynSpill) Dataset, or asks about evaluating this task. Reports mean hit rate@...
Evaluates a model's ability to predict missing links in knowledge graphs using only topological path information, without relying on entity embeddings. It tests inductive generalization by training on one graph and testing on a disjoint graph with unseen entities. Use when the user wants to benchmark on WN18RR, FB15K-237, NELL-995 (inductive versions v1-v4), or asks about evaluating this task. Reports Hits@1.
Evaluates parameter-efficient fine-tuning methods on natural language understanding and generation tasks, measuring how well they approximate full fine-tuning performance while using significantly fewer trainable parameters. Use when the user wants to benchmark on MNLI, SST2, WebNLG-challenge, CoQA, or asks about evaluating this task. Reports Accuracy.
Probes cross-lingual visual question answering on document images containing tables. It tests a model's ability to perform factual lookup, numerical comparison, aggregation, and structural reasoning across Bahasa Indonesia, English, Hindi, and Arabic. Use when the user wants to benchmark on IndoTabVQA, or asks about evaluating this task. Reports In-Match Accuracy.
Evaluates 3D object detection and BEV perception capabilities on indoor robotic platforms using LiDAR point clouds. It probes a model's ability to classify indoor objects and localize them with 3D bounding boxes, specifically highlighting the sim-to-real transfer gap in controlled indoor environments. Use when the user wants to benchmark on INDOOR-LiDAR, or asks about evaluating this task. Reports Mean IoU.
This benchmark evaluates Indonesian natural language understanding across 12 diverse tasks, including single-sentence classification, sentence-pair classification, and sequence labeling/tagging. It probes a model's ability to handle sentiment analysis, aspect-based sentiment, textual entailment, part-of-speech tagging, named entity recognition, keyphrase extraction, and question answering in Indonesian. Use when the user wants to benchmark on IndoNLU, or asks about evaluating this task. Repor...
Evaluates sequence labeling performance on Indonesian text by assigning part-of-speech tags to tokens. It probes morphological feature extraction, contextual understanding, and robustness to annotation inconsistencies and rare lexical categories. Use when the user wants to benchmark on IDN Tagged Corpus, or asks about evaluating this task. Reports F1.
Evaluates Indonesian NLP capabilities across morpho-syntax, semantics, and discourse. It probes token-level labeling (POS, NER), syntactic structure (dependency parsing), text classification (sentiment), generation (summarization), and discourse coherence (next tweet prediction, tweet ordering). Use when the user wants to benchmark on INDOLEM, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates the ability of LLMs to autoformalize natural language mathematical problems into correct Lean 4 theorems and subsequently prove them. It probes semantic equivalence, syntactic structural similarity, and automated theorem proving success rates on Olympiad-level geometry and algebra problems. Use when the user wants to benchmark on IndiMathBench, or asks about evaluating this task. Reports BEq.
Evaluates zero-shot cross-lingual transfer capabilities of language models on low-resource Indic languages by testing performance on nine natural language understanding tasks after training exclusively on English data. Use when the user wants to benchmark on IndicXTREME, or asks about evaluating this task. Reports task-specific accuracy/F1.
Evaluates multilingual natural language inference capabilities across 11 Indic languages, probing both intra-lingual reasoning and cross-lingual transfer performance of pre-trained language models. Use when the user wants to benchmark on IndicXNLI, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multilingual machine translation quality across 22 scheduled Indian languages and English. It probes a model's ability to handle diverse domains (news, web, conversation, legal, etc.) and both Indic-to-English and English-to-Indic translation directions in an n-way parallel setting. Use when the user wants to benchmark on IN22, FLORES-200, NTREX, WMT (2014, 2019, 2020), WAT (2020, 2021), UFAL, or asks about evaluating this task. Reports chrF++.
Evaluates large language models on multi-task language understanding across nine major Indic languages. It probes capabilities in reading comprehension, reasoning, and knowledge retention by adapting the English MMLU-Pro benchmark through machine translation and rigorous quality assurance. Use when the user wants to benchmark on IndicMMLU-Pro, or asks about evaluating this task. Reports Accuracy.