Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,985–7,008 of 20,835 skills
Evaluates the answer quality of Retrieval-Augmented Generation (RAG) systems across specialized domains. It measures how well generated responses address queries in terms of comprehensiveness, empowerment, diversity, and overall performance using pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on UltraDomain, or asks about evaluating this task. Reports win rate.
Evaluates histopathology image retrieval and classification performance using high-order texture features (Gram barcodes) extracted from CNN layers. It probes the model's ability to capture tissue texture patterns for accurate image matching and class prediction. Use when the user wants to benchmark on KimiaPath24, CRC, EMC, or asks about evaluating this task. Reports η_total, Accuracy.
Evaluates multimodal agents' ability to reason over personalized, device-scale file systems. It probes long-horizon cross-file retrieval, multimodal perception, and evidence-grounded factual retention under strict profile-isolation constraints. Use when the user wants to benchmark on HippoCamp, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multimodal physical reasoning and advanced problem-solving capabilities on international and regional physics Olympiad exams. It probes a model's ability to interpret complex diagrams, data, and text, perform multi-step logical derivations, and produce accurate solutions under strict, official scoring rubrics. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports exam score.
Evaluates multimodal large language models on authentic high school physics Olympiad problems, probing their ability to perform step-level physical reasoning, interpret complex diagrams and data plots, and solve problems across diverse physics subfields under Olympiad-level difficulty. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports Mean Normalized Score (MNS).
Evaluates multilingual historical text systems on person-place relation extraction, requiring temporal and geographical reasoning to classify relations as 'at' or 'isAt' with nuanced evidence levels (true, probable, false). It probes both extraction accuracy and reasoning quality in noisy, sparse corpora while also measuring computational efficiency. Use when the user wants to benchmark on HIPE-2026, or asks about evaluating this task. Reports macro-averaged Recall.
This benchmark probes a model's ability to detect adverse events (total hip replacement dislocation) from unstructured, free-text clinical narratives. It evaluates whether NLP models can correctly classify medical notes into dislocation status categories, handling complex negation, long-range dependencies, and multi-site anatomical references. Use when the user wants to benchmark on Radiology Notes, Telephone Notes, or asks about evaluating this task. Reports Kappa.
Evaluates multimodal large language models on vision-language tasks in Hindi and Telugu, measuring performance regression when transitioning from English to these Indian languages. It probes native-language visual question answering, mathematical reasoning, and multiple-choice comprehension across STEM and cultural domains. Use when the user wants to benchmark on HinTel-AlignBench, or asks about evaluating this task. Reports accuracy.
Evaluates cross-script machine translation quality for low-resource Indian languages (Hindi, Gujarati, Tamil) translating to English. It specifically probes how well models leverage back-translation augmented with quality and transliteration hints to handle noisy data and script conversion challenges. Use when the user wants to benchmark on IIT Bombay en-hi Corpus, WMT-2019 gu-en, TED2020, GNOME & Ubuntu, OPUS, WMT-2020 ta-en, GNOME, OPUS, WMT-2014 hi→en test, WMT-2019 gu→en test, WMT-2020 ta...
Evaluates machine translation models on code-mixed and noisy Hindi-English and Bengali-English text, measuring robustness to script variations, romanization, and synthetic noise. The protocol tests both in-domain performance on the HINMIX corpus and out-of-domain generalizability on LinCE, SpokenTutorial, and IITB Hi-En. It also assesses zero-shot transfer to unseen code-mixed Bengali-English translation. Use when the user wants to benchmark on HINMIX, or asks about evaluating this task. Repo...
Evaluates zero-shot information retrieval capabilities in Hindi across diverse domains and tasks. It probes how well multilingual embedding models and baselines rank relevant documents for Hindi queries without language-specific fine-tuning. Use when the user wants to benchmark on Hindi-BEIR, or asks about evaluating this task. Reports NDCG@10.
Evaluates the effectiveness of tree-based ensemble anomaly detectors in human-in-the-loop active learning settings. It measures how quickly an algorithm can discover anomalies by querying a limited budget of instances, comparing batch and streaming data paradigms. Use when the user wants to benchmark on Abalone, ANN-Thyroid-1v3, Cardiotocography, KDD-Cup-99, Mammography, Shuttle, Yeast, Covtype, Electricity, Weather, or asks about evaluating this task. Reports anomaly_discovery_rate.
Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the more recent HIKARI dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on HIKARI, or asks about evaluating this task. Reports F1S.
Evaluates a graph neural network's ability to predict particle velocities in particulate suspensions by learning many-body hydrodynamic interactions. It probes transferability across particle counts, external forcing types, and domain boundaries, alongside computational efficiency. Use when the user wants to benchmark on HIGNN-Training-Data, or asks about evaluating this task. Reports loss.
Evaluates models on hierarchical key performance indicator (KPI) extraction from SEC earnings filings, testing their ability to classify paragraph-level labels, perform token-level sequence labeling, and extract structured financial entities (tags, dates, currency, values) at varying granularities. Use when the user wants to benchmark on HiFi-KPI, HiFi-KPI Lite, or asks about evaluating this task. Reports aggregated macro F1.
Evaluates the inference throughput speedup of hierarchical speculative decoding against vanilla auto-regressive decoding and other acceleration baselines. It probes the method's ability to accelerate token generation across dialogue, summarization, code generation, and mathematical reasoning tasks without relying on auxiliary draft models. Use when the user wants to benchmark on ShareGPT, CNN/DM, XSum, HumanEval, GSM8K, or asks about evaluating this task. Reports Speedup (vs. Vanilla).
Evaluates the accuracy and stability of transformer head pruning methods (specifically HIES vs. baselines) across NLP, vision, and multimodal benchmarks at fixed sparsity ratios (10%, 30%, 50%). Use when the user wants to benchmark on GLUE (SST-2, CoLA, MRPC, QQP, STS-B, QNLI, MNLI, RTE), HellaSwag, Winogrande, ARC-e / ARC-c, OBQA, ImageNet1k, CIFAR-100, Food-101, Fashion MNIST, VizWiz-VQA, MM-Vet, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to perform fine-grained, taxonomy-aware visual recognition by predicting hierarchical biological labels (order, family, genus, species) from images. It specifically probes whether the model maintains logical consistency across taxonomic levels while accurately identifying leaf-level species, including generalization to unseen/novel categories. Use when the user wants to benchmark on iNaturalist-2021, TerraIncognita, or asks about evaluating this task. Reports Hiera...
Evaluates a model's ability to perform hierarchical multi-label image classification on remote sensing scenes. It probes how well the model captures label dependencies and hierarchy structures while predicting multiple overlapping categories per image. Use when the user wants to benchmark on UCM, AID, DFC-15, MLRSNet, or asks about evaluating this task. Reports AUPRC.
Evaluates human-object interaction recognition by decomposing activities into atomic body part states and reasoning hierarchically. Probes the model's ability to handle long-tail data and few-shot learning scenarios through compositional part-state representations. Use when the user wants to benchmark on HICO, or asks about evaluating this task. Reports mAP.
Evaluates the generalization and classification capabilities of foundational vision transformers on histopathology data across patch-level tissue classification, slide-level cancer subtyping, and nuclei segmentation tasks. Use when the user wants to benchmark on CRC-100K, MHIST, PCam, MSI-CRC, MSI-STAD, TIL-DET, BRCA, NSCLC, RCC, PanNuke, or asks about evaluating this task. Reports top-1 accuracy, AUC.
This benchmark evaluates language model alignment across four dimensions: helpfulness, honesty, harmlessness, and other. It measures how well a model's outputs adhere to human values and safety guidelines through pairwise comparison. Use when the user wants to benchmark on HHH_Alignment, or asks about evaluating this task. Reports Net Win Rate.
Evaluates an LLM agent's ability to execute profitable high-frequency trading decisions under strict latency constraints, balancing response speed with financial accuracy. The benchmark measures how well the model recognizes market patterns and executes trades within a fixed time window without degrading portfolio performance. Use when the user wants to benchmark on HFTBench, or asks about evaluating this task. Reports Daily Yield (%).
Evaluates the long-term phenotypic and genotypic dynamics of a heterogeneous cellular automaton with age constraints and local evolution, testing its ability to sustain open-ended innovation without stagnation. Use when the user wants to benchmark on Heterogeneous Life-Like CA Simulation, or asks about evaluating this task. Reports quantitative metrics.