Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,793–6,816 of 20,827 skills
Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning. Use when the user wants to benchmark on CrossSum-In, Flores-In, XQuAD-In, XorQA-In, or asks about evaluating this task. Reports Character-F1 (ChrF), SQuAD-style Token-F1.
Evaluates large language models' reasoning capabilities on authentic Indian high-stakes examination questions across STEM and humanities domains. It specifically probes bilingual reasoning, cross-lingual performance differentials, and the impact of prompting strategies (Zero-Shot, Few-Shot, Chain-of-Thought) on model accuracy. Use when the user wants to benchmark on IndicEval, or asks about evaluating this task. Reports exact-match accuracy.
This evaluation probes the out-of-vocabulary (OOV) intelligibility and perceptual quality of Indian Text-to-Speech systems. It measures how well models synthesize rare or unseen words while preserving speaker similarity and overall audio fidelity compared to ground-truth recordings. Use when the user wants to benchmark on IndicTTS, or asks about evaluating this task. Reports Intelligibility Error Rate (%).
Evaluates neural machine translation models for Indic languages by measuring translation quality against reference texts across multiple standard benchmarks. It specifically tests the effectiveness of training on the Samanantar parallel corpus compared to commercial systems and existing open-source baselines. Use when the user wants to benchmark on WAT2020 Indic task, WAT2021 Multi-IndicMT task, WMT test sets (2014, 2019, 2020), UFAL Entam, FLORES test set, PMIndia en-as testset, or asks abou...
This evaluation probes the multilingual instruction-following, natural language understanding, and generation capabilities of LLMs fine-tuned on 13 Indic languages. It measures performance on standardized academic benchmarks across NLU and NLG tasks, as well as real-world cultural relevance and helpfulness through pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on MMLU Indic (MMLU-I), ARC Indic (ARC-I), BoolQ Indic (BoolQ-I), TriviaQA Indic (TVQA-I), BeleBele (Bele),...
Evaluates data-driven regional weather forecasting models over India under varying boundary conditioning strategies. It probes the ability of architectures to accurately predict multi-variable meteorological fields at high resolution and assesses their robustness during extreme weather events like heatwaves. Use when the user wants to benchmark on IndiaWeatherBench, or asks about evaluating this task. Reports RMSE.
This benchmark evaluates fine-grained music information retrieval by measuring how well models match audio tracks to diverse text queries. It probes the ability to capture nuanced musical attributes, handle negations, and rank candidates based on graded relevance rather than binary matches. Use when the user wants to benchmark on IncompeBench, or asks about evaluating this task. Reports graded relevance (0-3).
Evaluates multilingual language understanding and regional/cultural knowledge across 44 languages using native-language exam questions. Probes models' ability to handle region-specific contexts without English bias or translation artifacts, and measures performance variance across languages and prompting strategies. Use when the user wants to benchmark on INCLUDE-base, INCLUDE-lite, or asks about evaluating this task. Reports accuracy.
Probes the ability of deep learning models to detect malignancy in mammograms and generalize across different imaging scanners and patient populations. It specifically tests whether injecting stable, multi-scale topological features improves robustness to domain shifts compared to standard grayscale inputs. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports patient-level AUC.
Evaluates the ability of deep multi-instance learning models to classify whole mammograms as benign or malignant without requiring region-of-interest (ROI) annotations. It probes patch-level malignancy prediction and whole-image classification robustness under sparse label conditions. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports Accuracy.
Evaluates a deep learning model's ability to detect and classify malignant lesions in mammograms. It measures classification accuracy at the breast level and detection/localization sensitivity against false positive rates. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports AUC.
Evaluates multi-class image classification capabilities for agricultural species, genus, family, and crop/weed distinction. Probes fine-grained visual recognition and taxonomic hierarchy understanding in plant identification. Use when the user wants to benchmark on iNatAg, or asks about evaluating this task. Reports Accuracy.
Diagnoses VLM capabilities in autonomous driving by evaluating perception, prediction, meta-planning via Q&A accuracy, and planning via trajectory prediction L2 error. Use when the user wants to benchmark on Impromptu VLA, or asks about evaluating this task. Reports Q&A Accuracy.
Evaluates a model's ability to perform arithmetic and grade-school math reasoning without generating explicit intermediate chain-of-thought steps. It measures both the exact-match accuracy of the final answer and the inference speed relative to a no-CoT baseline. Use when the user wants to benchmark on Multi-digit multiplication, GSM8K, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to solve Olympiad-level mathematical proof problems by generating rigorous solutions and iteratively refining them through self-verification. Use when the user wants to benchmark on IMO 2025, or asks about evaluating this task. Reports accuracy.
This benchmark probes an LLM's ability to produce logically sound, step-by-step mathematical reasoning for Olympiad-level problems. It specifically measures the gap between achieving the correct final answer and maintaining rigorous, fallacy-free solution processes. Use when the user wants to benchmark on IMO shortlist problems (2009-2023), or asks about evaluating this task. Reports Final Answer Accuracy (%), Correct|Correct Final Answer (%).
Evaluates the ability of vision models to perform interactive medical image segmentation using user prompts like clicks, bounding boxes, or text. It probes how well models generalize across different imaging modalities, anatomical structures, and interaction strategies (single vs. multi-turn). Use when the user wants to benchmark on IMed-361M, TotalSegmentator MRI dataset, ISLES dataset, or asks about evaluating this task. Reports Dice score.
Evaluates models on recognizing spontaneous emotional states from unscripted speech and text. It probes acoustic prosody through dimensional regression and categorical classification, as well as linguistic sentiment polarity in real-world sports interview contexts. Use when the user wants to benchmark on iMiGUE-Speech, or asks about evaluating this task. Reports Categorical Emotion Classification.
Evaluates text-and-image-to-image editing capabilities, including addition, removal, replacement, motion change, style transfer, background change, object extraction, and hybrid edits. It tests the model's capacity to modify existing images according to natural language instructions while preserving unedited regions. Use when the user wants to benchmark on ImgEdit-Bench, or asks about evaluating this task. Reports ImgEdit-Bench.
Evaluates deep learning models for imbalanced and long-tailed classification and regression in AI-aided drug discovery. It probes model robustness to severe class imbalance, open long-tailed distributions, and out-of-distribution chemical splits across graph, sequence, and fingerprint molecular representations. Use when the user wants to benchmark on HIV, SBAP, USPTO-50K, DrugBank, or asks about evaluating this task. Reports Balanced-Acc.
Tests the model's capability to generate fixed-length representations for variable-length documents containing multiple sentences. It probes whether the method can scale to longer texts and outperform traditional bag-of-words baselines on a large-scale sentiment classification benchmark. Use when the user wants to benchmark on IMDB dataset, or asks about evaluating this task. Reports error rate.
Evaluates the perceptual quality and naturalness of synthesized Malayalam speech generated by a multi-speaker TTS model trained on the IMaSC corpus. It probes the model's ability to capture agglutinative morphology, phonemic orthography, and diverse prosodic styles through subjective human listening tests. Use when the user wants to benchmark on IMaSC, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
Evaluates zero- and few-shot visual commonsense reasoning capabilities of language models and visually-augmented language models across 1,000 ImageNet categories using human-annotated QA pairs. Use when the user wants to benchmark on ImageNetVC, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates vision-language models' ability to generate structured, step-by-step reasoning (thinking tokens) and final answers for multimodal inputs. It probes reasoning coherence, logical progression, and alignment with reference synthetic reasoning traces. Use when the user wants to benchmark on ImageNet-Think-250K, or asks about evaluating this task. Reports BERTScore.