Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,827
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,745–6,768 of 20,827 skills

Intelligence Per Watt EvalA

Evaluates the efficiency of local LLM inference by combining task accuracy with energy consumption to compute Intelligence per Watt (IPW). It probes how model architecture, hardware acceleration, and numerical precision affect the trade-off between performance and power usage on real-world chat and reasoning tasks. Use when the user wants to benchmark on WildChat, NaturalReasoning, SuperGPQA, MMLU Pro, or asks about evaluating this task. Reports accuracy, intelligence per watt (IPW).

researchpythongo
0
3
Integrated Brier ScoreA

Evaluates the calibration and predictive accuracy of ensemble forecasting methods for time-to-event outcomes in meteorology. It compares how well different combination techniques predict the timing of events like the first hard freeze. Use when the user has predictions and gold and needs to compute Mean Integrated Brier Score (IBS).

researchpythongo
0
3
Integer Quantization EvalA

Evaluates the classification and detection accuracy, as well as inference latency, of neural networks quantized to 8-bit integer arithmetic on mobile ARM CPUs compared to floating-point baselines. Use when the user wants to benchmark on ImageNet, COCO, Face detection dataset, Face attributes dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Insure Dial EvalA

Probes phase-aware compliance verification and phase boundary detection in insurance benefit verification calls. It measures a model’s ability to accurately segment conversational phases under workflow-specific rules and apply rule-based compliance reasoning (Information and Procedural Compliance) to fixed spans. Use when the user wants to benchmark on INSURE-Dial, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Instructttseval Zh EvalA

Evaluates a text-to-speech model's ability to follow natural-language instructions for voice design, specifically controlling acoustic parameters, descriptive styles, and role-play characteristics. It measures how accurately synthesized speech adheres to explicit semantic and stylistic requirements. Use when the user wants to benchmark on InstructTTSEval-Zh, or asks about evaluating this task. Reports AVG.

researchpythonapi
0
3
Instructtts EvalA

Evaluates a text-to-speech model's capability to generate realistic vocal timbres that accurately follow complex natural-language style instructions. It probes fine-grained acoustic control, generalization to unstructured descriptions, and contextual role-play inference. Use when the user wants to benchmark on InstructTTSEval, or asks about evaluating this task. Reports Instruction-following accuracy (%).

researchpythongo
0
3
Instructpart EvalA

Evaluates Vision-Language Models' ability to perform fine-grained visual grounding and instruction reasoning for part segmentation. It probes whether models can infer task-relevant object parts from natural language instructions or oracle prompts, and assesses their capacity for affordance learning in human-robot interaction contexts. Use when the user wants to benchmark on InstructPart, or asks about evaluating this task. Reports gIoU.

researchpython
0
3
Instructir EvalA

Evaluates whether information retrieval models can accurately follow instance-specific, user-aligned instructions rather than generic task descriptions. It probes the robustness of retrievers to instruction variations and their ability to adapt to real-world search scenarios with diverse user contexts. Use when the user wants to benchmark on InstructIR, or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3
Instruction Tuning EvalA

Evaluates the instruction-following capability and alignment (helpfulness, honesty, harmlessness) of instruction-tuned LLMs on unseen tasks across English and Chinese. Use when the user wants to benchmark on User-Oriented-Instructions-252, Vicuna-Instructions-80, Unnatural Instructions, or asks about evaluating this task. Reports Relative Score (GPT-4).

researchpython
0
3
Instruction Robustness EvalA

Evaluates the zero-shot robustness of instruction-tuned language models to variations in instruction phrasing, even when instructions are semantically equivalent. It measures how well models maintain performance on unobserved instruction variants compared to observed ones. Use when the user wants to benchmark on MMLU, BBL, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Instruction Pretraining EvalA

Evaluates the generalization and domain-adaptive capabilities of language models pre-trained with instruction-augmented corpora. Probes zero/few-shot instruction following, general knowledge, and specialized performance in biomedicine and finance. Use when the user wants to benchmark on MMLU, PubMedQA, ChemProt, RCT, MQP, UMSLE, ConvFinQA, Headline, FiQA SA, FPB, NER, or asks about evaluating this task. Reports average task score.

researchpythongo
0
3
Instruction Nmt EvalA

Evaluates whether neural machine translation models can follow diverse natural language instructions (e.g., formality, voice, casing, simplification) without task-specific retraining, while maintaining general translation quality. Use when the user wants to benchmark on WMT'20 News Translation (EN-DE), Multi-30K, Custom Instruction Dataset, or asks about evaluating this task. Reports RR (%).

researchpythonapi
0
3
Instruction Editing EvalA

Evaluates instruction-guided image editing models on their ability to modify an input image according to a text prompt. It measures semantic alignment with the prompt and visual fidelity to a ground-truth edit, plus human preference in pairwise comparisons. Use when the user wants to benchmark on MagicBrush (MagBr), ZONE, or asks about evaluating this task. Reports CLIP-T.

researchpython
0
3
Instructgpt EvalA

Evaluates instruction-following alignment, truthfulness, toxicity, and bias in large language models. It measures how well model outputs match human preferences and public benchmark standards compared to base models. Use when the user wants to benchmark on API Prompt Distribution, TruthfulQA, RealToxicityPrompts, Winogender, CrowS-Pairs, or asks about evaluating this task. Reports winrate.

researchpythongo
0
3
Instructaudio EvalA

This evaluation probes a model's ability to generate speech and music conditioned on natural language instructions describing acoustic and musical attributes. It measures text-to-audio fidelity, attribute control accuracy, and perceptual quality across short-form generation tasks. Use when the user wants to benchmark on Seed-TTS benchmark, InstructAudio internal test set, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Instruct Tts EvalA

Evaluates a text-to-speech system's ability to follow complex natural-language instructions for acoustic parameter specification, descriptive style direction, and role-play scenarios. It probes fine-grained prosodic control, open-ended style inference, and high-level scenario-based emotional/character expression. Use when the user wants to benchmark on InstructTTSEval, or asks about evaluating this task. Reports accuracy.

researchpythonexpress
0
3
Instance Attribution EvalA

Evaluates the ability of different instance attribution methods to rank training data instances by their influence on a given test prediction, particularly focusing on identifying problematic training artifacts and comparing gradient-based versus similarity-based approaches. Use when the user wants to benchmark on SST-2, MNLI, HANS, or asks about evaluating this task. Reports Spearman Correlation.

researchpython
0
3
Inst It Bench EvalA

Evaluates a model's ability to perform fine-grained, instance-level understanding on images and videos. It probes spatial-temporal grounding, multi-level annotation comprehension (captions, temporal changes), and multiple-choice question answering over explicitly prompted visual regions. Use when the user wants to benchmark on Inst-IT Bench, or asks about evaluating this task. Reports average score.

researchpythongo
0
3
Inspiration Retrieval EvalA

Evaluates an LLM's ability to retrieve relevant prior research papers (inspirations) that can inform a given research question from a candidate pool. It measures how well models can surface novel, non-obvious knowledge links through iterative group-based selection. Use when the user wants to benchmark on ResearchBench Inspiration Retrieval, or asks about evaluating this task. Reports Hit Ratio.

researchpythongo
0
3
Insight O3 EvalA

Evaluates multimodal reasoning and generalized visual search capabilities, measuring how well models can locate and reason about high-information-density images using a multi-agent framework with a dedicated visual search agent. Use when the user wants to benchmark on V*-Bench, Tree-Bench, VisualProbe-Hard, HR-Bench, MME-RealWorld, O3-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Insectset459 EvalA

Evaluates multi-class bioacoustic classification of insect audio recordings into one of 459 species. It probes a model's robustness to severe class imbalance, highly variable sampling rates, and ultrasonic frequency ranges. Use when the user wants to benchmark on InsectSet459, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Inquire EvalA

Evaluates text-to-image retrieval capabilities of vision-language models on expert-level, ecologically grounded queries. It probes fine-grained visual understanding, domain-specific language comprehension, and ranking quality across multiple relevant images per query. Use when the user wants to benchmark on INQUIRE, or asks about evaluating this task. Reports AP@k (mAP@50).

researchpythonperformance
0
3
Innovatorbench EvalA

Evaluates AI agents' ability to conduct end-to-end LLM research across six domains: data construction, filtering, augmentation, loss/reward design, and scaffold construction. It probes long-horizon decision making, algorithmic robustness, resource management, and iterative code generation in a simulated research environment. Use when the user wants to benchmark on InnovatorBench, or asks about evaluating this task. Reports Best Score.

researchpythongo
0
3
Innovator Vl EvalA

Evaluates multimodal large language models across general vision, mathematical reasoning, and specialized scientific domains to measure visual perception, instruction following, and domain-specific knowledge retention. Use when the user wants to benchmark on AI2D, OCRBench, ChartQA, MMMU(Val), MMMU-Pro (Standard), MMStar, VStar-Bench, MMBench-EN, MME-RealWorld, DocVQA(Val), InfoVQA(Val), SEED-Bench, SEED-Bench-2-plus, RealWorldQA, MathVision, MathVerse, MathVista, WeMath, ScienceQA, RxnBench,...

researchpythongo
0
3