Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,827
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,697–6,720 of 20,827 skills

Jailbreak Audio Bench EvalA

This benchmark probes the safety alignment and jailbreak resilience of Large Audio-Language Models (LALMs). It specifically tests whether manipulating audio-specific hidden semantics—such as tone, intonation, emotion, and background noise—can bypass safety guardrails and elicit harmful responses more effectively than text-only prompts. Use when the user wants to benchmark on Jailbreak-AudioBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythonrails
0
3
Jailbreak Attack EvalA

This protocol evaluates the robustness of large language models against automated jailbreak attacks. It measures how effectively generated or human-crafted prompts can bypass safety filters to elicit prohibited or harmful responses. Use when the user wants to benchmark on 100 questions from two open datasets [6,37], or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpython
0
3
Iwslt2023 St EvalA

Evaluates automatic speech translation systems on long-form audio across offline, multilingual, and simultaneous conditions. Probes the model's ability to handle segmentation, resegmentation, and translation quality under varying acoustic and linguistic challenges. Use when the user wants to benchmark on IWSLT2023 TED Test Set, IWSLT2023 ACL Test Set, or asks about evaluating this task. Reports COMET.

researchpython
0
3
Iwslt2017 Nmt EvalA

Evaluates neural machine translation quality of character-level versus subword models across multiple language pairs. It probes morphological generalization, noise robustness, and the impact of sequence length expansion on training and inference efficiency. Use when the user wants to benchmark on IWSLT 2017, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
Ivy Fake EvalA

This benchmark evaluates multimodal AI-generated content (AIGC) detection and explainable reasoning capabilities. It probes a model's ability to classify images and videos as real or fake, and to generate natural-language explanations that localize and justify synthetic artifacts. Use when the user wants to benchmark on Ivy-Fake, GenImage, Chameleon, GenVideo, or asks about evaluating this task. Reports Accuracy (Acc).

researchpythongo
0
3
Iu Xray Report Gen EvalA

Evaluates a vision-language model's ability to generate clinically accurate and semantically coherent radiology reports from chest X-ray images. It probes the model's capacity for medical terminology usage, anatomical consistency, and structured clinical text generation. Use when the user wants to benchmark on IU X-ray, or asks about evaluating this task. Reports ROUGE-L.

researchpythonperformance
0
3
Iu Rr Radiology Report EvalA

Evaluates a model's ability to generate clinically accurate and structurally coherent radiology reports from multi-view chest X-ray images. It probes cross-modal alignment, medical terminology recall, and the model's capacity to synthesize findings and impressions from visual evidence. Use when the user wants to benchmark on IU-RR, or asks about evaluating this task. Reports BLEU-4.

researchpythontesting
0
3
Iteris Merging EvalA

Evaluates the effectiveness of iterative LoRA merging (IterIS) across text-to-image diffusion, vision-language, and large language models. It probes the model's ability to preserve multiple concepts or styles without mutual interference while maintaining generation quality and task-specific performance metrics. Use when the user wants to benchmark on CustomConcept101, DreamBooth, SentiCap, Emotion datasets (Emoint, EC, TEC, ISEAR, SUM), GLUE benchmark, or asks about evaluating this task. Repo...

researchpythongo
0
3
Ist Unbabel 2022 Qe EvalA

Evaluates machine translation quality estimation (QE) by predicting human quality scores at the sentence level and identifying error locations at the word level. It also assesses the model's ability to generate faithful explanations for predicted errors. Use when the user wants to benchmark on IST-Unbabel 2022 QE Shared Task, or asks about evaluating this task. Reports Spearman's rank correlation, Matthew's correlation coefficient (MCC), Recall@K (R@K).

researchpythongo
0
3
Isles24 Segmentation EvalA

Evaluates the ability of models to perform 3D medical image segmentation for stroke lesion (infarct) and vessel occlusion detection using longitudinal multimodal CT and MRI scans. Use when the user wants to benchmark on ISLES'24, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).

researchpythongit
0
3
Isign EvalA

Evaluates the accuracy of English text generation from Indian Sign Language (ISL) videos and pose sequences. It probes multimodal translation capabilities, specifically how well models align visual sign language signals with corresponding natural language references. Use when the user wants to benchmark on iSign, or asks about evaluating this task. Reports BLEU-4.

researchpythonexpress
0
3
Isic Ham Segmentation EvalA

Evaluates dermatologic image segmentation models by measuring how training on real versus synthetic data affects performance on held-out real test sets, and how model accuracy correlates with controllable synthetic image parameters like skin tone and lesion shape. Use when the user wants to benchmark on ISIC, HAM, or asks about evaluating this task. Reports Dice score.

researchpythonperformance
0
3
Isdrama EvalA

This evaluation protocol assesses the capability of multimodal speech synthesis models to generate high-fidelity, spatially accurate binaural audio from scripts, poses, and prompts. It probes content accuracy, speaker similarity, prosodic expressiveness, and precise spatial localization (interaural phase/level differences and angle/distance consistency). Use when the user wants to benchmark on MRSDrama, or asks about evaluating this task. Reports IPD MAE.

researchpythonexpress
0
3
Isafetybench EvalA

Probes vision-language models' ability to recognize routine and hazardous industrial actions in real-world videos under zero-shot conditions. It tests both single-label precision and multi-label recall in safety-critical contexts, evaluating how well models discriminate between semantically similar distractors and identify multiple concurrent actions. Use when the user wants to benchmark on iSafetyBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Isac Lawn Multimodal EvalA

Probes the capability of multimodal fusion and adaptive expert routing for integrated sensing and communication tasks in low-altitude wireless networks. Specifically, it evaluates how well models leverage synchronized visual, lidar, radar, GPS, and RF channel data to predict beam indices, estimate path loss, and track UAV trajectories under dynamic environmental conditions. Use when the user wants to benchmark on Public Multimodal ISAC Dataset for Low-Altitude Scenarios, or asks about evaluat...

researchpythongo
0
3
Isaacsim Kitchen EvalA

Evaluates a robot's ability to decompose high-level language instructions into executable task plans and execute them in a simulated kitchen environment. It jointly measures planning accuracy and low-level control success under strict time and spatial constraints. Use when the user wants to benchmark on IsaacSim Kitchen Benchmark, or asks about evaluating this task. Reports EM.

researchpython
0
3
Irt2 EvalA

Evaluates neural and baseline models on inductive link prediction and ranking tasks across knowledge graphs of varying scales. It probes the models' ability to map textual entity mentions to graph vertices and rank candidate entities based on combined textual and structural signals, particularly under data scarcity conditions. Use when the user wants to benchmark on IRT2, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Irsc EvalA

Evaluates embedding models on multilingual information retrieval tasks across five query types (query, title, part-of-paragraph, keyword, summary). It probes semantic comprehension and cross-lingual retrieval alignment in Retrieval-Augmented Generation (RAG) scenarios. Use when the user wants to benchmark on IRSC Benchmark, or asks about evaluating this task. Reports r@10.

researchpythongo
0
3
Irpapers EvalA

Evaluates the ability of multimodal and text-only models to retrieve relevant scientific paper pages and answer questions based on those pages. It probes retrieval depth, modality complementarity, and the impact of context quantity on RAG performance. Use when the user wants to benchmark on IRPAPERS, or asks about evaluating this task. Reports Recall@1.

researchpythongo
0
3
Irish English St EvalA

Evaluates end-to-end speech translation from Irish to English, specifically probing how synthetic audio data and augmentation techniques (noise, VAD) impact model performance in low-resource settings. Use when the user wants to benchmark on IWSLT-2023, FLEURS, Bitesize, SpokenWords, or asks about evaluating this task. Reports chrF++.

researchpythongit
0
3
Iris Benchmark EvalA

Probes fairness across understanding and generation tasks in Unified Multimodal Large Language Models (UMLLMs) by measuring Ideal Fairness, Real-world Fidelity, and Bias Inertia & Steerability across demographic attributes. It reveals systemic trade-offs, generation gaps, and personality splits that single-task or single-metric evaluations miss. Use when the user wants to benchmark on IRIS Benchmark, or asks about evaluating this task. Reports IRIS-Score.

researchpythongo
0
3
Irefvla EvalA

Evaluates a model's ability to ground referential language in 3D scenes when references are imperfect or ambiguous. It probes whether the model can correctly identify existing objects, detect non-existent references, and generate plausible alternative objects based on spatial and semantic reasoning. Use when the user wants to benchmark on IRef-VLA, or asks about evaluating this task. Reports score_sim.

researchpythongo
0
3
Ircad Liver EvalA

Evaluates medical image segmentation models on liver CT volumes, testing their ability to accurately delineate organ boundaries using interactive or automatic refinement techniques. The protocol measures how well models handle low-contrast boundaries and varying slice geometries in clinical imaging. Use when the user wants to benchmark on IRCAD, or asks about evaluating this task. Reports Dice coefficient.

researchpythongo
0
3
Ir Triplet EvalA

Evaluates the model's ability to capture semantic similarity between paragraphs for information retrieval. It tests whether fixed-length vector representations can effectively distinguish query-related documents from irrelevant ones using a triplet ranking protocol. Use when the user wants to benchmark on Information Retrieval (Paragraph Vectors), or asks about evaluating this task. Reports error rate.

researchpython
0
3