Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where each task provides an image-text bundle to evaluate multimodal understanding and citation-grounded report generation. Compared to prior setups, MMDR-Bench emp...
Scanned 9/9/2026
Install to Claude Code
npx -y skills add ADu2021/skillXiv --skill mmdeepresearch-bench-a-benchmark-for-multimodal --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Mmdeepresearch Bench A Benchmark For Multimodal?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/adu2021-mmdeepresearch-bench-a-benchmark-for-multimodal)More formats (shields.io, HTML) on the badges page.
---
name: mmdeepresearch-bench-a-benchmark-for-multimodal
title: "MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: "https://arxiv.org/abs/2601.12346"
keywords: [Agent, Benchmark]
description: "Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where each task provides an image-text bundle to evaluate multimodal understanding and citation-grounded report generation. Compared to prior setups, MMDR-Bench emphas..."
---
## Overview
This skill covers mmdeepresearch-bench: a benchmark for multimodal deep research agents. It addresses critical challenges in autonomous agent development.
## Key Concepts
The paper introduces novel approaches to:
- Agent evaluation and benchmarking
- Improving agent efficiency and reasoning
- Designing robust agent systems
## When to Use
Use this when working on:
- Agent-based systems and evaluation
- Autonomous reasoning and planning
- Multi-agent frameworks
## When NOT to Use
- Non-agent applications
- Tasks requiring implementation code (see the paper)
## References
- Paper: https://arxiv.org/abs/2601.12346
- PDF: https://arxiv.org/pdf/2601.12346
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!