Evaluating whether multimodal large language models truly understand long-form scientific papers remains challenging: answer-only metrics and synthetic 'Needle-In-A-Haystack' tests often reward answer matching without requiring a causal, evidence-linked reasoning trace in the document. We propose the 'Fish-in-the-Ocean' (FITO) paradigm, which requires models to construct explicit cross-modal evidence chains within native scientific documents. To operationalize FITO, we build SIN-Data, a scien...
Scanned 9/9/2026
Install to Claude Code
npx -y skills add ADu2021/skillXiv --skill sin-bench-tracing-native-evidence-chains-in-long --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Sin Bench Tracing Native Evidence Chains In Long?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/adu2021-sin-bench-tracing-native-evidence-chains-in-long)More formats (shields.io, HTML) on the badges page.
---
name: sin-bench-tracing-native-evidence-chains-in-long
title: "SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Semantic Integration"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: "https://arxiv.org/abs/2601.10108"
keywords: [Agents, Benchmarking]
description: "Evaluating whether multimodal large language models truly understand long-form scientific papers remains challenging: answer-only metrics and synthetic 'Needle-In-A-Haystack' tests often reward answer matching without requiring a causal, evidence-linked reasoning trace in the document. We propose the 'Fish-in-the-Ocean' (FITO) paradigm, which requires models to construct explicit cross-modal evidence chains within native scientific documents. To operationalize FITO, we build SIN-Data, a scientif..."
---
## Problem
SIN-Bench addresses key challenges in autonomous agent development. This paper provides solutions for evaluating, building, or improving agent systems.
## Key Approach
The paper introduces a novel framework, methodology, or benchmark for sin-bench. The core contributions include:
1. Systematic framework or benchmark for agent evaluation and development
2. Empirical findings on agent performance, efficiency, or capabilities
3. Generalizable principles applicable across domains
## When to Use
Use this skill when you need to:
- Evaluate or benchmark autonomous agent systems
- Understand best practices in agent design and evaluation
- Learn empirical results on agent performance
- Improve agent efficiency, reasoning, or capabilities
## When NOT to Use
- For non-agent-related tasks
- When seeking quick implementation code (see the paper for details)
- For general knowledge unrelated to autonomous agents
## Resources
- ArXiv Abstract: https://arxiv.org/abs/2601.10108
- Full PDF: https://arxiv.org/pdf/2601.10108
- HTML Version: https://arxiv.org/html/2601.10108
See the paper for comprehensive methodology, experimental protocols, benchmarks, and implementation details.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!