Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill otb99-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Otb99 Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-otb99-eval)More formats (shields.io, HTML) on the badges page.
---
name: otb99-eval
description: Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.
metadata:
skill_kind: dataset_eval
source_arxiv: 2508.05221
bibtex_key: wang2025reasoningtrack
confidence: high
---
# otb99-eval
> ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking — Xiao Wang et al. (2025) (arXiv:2508.05221, 2025)
## What this evaluates
Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations.
## Datasets
- **OTB99** — total 99; splits: train (51), test (48)
## Metrics
- `PR` **(primary)** — range: percent
- Precision Rate: percentage of frames where the center distance is within 20 pixels.
- `AUC` — range: percent
- Area Under the Curve: area under the precision plot across varying distance thresholds.
## Input / output format
**Input**: Video frames, initial language description, and ground-truth bounding boxes.
**Output**: Predicted bounding box coordinates per frame.
## Scoring recipe
```python
For each frame, compute center distance d.
PR = (count(d < 20) / N) * 100
AUC = integral of precision plot over thresholds [0, 50]
```
## Common pitfalls
- Text annotations are added post-hoc to OTB100, so baseline methods may not be trained with language guidance.
- Short-term tracking may not stress-test long-horizon language updates.
## Evidence (verbatim from paper)
> OTB99 consists of 99 video sequences, divided into 51 for training and 48 for testing. As shown in Table[V], our method achieves the best performance in two metrics, with an AUC of 71.11% and PR of 95.58%.
## Citation
```bibtex
@misc{wang2025reasoningtrack,
title={ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking},
author={Xiao Wang et al. (2025)},
year={2025},
note={arXiv:2508.05221}
}
```
- arXiv: 2508.05221
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!