Evaluates a vision-language model's ability to navigate and complete multi-step tasks on mobile GUIs. It measures step-level action correctness, full trajectory success, and robustness to intermediate errors in a simulated Android environment. Use when the user wants to benchmark on Chinese Mobile Agent Benchmark, or asks about evaluating this task. Reports Accuracy (Acc.).
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill mobile-r1-benchmark-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Mobile R1 Benchmark Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-mobile-r1-benchmark-eval)More formats (shields.io, HTML) on the badges page.
---
name: mobile-r1-benchmark-eval
description: Evaluates a vision-language model's ability to navigate and complete multi-step tasks on mobile GUIs. It measures step-level action correctness, full trajectory success, and robustness to intermediate errors in a simulated Android environment. Use when the user wants to benchmark on Chinese Mobile Agent Benchmark, or asks about evaluating this task. Reports Accuracy (Acc.).
metadata:
skill_kind: dataset_eval
source_arxiv: 2506.20332
bibtex_key: gu2025mobiler1
confidence: high
---
# mobile-r1-benchmark-eval
> Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards — Gu et al. (2025) (arXiv:2506.20332, 2025)
## What this evaluates
Evaluates a vision-language model's ability to navigate and complete multi-step tasks on mobile GUIs. It measures step-level action correctness, full trajectory success, and robustness to intermediate errors in a simulated Android environment.
## Datasets
- **Chinese Mobile Agent Benchmark** — total 500; splits: test (500)
## Metrics
- `Accuracy (Acc.)` **(primary)** — range: percent
- Percentage of steps in a trajectory where both the predicted action format and action type match the ground truth definitions.
- `Task Success Ratio (Task Succ.)` — range: percent
- Percentage of trajectories where every single step is executed entirely correctly without any deviations.
- `Tail Success Ratio (Tail Succ)` — range: percent
- Percentage of trajectories that ultimately complete the final task goal, regardless of intermediate errors or deviations.
- `Action Argument Error Number (Avg Err.)` — range: other
- Average count of incorrect action arguments per trajectory.
## Input / output format
**Input**: Task instruction and sequential mobile GUI screenshots (visual state) provided to a VLM.
**Output**: A sequence of structured actions per step, specifying action type and required parameters (e.g., click coordinates, text input).
## Scoring recipe
```python
def score_trajectory(pred_steps, gold_steps):
step_acc = sum(1 for p, g in zip(pred_steps, gold_steps) if p.format == g.format and p.action == g.action) / len(gold_steps)
task_succ = 1.0 if all(p.format == g.format and p.action == g.action for p, g in zip(pred_steps, gold_steps)) else 0.0
tail_succ = 1.0 if pred_steps[-1].final_state == gold_steps[-1].goal_state else 0.0
err_count = sum(1 for p, g in zip(pred_steps, gold_steps) if p.action != g.action)
return step_acc * 100, task_succ * 100, tail_succ * 100, err_count
```
## Common pitfalls
- Confusing step-level Accuracy with Task Success Ratio; the former allows intermediate failures while the latter requires 100% step correctness.
- Tail Success Ratio explicitly ignores intermediate deviations, so evaluators must track the final task completion state rather than penalizing every wrong click.
- Avg Err. is a raw count of incorrect actions, not a normalized ratio, so it should not be averaged across trajectories without weighting by trajectory length.
## Evidence (verbatim from paper)
> We evaluate the model’s performance using the following metrics: Accuracy (Acc.): The probability of correctly performing each step in a trajectory, correct if both format and action match definitions R_F and R_Act. • Task Success Ratio (Task Succ.): The probability of a complete trajectory being executed entirely correctly. • Tail Success Ratio (Tail Succ): The probability that the task within a trajectory is ultimately completed successfully, regardless of intermediate errors or deviations. • Action Argument Error Number (Avg Err.): The count of errors of incorrect action.
## Citation
```bibtex
@misc{gu2025mobiler1,
title={Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards},
author={Gu et al. (2025)},
year={2025},
note={arXiv:2506.20332}
}
```
- arXiv: 2506.20332
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!