Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits. Use when the user wants to benchmark on WebWalkerQA, or asks about evaluating this task. Reports accuracy.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill webwalker-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Webwalker Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-webwalker-eval)More formats (shields.io, HTML) on the badges page.
---
name: webwalker-eval
description: Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits. Use when the user wants to benchmark on WebWalkerQA, or asks about evaluating this task. Reports accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2501.07572
bibtex_key: wu2025webwalker
confidence: high
---
# webwalker-eval
> WebWalker: Benchmarking LLMs in Web Traversal — Jialong Wu et al. (2025) (arXiv:2501.07572, 2025)
## What this evaluates
Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits.
## Datasets
- **WebWalkerQA** — total ?; splits: Single-source QA - Easy (-1), Single-source QA - Medium (-1), Single-source QA - Hard (-1), Multi-source QA - Easy (-1), Multi-source QA - Medium (-1), Multi-source QA - Hard (-1); repo https://github.com/Alibaba-NLP/WebAgent
## Metrics
- `accuracy` **(primary)** — range: percent
- Percentage of correctly answered queries out of the total number of queries in a split. Calculated as (correct predictions / total queries) * 100.
- `action_count` — range: other
- Average number of navigation steps (clicks) taken by the agent, counted only for executions that ultimately yield the correct answer.
## Input / output format
**Input**: A natural language query requiring information extraction from a specific, multi-layered real-world website. The agent receives the current page's text and interactive elements (buttons/links) to decide the next click.
**Output**: A sequence of click actions on the website interface, terminated when the agent submits an answer or reaches the maximum step limit (K=15).
## Scoring recipe
```python
def score(predictions, golds, max_steps=15):
correct = 0
action_counts = []
for pred, gold in zip(predictions, golds):
if pred.answer == gold.answer:
correct += 1
if pred.steps <= max_steps:
action_counts.append(pred.steps)
accuracy = (correct / len(golds)) * 100
avg_actions = sum(action_counts) / len(action_counts) if action_counts else 0
return {'accuracy': accuracy, 'action_count': avg_actions}
```
## Common pitfalls
- Action count (A.C.) is only averaged over correct executions, which may overstate efficiency for models that succeed quickly but fail often.
- The hard cap of 15 steps severely penalizes tasks requiring deeper vertical exploration, making accuracy drops at higher difficulties partly a function of the step limit rather than pure reasoning failure.
- Real-world website navigation requires handling dynamic layouts and non-standard button labels, which can cause failures unrelated to the model's logical reasoning.
## Evidence (verbatim from paper)
> We limit the number of actions K for the explorer agent to 15, meaning that the explorer agent can explore at most 15 steps. We only counted the action count (A.C.) from correct executions, and as the model size increases, the A.C. grows, indicating that larger LLMs have enhanced long-range information-seeking ability. Even the best-performing WebWalker using GPT-4o as its backbone does not surpass 40%, highlighting the challenge posed by WebWalkerQA.
## Citation
```bibtex
@misc{wu2025webwalker,
title={WebWalker: Benchmarking LLMs in Web Traversal},
author={Jialong Wu et al. (2025)},
year={2025},
note={arXiv:2501.07572}
}
```
- arXiv: 2501.07572
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!