LLMS Central - The Robots.txt for AI

claw-bench.com

Last updated: 7/28/2026valid

Independent Directory - Important Information

This llms.txt file was publicly accessible and retrieved from claw-bench.com. LLMS Central does not claim ownership of this content and hosts it for informational purposes only to help AI systems discover and respect website policies.

This listing is not an endorsement by claw-bench.com and they have not sponsored this page. We are an independent directory service with no affiliation to the listed domain.

Copyright & Terms: Users should respect the original terms of service of claw-bench.com. If you believe there is a copyright or terms of service violation, please contact us at support@llmscentral.com for prompt removal. Domain owners can also claim their listing.

Current llms.txt Content

# ClawBench

> ClawBench is a comprehensive benchmark for evaluating AI browser agents on 130 real-world everyday online tasks across 130+ live websites and 15 life categories. It captures 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), and scores with a two-stage evaluator: HTTP-request interception + LLM judge. The best-performing model (GLM 5.1) achieves 51.7% interception rate, revealing a large gap between current AI agents and human-level web task completion.

## Links

- [Paper](https://arxiv.org/abs/2604.08523): ClawBench: Can AI Agents Complete Everyday Online Tasks? (arXiv:2604.08523)
- [PDF](https://arxiv.org/pdf/2604.08523): Full paper PDF
- [Website](https://claw-bench.com): Interactive leaderboard, task browser, trace viewer, and agent demo gallery
- [GitHub](https://github.com/TIGER-AI-Lab/ClawBench): Source code — framework, evaluators, test driver, and Chrome extension
- [Dataset](https://huggingface.co/datasets/TIGER-Lab/ClawBenchV2Trace): V2 trace bundles on Hugging Face
- [Hugging Face Papers](https://huggingface.co/papers/2604.08523): Community discussion page
- [PyPI](https://pypi.org/project/clawbench-eval/): Install with `pip install clawbench-eval`

## Interactive Pages

- [Leaderboard](https://claw-bench.com/#results): Model rankings + per-task × per-model heatmap
- [Difficulty](https://claw-bench.com/#difficulty): 130 tasks sorted by cross-model pass rate, filterable by category
- [Categories](https://claw-bench.com/#categories): 15 life categories spanning daily life, work, dev, social, academic, travel, and more
- [Task Detail](https://claw-bench.com/#task/1): Per-task page with full instruction + per-model pass/fail
- [Model Profile](https://claw-bench.com/#model/glm-5.1): Per-model page with category strengths, unique solves, failures
- [Traces](https://claw-bench.com/#traces): Step-by-step agent execution with screenshots and HTTP requests
- [Gallery](https://claw-bench.com/#gallery): Screenshot slideshows of successful agent runs
- [Compare](https://claw-bench.com/#compare): How ClawBench compares to WebArena, REAL Bench, Claw-Eval
- [Contribute](https://claw-bench.com/#contribute): Community task proposal form (Low / Mid / High tier)

## Key Facts

- 130 tasks across 130+ live websites in 15 life categories
- Life categories include: food delivery, travel booking, job applications, shopping, housing search, email and calendar management, academic research, software development, learning platforms, office tasks, finance, entertainment, home services, pets, government services, and more
- 5 layers of behavioral data: session replay (rrweb), screenshots, HTTP traffic, agent reasoning traces, browser actions
- Two-stage scoring: HTTP-request interception (deterministic) + LLM judge (agentic)
- Request interceptor prevents irreversible real-world actions (payments, form submissions) during evaluation
- 6 models evaluated: GLM 5.1, DeepSeek V4 Pro, GLM 4.5 Air, DeepSeek V4 Flash, Claude Opus 4.7, GPT-5.5
- Apache 2.0 license
- COLM 2026 submission
- 21 authors from 11 institutions (UBC, Vector Institute, Etude AI, CMU, U Waterloo, SJTU, UniPat AI, ZJU, HKUST, Tsinghua, Netmind.ai)

## Leaderboard (V2 Interception Rate %)

| Model | Provider | Score |
|-------|----------|-------|
| GLM 5.1 | Zhipu AI | 51.7% |
| DeepSeek V4 Pro | DeepSeek | 43.8% |
| GLM 4.5 Air | Zhipu AI | 4.6% |
| DeepSeek V4 Flash | DeepSeek | 3.1% |
| Claude Opus 4.7 | Anthropic | 0.0% |
| GPT-5.5 | OpenAI | 0.0% |

## What Makes ClawBench Different

- **Real websites, not simulations**: Tasks run on 130+ actual live platforms (Airbnb, Uber Eats, Coursera, Indeed, etc.), not synthetic environments
- **Everyday tasks**: Booking flights, ordering groceries, applying for jobs, scheduling appointments — tasks people actually do online
- **Safe evaluation**: Request interceptor blocks the final HTTP request before irreversible actions, allowing evaluation on production websites without side effects
- **Rich behavioral data**: 5 complementary data layers enable fine-grained analysis of where and why agents fail
- **Two-stage scoring**: Deterministic HTTP interception for objective pass/fail, plus LLM judge for nuanced evaluation

## Citation

```bibtex
@article{zhang2026clawbench,
  title={ClawBench: Can AI Agents Complete Everyday Online Tasks?},
  author={Zhang, Yuxuan and Wang, Yubo and Zhu, Yipeng and Du, Penghui and Miao, Junwen and Lu, Xuan and Xu, Wendong and Hao, Yunzhuo and Cai, Songcheng and Wang, Xiaochen and Zhang, Huaisong and Wu, Xian and Lu, Yi and Lei, Minyi and Zou, Kai and Yin, Huifeng and Nie, Ping and Chen, Liang and Jiang, Dongfu and Chen, Wenhu and Allen, Kelsey R.},
  journal={arXiv preprint arXiv:2604.08523},
  year={2026}
}
```

Version History

Version 17/28/2026, 4:57:46 PMvalid
4808 bytes

Categories

entertainmentfinancetravelfoodsocial

Visit Website

Explore the original website and see their AI training policy in action.

Visit claw-bench.com

Content Types

pages

Recent Access

No recent access

API Access

Canonical URL:
https://llmscentral.com/claw-bench.com/llms.txt
API Endpoint:
/api/llms?domain=claw-bench.com