llms.txt Content
# Steel Agent Leaderboard
> Benchmark hub with canonical benchmark leaderboard pages.
Maintained by Steel (https://steel.dev).
## Leaderboard
- [WebVoyager](https://leaderboard.steel.dev/leaderboards/webvoyager/): WebVoyager benchmark leaderboard for AI browser agents on 643 live-web tasks across 15 popular websites, with source-linked scores and methodology notes.
- [BrowseComp](https://leaderboard.steel.dev/leaderboards/browsecomp/): BrowseComp leaderboard for agentic web research systems solving OpenAI's hard-to-find short-answer browsing benchmark, with sourced scores and setup notes.
- [DRACO](https://leaderboard.steel.dev/leaderboards/draco/): DRACO leaderboard for deep research systems on Perplexity's benchmark: 100 expert-graded research tasks across 10 domains, with sourced scores and notes on the grading judge.
- [WebArena](https://leaderboard.steel.dev/leaderboards/webarena/): WebArena leaderboard for autonomous browser agents evaluated on reproducible, self-hosted web tasks across shopping, forum, GitLab, CMS, map, and wiki environments.
- [SWE-bench Verified](https://leaderboard.steel.dev/leaderboards/swe-bench-verified/): SWE-bench Verified leaderboard for coding agents resolving 500 human-filtered real GitHub issues with Docker-based test execution.
- [Aider](https://leaderboard.steel.dev/leaderboards/aider/): Aider leaderboard ranking LLMs on the Aider Polyglot benchmark: 225 of the hardest Exercism exercises across C++, Go, Java, JavaScript, Python, and Rust, scored inside Aider's real edit loop.
- [OSWorld](https://leaderboard.steel.dev/leaderboards/osworld/): OSWorld leaderboard for multimodal computer-use agents completing 369 real desktop tasks with execution-based verification.
- [OSWorld 2.0](https://leaderboard.steel.dev/leaderboards/osworld-2/): OSWorld 2.0 leaderboard for computer-use agents on 108 long-horizon real-world desktop workflows that take human users a median of about 1.6 hours.
- [GAIA](https://leaderboard.steel.dev/leaderbo