Browser Agent Research
WebMCP & Browser Agent Evaluation
Built browser-native tool suites and the evaluation infrastructure used to study how LLM agents act across real web environments.
Problem
Browser agents lost reliability when tool dispatch and element IDs drifted across live pages.
Built
WebMCP tools, a Python evaluation harness, and five WebArena study modes across 9+ sites.
Result
Traced failures across the stack and made the evaluation signal materially cleaner.
- 9+ sites
- tool-enabled environments
- 10%
- avg. success rate increase
- 40%
- token reduction vs. baseline

