Skip to content

9.5 Computer Use & Browser Agents

Agents that drive a screen or a browser: what works, the benchmark-versus-real-task gap, and when a tool API beats screen-driving.

Also called: computer use, browser agents, GUI agents, web agents.

Stub: scaffolding, not finished writing · planned for a later release. The skeleton below shows the beats this chapter will hit. Contributions welcome.

Why you'd reach for it

The problem, what breaks without it, and when you need it. To be written.

What it actually is

A crisp definition, the maturity call argued with cited evidence, and how it differs from its neighbours. To be written.

How to do it

The mechanics, with public tools named plainly where the reader can verify them; code only where a tested listing earns its place. To be written.

Gotchas

The real costs and when not to, feeding the Anti-Patterns Catalog. To be written.

In short

A weighted recommendation: what you would actually do. To be written.

Maturity: Established (exists) / Contested (reliability) (the capability ships from every major lab; the reliability on real tasks does not yet) · Grounding: research · Last reviewed: 2026-08

Sources

Citations added as the chapter is written. Every non-obvious claim gets a footnote.

See also

Related chapters.