Computer Use
Computer use lets an AI system operate a user interface through actions such as clicking, typing, scrolling, and reading screenshots. It gives an agent a way to act on software that may not expose a suitable API.
The model selects an interface action, an executor performs it, and the next observation shows the result. Some browser systems also expose structured page information. Screenshot-based desktop control is one approach, not the only way to automate an interface.
Dhruv Batra argues in episode 63 at 2:20 that much of the web will continue to need agents that operate existing interfaces. That is his infrastructure thesis. It helps explain the use case without establishing that interface automation is reliable for every website.
Example: after entering a form, inspect the confirmation and stored result rather than assuming a click completed the transaction. Account permissions still apply, and untrusted page text can contain prompt injection. The practical decision is whether UI access provides needed capability with acceptable reliability and oversight.
Sources
- Claude Platform: Computer use — Documents the action/observation loop and security considerations for interface control.
Go deeper
- Why won't most websites get APIs for AI agents? AI, decoded · Why Most Websites Won't Get APIs for AI Agents
- How does an AI agent decide which tool to use? AI, decoded · How AI Agents Use Tools
- Anthropic: Computer-use demo docs
Trace a working reference implementation from interface actions to new observations.