Adds two opt-in agent tools for controlling a computer:
- computer_use: pixel-based (screenshot + mouse/keyboard) via a desktop
(pyautogui) or browser (playwright) backend.
- browser: DOM/accessibility-based web automation (act by element ref),
reliable across ANY tool-calling model, not just vision/CU-trained ones.
Core enabler in the runner: a tool may return image content blocks, which
are split out and delivered to the model as a follow-up user message
(_split_tool_result_media), so screenshots reach any vision provider
(e.g. via OpenRouter/openai-compat) without provider-specific code.
Both tools are OFF by default (tools.computerUse.enable / tools.browser.enable),
are not exposed to subagents, and the browser tool supports an allowed_domains
allowlist. Heavy deps (pyautogui/pillow/playwright) are an optional [computer-use] extra.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>