Compare commits

..
Author SHA1 Message Date
Xubin Ren 612e714479 feat(webui): add computer use controls 2026-08-09 02:40:17 +09:00
Xubin Ren a739185740 fix(tools): bound computer use runtime state 2026-08-09 01:27:22 +09:00
Xubin Ren 9859e02215 fix(tools): harden computer use integration 2026-08-09 01:12:13 +09:00
95160a304d feat(tools): add model-agnostic computer use (computer_use + browser tools)
Adds two opt-in agent tools for controlling a computer:
- computer_use: pixel-based (screenshot + mouse/keyboard) via a desktop
  (pyautogui) or browser (playwright) backend.
- browser: DOM/accessibility-based web automation (act by element ref),
  reliable across ANY tool-calling model, not just vision/CU-trained ones.

Core enabler in the runner: a tool may return image content blocks, which
are split out and delivered to the model as a follow-up user message
(_split_tool_result_media), so screenshots reach any vision provider
(e.g. via OpenRouter/openai-compat) without provider-specific code.

Both tools are OFF by default (tools.computerUse.enable / tools.browser.enable),
are not exposed to subagents, and the browser tool supports an allowed_domains
allowlist. Heavy deps (pyautogui/pillow/playwright) are an optional [computer-use] extra.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-09 01:12:13 +09:00
44 changed files with 2773 additions and 8 deletions
+2 -1
View File
@@ -125,6 +125,7 @@ Important files:
| Shell execution | `nanobot/agent/tools/shell.py` |
| Filesystem tools | `nanobot/agent/tools/filesystem.py` |
| Web search/fetch | `nanobot/agent/tools/web.py` |
| Browser and computer use | `nanobot/agent/tools/browser_tool.py`, `nanobot/agent/tools/computer_use.py` |
| MCP tools | `nanobot/agent/tools/mcp.py` |
| Cron | `nanobot/agent/tools/cron.py`, `nanobot/cron/` |
| Image generation | `nanobot/agent/tools/image_generation.py` |
@@ -188,7 +189,7 @@ Security-sensitive code paths include:
|---|---|
| Workspace scope | `nanobot/security/workspace_access.py`, `nanobot/security/workspace_policy.py` |
| Shell sandboxing | `nanobot/agent/tools/shell.py` |
| SSRF/network checks | `nanobot/security/network.py`, `nanobot/agent/tools/web.py` |
| SSRF/network checks | `nanobot/security/network.py`, `nanobot/agent/tools/web.py`, `nanobot/agent/tools/computer_use_backends/browser_playwright.py` |
| PTH guard and CLI startup security | `nanobot/security/` and CLI entrypoints |
| Channel access control | channel config in `nanobot/channels/*.py` |
+57
View File
@@ -42,6 +42,7 @@ the focused guides first and come back here for exact fields and defaults.
| Add fallback chains | [Model Fallbacks](#model-fallbacks) |
| Configure voice transcription | [Transcription Settings](#transcription-settings) |
| Tune channel defaults | [Channel Settings](#channel-settings) |
| Enable browser or desktop control | [Browser and Computer Use](#browser-and-computer-use) |
| Configure web search and fetch | [Web Tools](#web-tools) |
| Enable image generation | [Image Generation](#image-generation) |
| Add MCP servers | [MCP](#mcp-model-context-protocol) |
@@ -1670,6 +1671,62 @@ When a channel `send()` raises, nanobot retries at the channel-manager layer. By
>
> If a channel is completely unreachable, nanobot cannot notify the user through that same channel. Watch logs for `Failed to send to {channel} after N attempts` to spot persistent delivery failures.
## Browser and Computer Use
Browser and desktop control are optional and disabled by default. Install their runtime first:
```bash
pip install 'nanobot-ai[computer-use]'
playwright install chromium
```
For normal web interaction, prefer the DOM-based `browser` tool. It gives the model numbered
element references and works without vision. Use `computer_use` when the model must see and act
on pixels; its `desktop` backend controls the real local machine, while its `browser` backend
controls an isolated Playwright page.
```json
{
"tools": {
"browser": {
"enable": true,
"allowedDomains": ["example.com"]
},
"computerUse": {
"enable": false,
"backend": "desktop"
}
}
}
```
| Option | Default | Description |
|---|---|---|
| `tools.browser.enable` | `false` | Register the DOM-based `browser` tool |
| `tools.browser.allowedDomains` | `[]` | Optional top-level navigation allowlist; entries include subdomains |
| `tools.browser.includeScreenshot` | `false` | Attach a screenshot after browser actions |
| `tools.browser.maxSessions` | `8` | Maximum retained browser sessions; least-recently-used state is closed first |
| `tools.computerUse.enable` | `false` | Register pixel-based `computer_use` |
| `tools.computerUse.backend` | `"desktop"` | `"desktop"` or `"browser"` |
| `tools.computerUse.allowedDomains` | `[]` | Navigation allowlist for the browser backend |
| `tools.computerUse.targetWidth` / `targetHeight` | `1280` / `800` | Maximum screenshot dimensions exposed to the model |
| `tools.computerUse.maxSessions` | `8` | Maximum retained sessions for the browser backend |
Each nanobot session gets separate browser state. Browser HTTP and WebSocket traffic passes
through the shared SSRF policy; local, private, link-local, and metadata targets are blocked
unless explicitly permitted with `tools.ssrfWhitelist`. When `maxSessions` is reached, the
least-recently-used browser state is closed. `file:` URLs are not accepted.
> [!IMPORTANT]
> Browser URL checks are defense in depth, not an egress sandbox: Chromium performs its own DNS
> resolution after validation. Use OS/container network isolation when browsing hostile pages.
> [!WARNING]
> The desktop backend can click, type, and change state outside the workspace. Enabling it is an
> explicit trust decision: use a trusted model and input source, and run nanobot in a disposable
> OS account or VM when unattended. The workspace restriction is not an OS sandbox. Desktop text
> input supports ASCII key events; use the browser backend when Unicode text input is required.
## Web Tools
nanobot incorporates basic tools for accessing the web. These include searching via APIs, and fetching arbitrary web pages in Markdown format. They are enabled by default, and can be configured in `~/.nanobot/config.json` under `tools.web`.
+44
View File
@@ -33,6 +33,11 @@ COMPACTABLE_TOOLS = frozenset({
"read_file", "exec", "grep", "find_files",
"web_search", "web_fetch", "list_dir", "list_exec_sessions",
})
VISUAL_TOOLS = frozenset({"browser", "computer_use"})
STALE_SCREENSHOT_PLACEHOLDER = {
"type": "text",
"text": "[Earlier screenshot omitted; use the latest screenshot from this tool.]",
}
# read_file is the recovery path for persisted results; exempting it prevents persist->read->persist loops.
TOOL_RESULT_OFFLOAD_EXEMPT_TOOLS = frozenset({"read_file"})
BACKFILL_CONTENT = "[Tool result unavailable — call was interrupted or lost]"
@@ -41,6 +46,12 @@ PLACEHOLDER_TEXTS = frozenset({
})
def _is_image_block(value: object) -> bool:
if not isinstance(value, dict):
return False
return cast(dict[str, Any], value).get("type") in {"image_url", "input_image"}
def _tool_call_name_is_valid(tool_call: Any) -> bool:
"""Whether a persisted OpenAI-style tool_call carries a usable name.
@@ -84,6 +95,10 @@ class ContextGovernor:
updated = self.drop_orphan_tool_results(updated)
updated = self.backfill_missing_tool_results(updated)
updated = self.apply_tool_result_budget(config, updated)
updated = self.drop_stale_visual_tool_images(
updated,
start_index=config.inflight_start_index,
)
updated = self.compact_inflight_overflow(config, updated, compacted_tool_call_ids)
updated = self.snip_history(config, updated)
updated = self.drop_orphan_tool_results(updated)
@@ -326,6 +341,35 @@ class ContextGovernor:
updated[idx]["content"] = normalized
return updated
@staticmethod
def drop_stale_visual_tool_images(
messages: list[dict[str, Any]],
*,
start_index: int,
) -> list[dict[str, Any]]:
"""Keep only the latest in-flight screenshot from each visual tool."""
seen: set[str] = set()
updated = messages
for idx in range(len(messages) - 1, start_index - 1, -1):
message = messages[idx]
name = str(message.get("name") or "")
content = message.get("content")
if message.get("role") != "tool" or name not in VISUAL_TOOLS:
continue
if not isinstance(content, list):
continue
content_blocks = cast(list[object], content)
blocks = [block for block in content_blocks if not _is_image_block(block)]
if len(blocks) == len(content_blocks):
continue
if name not in seen:
seen.add(name)
continue
if updated is messages:
updated = [dict(item) for item in messages]
updated[idx]["content"] = [dict(STALE_SCREENSHOT_PLACEHOLDER), *blocks]
return updated
def compact_inflight_overflow(
self,
config: ContextGovernanceConfig,
+1
View File
@@ -1397,6 +1397,7 @@ class AgentLoop:
cleanup_steps = (
self.subagents.close,
self._exec_session_manager.close_all,
*(() if not hasattr(self, "tools") else (self.tools.close,)),
lambda: agent_context.close_mcp(self),
)
for cleanup in cleanup_steps:
+4
View File
@@ -220,6 +220,10 @@ class Tool(ABC):
"""Return optional per-turn prompt context owned by this tool."""
return None
async def close(self) -> None:
"""Release resources owned by the tool. Safe to call repeatedly."""
return None
@abstractmethod
async def execute(self, **kwargs: Any) -> Any:
"""Run the tool; return content, or ``ToolResult.error(...)`` for failures."""
+280
View File
@@ -0,0 +1,280 @@
"""DOM-based browser automation by element reference."""
# pyright: reportIncompatibleMethodOverride=false
from __future__ import annotations
import asyncio
from typing import Any
from pydantic import Field
from nanobot.agent.tools.base import Tool, tool_parameters
from nanobot.agent.tools.computer_use_backends.base import SessionBackendPool
from nanobot.agent.tools.context import ToolContext
from nanobot.agent.tools.schema import (
BooleanSchema,
IntegerSchema,
StringSchema,
tool_parameters_schema,
)
from nanobot.config_base import Base
from nanobot.utils.helpers import build_image_content_blocks
_ACTIONS = [
"navigate",
"snapshot",
"click",
"type",
"select",
"scroll",
"key",
"back",
"read_text",
]
class BrowserToolConfig(Base):
"""browser (DOM) tool configuration."""
enable: bool = False
start_url: str = "about:blank"
headless: bool = True
width: int = Field(default=1280, ge=320, le=4096)
height: int = Field(default=800, ge=240, le=4096)
allowed_domains: list[str] = Field(default_factory=list)
include_screenshot: bool = False
max_elements: int = Field(default=200, ge=1, le=1000)
max_sessions: int = Field(default=8, ge=1, le=64)
def _format_elements(elements: list[dict[str, Any]]) -> str:
if not elements:
return "Interactive elements: (none found — try scrolling or read_text)"
lines: list[str] = []
for e in elements:
tag = str(e.get("tag") or "")
typ = str(e.get("type") or "")
label = tag + (f"[{typ}]" if typ else "")
line = f"[{e.get('ref')}] {label}"
name = str(e.get("name") or "").strip()
if name:
line += f' "{name}"'
href = str(e.get("href") or "")
if href and tag == "a":
line += f" -> {href[:60]}"
lines.append(line)
return "Interactive elements (act with the [ref] number):\n" + "\n".join(lines)
@tool_parameters(
tool_parameters_schema(
action=StringSchema("The action to perform.", enum=_ACTIONS),
ref=IntegerSchema(
description="Element ref number from the latest snapshot (click/type/select).",
minimum=1,
nullable=True,
),
text=StringSchema(
"Text to type (action=type) or key/combo like 'Enter'/'ctrl+a' (action=key).",
nullable=True,
),
url=StringSchema("URL to open (action=navigate).", nullable=True),
value=StringSchema("Option value/label to choose (action=select).", nullable=True),
submit=BooleanSchema(description="Press Enter after typing (action=type).", nullable=True),
scroll_direction=StringSchema(
"Scroll direction (action=scroll).", enum=["up", "down", "left", "right"], nullable=True
),
scroll_amount=IntegerSchema(
description="Scroll clicks (action=scroll).",
minimum=1,
maximum=100,
nullable=True,
),
required=["action"],
)
)
class BrowserTool(Tool):
"""Browse and act on web pages by element ref (DOM-based, works with any model)."""
_scopes = {"core"}
name = "browser" # pyright: ignore[reportIncompatibleMethodOverride, reportAssignmentType]
description = ( # pyright: ignore[reportIncompatibleMethodOverride, reportAssignmentType]
"Control a web browser by acting on page elements by their [ref] number. "
"Each call returns the current page URL plus a fresh numbered list of the page's "
"interactive elements; pick a [ref] to click/type/select — no pixel coordinates "
"needed. A page may already be open: call 'snapshot' FIRST to see it. Only use "
"'navigate' for a specific URL you were explicitly given — never guess a URL. "
"Move between pages by clicking links/buttons via their [ref]. Use 'read_text' to "
"read page text. Re-read the element list after each action; refs are reassigned."
)
config_key = "browser"
@classmethod
def config_cls(cls) -> type[BrowserToolConfig]:
return BrowserToolConfig
@classmethod
def enabled(cls, ctx: ToolContext) -> bool:
return bool(ctx.config.browser.enable)
@classmethod
def create(cls, ctx: ToolContext) -> Tool:
return cls(ctx.config.browser)
def __init__(
self,
config: BrowserToolConfig | None = None,
*,
backend_impl: Any = None,
) -> None:
self.config = config or BrowserToolConfig()
runtime = None
if backend_impl is None:
from nanobot.agent.tools.computer_use_backends.browser_playwright import BrowserRuntime
runtime = BrowserRuntime(headless=self.config.headless)
self._runtime = runtime
self._execution_lock = asyncio.Lock()
self._backends = SessionBackendPool(
self._make_backend,
backend_impl,
max_backends=self.config.max_sessions,
finalizer=runtime.close if runtime is not None else None,
)
@property
def read_only(self) -> bool:
return False
@property
def exclusive(self) -> bool:
return True
def _make_backend(self) -> Any:
from nanobot.agent.tools.computer_use_backends.browser_playwright import BrowserBackend
return BrowserBackend(
width=self.config.width,
height=self.config.height,
start_url=self.config.start_url,
allowed_domains=self.config.allowed_domains,
runtime=self._runtime,
)
@staticmethod
def _req_ref(params: dict[str, Any], action: str) -> Any:
ref = params.get("ref")
if ref is None:
raise ValueError(f"action '{action}' requires an element 'ref' from the snapshot")
return ref
async def _dispatch(self, backend: Any, action: str, p: dict[str, Any]) -> tuple[str, str | None]:
"""Return (status, direct_text). If direct_text is set, it is returned as-is
(no snapshot appended)."""
if action == "navigate":
url = p.get("url")
if not url:
raise ValueError("action 'navigate' requires 'url'")
await backend.navigate(str(url))
return f"Navigated to {url}", None
if action == "snapshot":
return "Snapshot of the current page", None
if action == "click":
ref = self._req_ref(p, action)
await backend.click_ref(ref)
return f"Clicked element [{ref}]", None
if action == "type":
ref = self._req_ref(p, action)
text = p.get("text")
if text is None:
raise ValueError("action 'type' requires 'text'")
submit = bool(p.get("submit"))
await backend.fill_ref(ref, str(text), submit=submit)
return f"Typed into [{ref}]" + (" and pressed Enter" if submit else ""), None
if action == "select":
ref = self._req_ref(p, action)
value = p.get("value")
if value is None:
raise ValueError("action 'select' requires 'value'")
await backend.select_ref(ref, str(value))
return f"Selected '{value}' in [{ref}]", None
if action == "scroll":
direction = str(p.get("scroll_direction") or "down").lower()
if direction not in ("up", "down", "left", "right"):
raise ValueError("'scroll_direction' must be up/down/left/right")
await backend.scroll_page(direction, int(p.get("scroll_amount") or 3))
return f"Scrolled {direction}", None
if action == "key":
combo = p.get("text")
if not combo:
raise ValueError("action 'key' requires 'text' (e.g. 'Enter')")
await backend.key(str(combo))
return f"Pressed {combo}", None
if action == "back":
await backend.go_back()
return "Navigated back", None
if action == "read_text":
txt = await backend.read_text()
return "", f"Page text:\n{txt}"
raise ValueError(f"unknown action '{action}'")
async def execute(self, action: str | None = None, **kwargs: Any) -> Any:
async with self._execution_lock:
return await self._execute(action, **kwargs)
async def _execute(self, action: str | None = None, **kwargs: Any) -> Any:
action = (action or "").strip()
if action not in _ACTIONS:
return f"Error: unknown action '{action}'. Valid actions: {', '.join(_ACTIONS)}"
try:
backend = await self._backends.get()
except ImportError as exc:
return f"Error: {exc}"
except Exception as exc:
return f"Error: could not initialize browser backend: {type(exc).__name__}: {exc}"
try:
status, direct = await self._dispatch(backend, action, kwargs)
if blocked := getattr(backend, "pop_blocked_navigation", lambda: None)():
raise ValueError(f"navigation was blocked: {blocked}")
except ValueError as exc:
return f"Error: {exc}"
except Exception as exc:
return f"Error executing browser '{action}': {type(exc).__name__}: {exc}"
if direct is not None:
return direct
try:
elements = await backend.dom_snapshot(self.config.max_elements)
snapshot = _format_elements(elements)
except Exception as exc:
snapshot = f"(could not read page elements: {type(exc).__name__}: {exc})"
try:
current = await backend.current_url()
except Exception:
current = ""
header = f"{status}\nCurrent page: {current}" if current else status
text_out = f"{header}\n\n{snapshot}"
if self.config.include_screenshot:
try:
png = await backend.screenshot()
return build_image_content_blocks(png, "image/png", "", text_out)
except Exception:
return text_out
return text_out
async def close(self) -> None:
await self._backends.close()
+327
View File
@@ -0,0 +1,327 @@
"""Screenshot-based computer control."""
# pyright: reportIncompatibleMethodOverride=false
from __future__ import annotations
import asyncio
import io
from typing import Any, Literal
from pydantic import Field
from nanobot.agent.tools.base import Tool, tool_parameters
from nanobot.agent.tools.computer_use_backends.base import SessionBackendPool
from nanobot.agent.tools.context import ToolContext
from nanobot.agent.tools.schema import (
IntegerSchema,
NumberSchema,
StringSchema,
tool_parameters_schema,
)
from nanobot.config_base import Base
from nanobot.utils.helpers import build_image_content_blocks
_ACTIONS = [
"screenshot",
"left_click",
"right_click",
"middle_click",
"double_click",
"triple_click",
"mouse_move",
"left_click_drag",
"scroll",
"type",
"key",
"wait",
"navigate",
]
_CLICK_BUTTONS = {
"left_click": "left",
"double_click": "left",
"triple_click": "left",
"right_click": "right",
"middle_click": "middle",
}
_CLICK_COUNTS = {"double_click": 2, "triple_click": 3}
_MAX_WAIT_S = 10.0
class ComputerUseToolConfig(Base):
"""computer_use tool configuration."""
enable: bool = False
backend: Literal["desktop", "browser"] = "desktop"
target_width: int = Field(default=1280, ge=320, le=4096)
target_height: int = Field(default=800, ge=240, le=4096)
allowed_domains: list[str] = Field(default_factory=list)
start_url: str = "about:blank"
headless: bool = True
max_sessions: int = Field(default=8, ge=1, le=64)
def _fit_size(width: int, height: int, max_width: int, max_height: int) -> tuple[int, int]:
if width <= 0 or height <= 0:
return max(1, max_width), max(1, max_height)
scale = min(max_width / width, max_height / height, 1.0)
return max(1, round(width * scale)), max(1, round(height * scale))
def _scale_point(
x: int,
y: int,
source: tuple[int, int],
target: tuple[int, int],
) -> tuple[int, int]:
width, height = source
target_width, target_height = target
real_x = round(x * width / target_width) if target_width else x
real_y = round(y * height / target_height) if target_height else y
return (
max(0, min(real_x, max(0, width - 1))),
max(0, min(real_y, max(0, height - 1))),
)
@tool_parameters(
tool_parameters_schema(
action=StringSchema("The action to perform.", enum=_ACTIONS),
x=IntegerSchema(
description="X coordinate in the pixel space of the screenshot you were last shown.",
nullable=True,
),
y=IntegerSchema(
description="Y coordinate in the pixel space of the screenshot you were last shown.",
nullable=True,
),
text=StringSchema(
"Text to type (action=type; desktop supports ASCII), or a key/combo like "
"'ctrl+s' or 'Enter' (action=key).",
nullable=True,
),
scroll_direction=StringSchema(
"Scroll direction (action=scroll).", enum=["up", "down", "left", "right"], nullable=True
),
scroll_amount=IntegerSchema(
description="Number of scroll clicks (action=scroll).",
minimum=1,
maximum=100,
nullable=True,
),
duration=NumberSchema(
description="Seconds to wait (action=wait).",
minimum=0,
maximum=_MAX_WAIT_S,
nullable=True,
),
url=StringSchema("URL to open (action=navigate, browser backend only).", nullable=True),
required=["action"],
)
)
class ComputerUseTool(Tool):
"""Control a computer (desktop or browser) by looking at screenshots and acting."""
_scopes = {"core"} # never exposed to subagents — security-sensitive
name = "computer_use" # pyright: ignore[reportIncompatibleMethodOverride, reportAssignmentType]
description = ( # pyright: ignore[reportIncompatibleMethodOverride, reportAssignmentType]
"Control a computer via screenshots and mouse/keyboard. Each call performs ONE "
"action and returns a fresh screenshot of the resulting screen. Coordinates (x, y) "
"are in the pixel space of the screenshot you were last shown (top-left is 0,0). "
"The 'browser' backend additionally supports the 'navigate' action. Always start "
"with a 'screenshot' to see the screen, then act based on what you observe; after "
"each action re-check the new screenshot before the next step."
)
config_key = "computer_use"
@classmethod
def config_cls(cls) -> type[ComputerUseToolConfig]:
return ComputerUseToolConfig
@classmethod
def enabled(cls, ctx: ToolContext) -> bool:
return bool(ctx.config.computer_use.enable)
@classmethod
def create(cls, ctx: ToolContext) -> Tool:
return cls(ctx.config.computer_use)
def __init__(
self,
config: ComputerUseToolConfig | None = None,
*,
backend_impl: Any = None,
) -> None:
self.config = config or ComputerUseToolConfig()
runtime = None
if backend_impl is None and self.config.backend == "browser":
from nanobot.agent.tools.computer_use_backends.browser_playwright import BrowserRuntime
runtime = BrowserRuntime(headless=self.config.headless)
self._runtime = runtime
self._execution_lock = asyncio.Lock()
self._backends = SessionBackendPool(
self._make_backend,
backend_impl,
max_backends=1 if self.config.backend == "desktop" else self.config.max_sessions,
finalizer=runtime.close if runtime is not None else None,
)
@property
def read_only(self) -> bool:
return False
@property
def exclusive(self) -> bool:
# Stateful single environment; must not run alongside other tools.
return True
def _make_backend(self) -> Any:
if self.config.backend == "browser":
from nanobot.agent.tools.computer_use_backends.browser_playwright import BrowserBackend
return BrowserBackend(
width=self.config.target_width,
height=self.config.target_height,
start_url=self.config.start_url,
allowed_domains=self.config.allowed_domains,
runtime=self._runtime,
)
from nanobot.agent.tools.computer_use_backends.desktop_pyautogui import DesktopBackend
return DesktopBackend()
@staticmethod
def _downscale_png(png: bytes, target: tuple[int, int]) -> bytes:
try:
from PIL import Image # noqa: PLC0415
except Exception as exc:
raise ImportError(
"Pillow is required for computer_use. Install: pip install 'nanobot-ai[computer-use]'"
) from exc
tw, th = target
with Image.open(io.BytesIO(png)) as img:
if (img.width, img.height) == (tw, th):
return png
resized = img.convert("RGB").resize((tw, th)) # pyright: ignore[reportUnknownMemberType]
out = io.BytesIO()
resized.save(out, format="PNG")
return out.getvalue()
async def _dispatch(
self,
backend: Any,
action: str,
params: dict[str, Any],
source: tuple[int, int],
target: tuple[int, int],
) -> str:
def _xy() -> tuple[int, int]:
x, y = params.get("x"), params.get("y")
if x is None or y is None:
raise ValueError(f"action '{action}' requires integer 'x' and 'y'")
return _scale_point(int(x), int(y), source, target)
if action == "screenshot":
return "Took a screenshot"
if action == "wait":
duration = params.get("duration")
secs = 1.0 if duration is None else float(duration)
secs = max(0.0, min(secs, _MAX_WAIT_S))
await asyncio.sleep(secs)
return f"Waited {secs:g}s"
if action in _CLICK_BUTTONS:
rx, ry = _xy()
await backend.click(rx, ry, _CLICK_BUTTONS[action], _CLICK_COUNTS.get(action, 1))
return f"{action} at ({rx}, {ry})"
if action == "mouse_move":
rx, ry = _xy()
await backend.move(rx, ry)
return f"Moved to ({rx}, {ry})"
if action == "left_click_drag":
rx, ry = _xy()
await backend.drag(rx, ry)
return f"Dragged to ({rx}, {ry})"
if action == "scroll":
rx, ry = _xy()
direction = str(params.get("scroll_direction") or "down").lower()
if direction not in ("up", "down", "left", "right"):
raise ValueError("'scroll_direction' must be up/down/left/right")
amount = int(params.get("scroll_amount") or 3)
await backend.scroll(rx, ry, direction, amount)
return f"Scrolled {direction} by {amount} at ({rx}, {ry})"
if action == "type":
text = params.get("text")
if not text:
raise ValueError("action 'type' requires 'text'")
await backend.type_text(str(text))
return f"Typed {len(str(text))} characters"
if action == "key":
combo = params.get("text")
if not combo:
raise ValueError("action 'key' requires 'text' (e.g. 'ctrl+s')")
await backend.key(str(combo))
return f"Pressed {combo}"
if action == "navigate":
url = params.get("url")
if not url:
raise ValueError("action 'navigate' requires 'url'")
await backend.navigate(str(url))
return f"Navigated to {url}"
raise ValueError(f"unknown action '{action}'")
async def execute(self, action: str | None = None, **kwargs: Any) -> Any:
async with self._execution_lock:
return await self._execute(action, **kwargs)
async def _execute(self, action: str | None = None, **kwargs: Any) -> Any:
action = (action or "").strip()
if action not in _ACTIONS:
return f"Error: unknown action '{action}'. Valid actions: {', '.join(_ACTIONS)}"
try:
backend = await self._backends.get()
real_w, real_h = await backend.dimensions()
except ImportError as exc:
return f"Error: {exc}"
except Exception as exc:
return f"Error: could not initialize computer_use backend: {type(exc).__name__}: {exc}"
source = (real_w, real_h)
target = _fit_size(real_w, real_h, self.config.target_width, self.config.target_height)
try:
status = await self._dispatch(backend, action, kwargs, source, target)
if blocked := getattr(backend, "pop_blocked_navigation", lambda: None)():
raise ValueError(f"navigation was blocked: {blocked}")
except ValueError as exc:
return f"Error: {exc}"
except NotImplementedError as exc:
return f"Error: {exc}"
except Exception as exc:
return f"Error executing computer_use '{action}': {type(exc).__name__}: {exc}"
# Return a fresh screenshot so the model sees the result of its action.
try:
png = await backend.screenshot()
png = self._downscale_png(png, target)
except ImportError as exc:
return f"Error: {exc}"
except Exception as exc:
return f"{status}\n(Could not capture screenshot: {type(exc).__name__}: {exc})"
label = f"{status} | screen {target[0]}x{target[1]} ({backend.environment})"
return build_image_content_blocks(png, "image/png", "", label)
async def close(self) -> None:
await self._backends.close()
@@ -0,0 +1 @@
"""Computer-use backend adapters."""
@@ -0,0 +1,139 @@
"""Backend interface for the ``computer_use`` tool.
A backend is the *actuator* + *screenshot source* for one execution environment
(the local desktop, a headless browser, a VM, ...). The tool layer owns the
agent loop, coordinate scaling, screenshot downscaling and safety gating; a
backend only has to perform primitive actions and grab a screenshot.
Coordinate contract: every ``x``/``y`` passed to a backend is already in **real
device pixels** (the same pixel space as :meth:`screenshot`). The tool scales the
model's target-space coordinates to real pixels before calling the backend, so
backends never deal with the downscaled space.
"""
from __future__ import annotations
import asyncio
from abc import ABC, abstractmethod
from collections import OrderedDict
from collections.abc import Awaitable, Callable
from typing import Any
from nanobot.agent.tools.context import current_request_session_key
class ComputerBackend(ABC):
"""Primitive GUI actions + screenshot for one execution environment."""
#: "desktop" or "browser" — surfaced to the model so it knows the context.
environment: str = "desktop"
@abstractmethod
async def dimensions(self) -> tuple[int, int]:
"""Return the real screenshot pixel size as ``(width, height)``."""
@abstractmethod
async def screenshot(self) -> bytes:
"""Return a PNG screenshot of the current screen at real pixel size."""
@abstractmethod
async def click(self, x: int, y: int, button: str = "left", count: int = 1) -> None:
"""Click at ``(x, y)``. ``button`` in {left,right,middle}; ``count`` for double/triple."""
@abstractmethod
async def move(self, x: int, y: int) -> None:
"""Move the cursor to ``(x, y)`` without clicking."""
@abstractmethod
async def drag(self, x: int, y: int) -> None:
"""Press at the current cursor position and drag to ``(x, y)``, then release."""
@abstractmethod
async def scroll(self, x: int, y: int, direction: str, amount: int) -> None:
"""Scroll at ``(x, y)``. ``direction`` in {up,down,left,right}; ``amount`` in clicks."""
@abstractmethod
async def type_text(self, text: str) -> None:
"""Type ``text`` at the current focus."""
@abstractmethod
async def key(self, combo: str) -> None:
"""Press a key or combo, e.g. ``"ctrl+s"`` / ``"Enter"`` (backend-specific syntax)."""
async def navigate(self, url: str) -> None:
"""Navigate to ``url`` (browser backends only)."""
raise NotImplementedError(
f"'navigate' is not supported by the {self.environment} backend"
)
async def close(self) -> None:
"""Release any resources (browser process, etc.). Safe to call repeatedly."""
return None
class SessionBackendPool:
"""Keep stateful backends isolated by nanobot session."""
def __init__(
self,
factory: Callable[[], Any],
injected: Any = None,
*,
max_backends: int = 8,
finalizer: Callable[[], Awaitable[None]] | None = None,
) -> None:
if max_backends < 1:
raise ValueError("max_backends must be at least 1")
self._factory = factory
self._injected = injected
self._max_backends = max_backends
self._finalizer = finalizer
self._backends: OrderedDict[str, Any] = OrderedDict()
self._lock = asyncio.Lock()
self._closed = False
async def get(self) -> Any:
async with self._lock:
if self._closed:
raise RuntimeError("computer-use backend pool is closed")
if self._injected is not None:
return self._injected
key = current_request_session_key() or "default"
backend = self._backends.get(key)
if backend is not None:
self._backends.move_to_end(key)
return backend
if len(self._backends) >= self._max_backends:
_, stale = self._backends.popitem(last=False)
await stale.close()
backend = self._factory()
self._backends[key] = backend
return backend
async def close(self) -> None:
async with self._lock:
if self._closed:
return
self._closed = True
backends = (
[self._injected]
if self._injected is not None
else list(self._backends.values())
)
self._injected = None
self._backends.clear()
finalizer, self._finalizer = self._finalizer, None
results = await asyncio.gather(
*(backend.close() for backend in backends if backend is not None),
return_exceptions=True,
)
errors = [result for result in results if isinstance(result, BaseException)]
if finalizer is not None:
try:
await finalizer()
except BaseException as exc:
errors.append(exc)
if len(errors) == 1:
raise errors[0]
if errors:
raise BaseExceptionGroup("failed to close computer-use backends", errors)
@@ -0,0 +1,382 @@
"""Playwright backend shared by browser and computer_use."""
from __future__ import annotations
import asyncio
import importlib
from collections.abc import Sequence
from typing import Any, cast
from urllib.parse import urlparse, urlunparse
from loguru import logger
from nanobot.agent.tools.computer_use_backends.base import ComputerBackend
from nanobot.security.network import validate_url_target
_MISSING = (
"Browser computer-use backend needs 'playwright'. Install with: "
"pip install 'nanobot-ai[computer-use]' && playwright install chromium"
)
_SCROLL_PIXELS = 100 # one "scroll click" ~= this many pixels
# Tags visible interactive elements with data-nanobot-ref and returns a compact
# list. Refs are reassigned per call. Used by DOM/accessibility mode.
_SNAPSHOT_JS = r"""
(max) => {
const SEL = 'a,button,input,textarea,select,[role=button],[role=link],[role=checkbox],[role=radio],[role=tab],[role=menuitem],[role=switch],[onclick],[contenteditable=""],[contenteditable=true]';
const out = [];
let ref = 0;
for (const el of document.querySelectorAll(SEL)) {
const r = el.getBoundingClientRect();
const s = getComputedStyle(el);
if (r.width <= 0 || r.height <= 0) continue;
if (s.visibility === 'hidden' || s.display === 'none' || s.opacity === '0') continue;
ref++;
el.setAttribute('data-nanobot-ref', String(ref));
let name = (el.getAttribute('aria-label') || el.innerText || el.value ||
el.getAttribute('placeholder') || el.getAttribute('name') ||
el.getAttribute('title') || '');
name = name.replace(/\s+/g, ' ').trim().slice(0, 120);
out.push({
ref: ref,
tag: el.tagName.toLowerCase(),
role: el.getAttribute('role') || '',
type: el.getAttribute('type') || '',
name: name,
href: el.getAttribute('href') || ''
});
if (out.length >= max) break;
}
return out;
}
"""
# CUA/xdotool-ish modifier names -> Playwright modifiers.
_MODIFIERS = {
"ctrl": "Control", "control": "Control",
"alt": "Alt", "option": "Alt",
"shift": "Shift",
"cmd": "Meta", "meta": "Meta", "super": "Meta", "win": "Meta",
}
# Common single-key names -> Playwright key names.
_KEYS = {
"return": "Enter", "enter": "Enter", "tab": "Tab", "esc": "Escape",
"escape": "Escape", "backspace": "Backspace", "delete": "Delete",
"space": "Space", "up": "ArrowUp", "down": "ArrowDown",
"left": "ArrowLeft", "right": "ArrowRight",
"page_down": "PageDown", "pagedown": "PageDown",
"page_up": "PageUp", "pageup": "PageUp", "home": "Home", "end": "End",
}
def _validate_browser_url(
url: str,
allowed_domains: Sequence[str] = (),
*,
navigation: bool = True,
) -> tuple[bool, str]:
if url == "about:blank":
return True, ""
parsed = urlparse(url)
if not navigation and parsed.scheme in {"blob", "data"}:
return True, ""
target = url
if parsed.scheme in {"ws", "wss"}:
target = urlunparse(parsed._replace(scheme="https" if parsed.scheme == "wss" else "http"))
if navigation and allowed_domains:
host = (parsed.hostname or "").rstrip(".").lower()
allowed = any(
normalized and (host == normalized or host.endswith(f".{normalized}"))
for domain in allowed_domains
if (normalized := domain.strip().lstrip(".").rstrip(".").lower())
)
if not allowed:
return False, f"host {host or '<missing>'} is not in allowed_domains"
return validate_url_target(target)
def _playwright_key(combo: str) -> str:
parts = [p.strip() for p in combo.split("+") if p.strip()]
out: list[str] = []
for part in parts:
low = part.lower()
if low in _MODIFIERS:
out.append(_MODIFIERS[low])
elif low in _KEYS:
out.append(_KEYS[low])
elif len(part) == 1:
out.append(part)
else:
out.append(part.capitalize())
return "+".join(out)
class BrowserRuntime:
"""One lazily started browser process shared by isolated session contexts."""
def __init__(self, *, headless: bool = True) -> None:
self._headless = headless
self._lock = asyncio.Lock()
self._playwright: Any = None
self._browser: Any = None
async def get(self) -> Any:
if self._browser is not None:
return self._browser
async with self._lock:
if self._browser is not None:
return self._browser
try:
playwright = importlib.import_module("playwright.async_api")
async_playwright = cast(Any, playwright).async_playwright
except ImportError as exc:
raise ImportError(_MISSING) from exc
self._playwright = await async_playwright().start()
try:
self._browser = await self._playwright.chromium.launch(
headless=self._headless
)
except BaseException:
await self.close()
raise
return self._browser
async def close(self) -> None:
browser, playwright = self._browser, self._playwright
self._browser = self._playwright = None
errors: list[BaseException] = []
closers = (
browser.close if browser is not None else None,
playwright.stop if playwright is not None else None,
)
for close in closers:
if close is None:
continue
try:
await close()
except BaseException as exc:
errors.append(exc)
if len(errors) == 1:
raise errors[0]
if errors:
raise BaseExceptionGroup("failed to close browser runtime", errors)
class BrowserBackend(ComputerBackend):
environment = "browser"
def __init__(
self,
*,
width: int = 1280,
height: int = 800,
headless: bool = True,
start_url: str = "about:blank",
allowed_domains: Sequence[str] = (),
runtime: BrowserRuntime | None = None,
) -> None:
self._width = width
self._height = height
self._start_url = start_url
self._allowed_domains = tuple(allowed_domains)
self._runtime = runtime or BrowserRuntime(headless=headless)
self._owns_runtime = runtime is None
self._context: Any = None
self._page: Any = None
self._last_pos = (0, 0)
self._blocked_navigation: str | None = None
async def _require_url(self, url: str, label: str) -> None:
ok, error = await asyncio.to_thread(
_validate_browser_url,
url,
self._allowed_domains,
)
if not ok:
raise ValueError(f"{label} is blocked: {error}")
async def _route_request(self, route: Any) -> None:
request = route.request
navigation = bool(request.is_navigation_request())
ok, error = await asyncio.to_thread(
_validate_browser_url,
request.url,
self._allowed_domains,
navigation=navigation,
)
if ok:
await route.continue_()
return
if navigation:
self._blocked_navigation = error
logger.warning("Blocked browser request to {}: {}", request.url, error)
await route.abort("blockedbyclient")
async def _route_web_socket(self, web_socket: Any) -> None:
ok, error = await asyncio.to_thread(
_validate_browser_url,
web_socket.url,
self._allowed_domains,
navigation=False,
)
if not ok:
logger.warning("Blocked browser WebSocket to {}: {}", web_socket.url, error)
await web_socket.close(code=1008, reason="Blocked by nanobot network policy")
return
await web_socket.connect_to_server()
def pop_blocked_navigation(self) -> str | None:
error = self._blocked_navigation
self._blocked_navigation = None
return error
async def _ensure(self) -> Any:
if self._page is not None:
return self._page
await self._require_url(self._start_url, "start_url")
try:
browser = await self._runtime.get()
self._context = await browser.new_context(
viewport={"width": self._width, "height": self._height},
device_scale_factor=1,
service_workers="block",
)
await self._context.route("**/*", self._route_request)
await self._context.route_web_socket("**/*", self._route_web_socket)
self._page = await self._context.new_page()
if self._start_url != "about:blank":
await self._page.goto(self._start_url)
return self._page
except BaseException:
await self.close()
raise
async def dimensions(self) -> tuple[int, int]:
await self._ensure()
vp = self._page.viewport_size or {"width": self._width, "height": self._height}
return vp["width"], vp["height"]
async def screenshot(self) -> bytes:
page = await self._ensure()
return await page.screenshot()
async def click(self, x: int, y: int, button: str = "left", count: int = 1) -> None:
page = await self._ensure()
await page.mouse.click(x, y, button=button, click_count=count)
self._last_pos = (x, y)
async def move(self, x: int, y: int) -> None:
page = await self._ensure()
await page.mouse.move(x, y)
self._last_pos = (x, y)
async def drag(self, x: int, y: int) -> None:
page = await self._ensure()
sx, sy = self._last_pos
await page.mouse.move(sx, sy)
await page.mouse.down()
await page.mouse.move(x, y)
await page.mouse.up()
self._last_pos = (x, y)
async def scroll(self, x: int, y: int, direction: str, amount: int) -> None:
page = await self._ensure()
await page.mouse.move(x, y)
pixels = max(1, amount) * _SCROLL_PIXELS
dx = pixels if direction == "right" else -pixels if direction == "left" else 0
dy = pixels if direction == "down" else -pixels if direction == "up" else 0
await page.mouse.wheel(dx, dy)
async def type_text(self, text: str) -> None:
page = await self._ensure()
await page.keyboard.type(text)
async def key(self, combo: str) -> None:
page = await self._ensure()
key = _playwright_key(combo)
if key:
await page.keyboard.press(key)
async def navigate(self, url: str) -> None:
await self._require_url(url, "navigation")
page = await self._ensure()
await page.goto(url)
self._last_pos = (0, 0)
# --- DOM / accessibility mode (act by element ref, not pixels) ---
async def dom_snapshot(self, max_elements: int = 200) -> list[dict[str, Any]]:
"""Tag visible interactive elements with ``data-nanobot-ref`` and return them.
Each entry: ``{ref, tag, role, type, name, href}``. Refs are reassigned on
every snapshot, so callers should act on the latest snapshot.
"""
page = await self._ensure()
return cast(list[dict[str, Any]], await page.evaluate(_SNAPSHOT_JS, max_elements))
def _ref_selector(self, ref: int) -> str:
return f'[data-nanobot-ref="{int(ref)}"]'
async def click_ref(self, ref: int) -> None:
page = await self._ensure()
await page.click(self._ref_selector(ref), timeout=5000)
async def fill_ref(self, ref: int, text: str, submit: bool = False) -> None:
page = await self._ensure()
sel = self._ref_selector(ref)
await page.fill(sel, text, timeout=5000)
if submit:
await page.press(sel, "Enter")
async def select_ref(self, ref: int, value: str) -> None:
page = await self._ensure()
sel = self._ref_selector(ref)
try:
await page.select_option(sel, value, timeout=3000)
except Exception:
# Models usually pass the visible label, not the option value.
await page.select_option(sel, label=value, timeout=3000)
async def scroll_page(self, direction: str, amount: int) -> None:
page = await self._ensure()
pixels = max(1, amount) * _SCROLL_PIXELS
dx = pixels if direction == "right" else -pixels if direction == "left" else 0
dy = pixels if direction == "down" else -pixels if direction == "up" else 0
await page.evaluate("([x, y]) => window.scrollBy(x, y)", [dx, dy])
async def go_back(self) -> None:
page = await self._ensure()
await page.go_back()
async def read_text(self, max_chars: int = 4000) -> str:
page = await self._ensure()
txt = await page.evaluate("() => document.body ? document.body.innerText : ''")
return (txt or "")[:max_chars]
async def current_url(self) -> str:
page = await self._ensure()
return page.url
async def close(self) -> None:
context = self._context
self._context = self._page = None
error: BaseException | None = None
if context is not None:
try:
await context.close()
except BaseException as exc:
error = exc
if self._owns_runtime:
try:
await self._runtime.close()
except BaseException as exc:
if error is not None:
raise BaseExceptionGroup("failed to close browser backend", [error, exc])
raise
if error is not None:
raise error
@@ -0,0 +1,134 @@
"""PyAutoGUI desktop backend with HiDPI coordinate correction."""
from __future__ import annotations
import asyncio
import io
from typing import Any
from nanobot.agent.tools.computer_use_backends.base import ComputerBackend
_MISSING = (
"Desktop computer-use backend needs 'pyautogui' and 'pillow'. "
"Install with: pip install 'nanobot-ai[computer-use]'"
)
# xdotool/CUA-style key names -> PyAutoGUI key names.
_KEY_ALIASES = {
"return": "enter",
"ctrl": "ctrl",
"control": "ctrl",
"cmd": "command",
"super": "win",
"win": "win",
"page_down": "pagedown",
"page_up": "pageup",
"pagedown": "pagedown",
"pageup": "pageup",
"esc": "esc",
"escape": "esc",
}
class DesktopBackend(ComputerBackend):
environment = "desktop"
def __init__(self) -> None:
self._pg: Any = None
self._ratio_x = 1.0
self._ratio_y = 1.0
self._dims: tuple[int, int] | None = None
def _ensure(self) -> Any:
if self._pg is not None:
return self._pg
try:
import pyautogui # noqa: PLC0415
except Exception as exc: # ImportError, or platform display errors
raise ImportError(_MISSING) from exc
self._pg = pyautogui
return pyautogui
def _grab_png_and_size(self) -> tuple[bytes, int, int]:
pg = self._ensure()
img = pg.screenshot()
buf = io.BytesIO()
img.save(buf, format="PNG")
width, height = img.size
# Refresh logical<->physical ratio from the actual grab.
try:
logical_w, logical_h = pg.size()
self._ratio_x = (logical_w / width) if width else 1.0
self._ratio_y = (logical_h / height) if height else 1.0
except Exception:
self._ratio_x = self._ratio_y = 1.0
self._dims = (width, height)
return buf.getvalue(), width, height
def _to_logical(self, x: int, y: int) -> tuple[int, int]:
return round(x * self._ratio_x), round(y * self._ratio_y)
async def dimensions(self) -> tuple[int, int]:
if self._dims is not None:
return self._dims
_, w, h = await asyncio.to_thread(self._grab_png_and_size)
return w, h
async def screenshot(self) -> bytes:
png, _, _ = await asyncio.to_thread(self._grab_png_and_size)
return png
async def click(self, x: int, y: int, button: str = "left", count: int = 1) -> None:
pg = self._ensure()
lx, ly = self._to_logical(x, y)
await asyncio.to_thread(pg.click, lx, ly, clicks=count, button=button)
async def move(self, x: int, y: int) -> None:
pg = self._ensure()
lx, ly = self._to_logical(x, y)
await asyncio.to_thread(pg.moveTo, lx, ly)
async def drag(self, x: int, y: int) -> None:
pg = self._ensure()
lx, ly = self._to_logical(x, y)
await asyncio.to_thread(
pg.dragTo,
lx,
ly,
duration=0.3,
tween=pg.easeInOutQuad,
button="left",
)
async def scroll(self, x: int, y: int, direction: str, amount: int) -> None:
pg = self._ensure()
lx, ly = self._to_logical(x, y)
clicks = max(1, amount)
await asyncio.to_thread(pg.moveTo, lx, ly)
if direction in ("up", "down"):
await asyncio.to_thread(pg.scroll, clicks if direction == "up" else -clicks)
else:
await asyncio.to_thread(pg.hscroll, clicks if direction == "right" else -clicks)
async def type_text(self, text: str) -> None:
if not text.isascii():
raise ValueError(
"desktop text input supports ASCII key events only; "
"use the browser backend for Unicode text"
)
pg = self._ensure()
await asyncio.to_thread(pg.typewrite, text, 0.01)
async def key(self, combo: str) -> None:
pg = self._ensure()
keys = [
_KEY_ALIASES.get(part.strip().lower(), part.strip().lower())
for part in combo.split("+")
if part.strip()
]
if not keys:
return
if len(keys) == 1:
await asyncio.to_thread(pg.press, keys[0])
else:
await asyncio.to_thread(pg.hotkey, *keys)
+3
View File
@@ -187,5 +187,8 @@ class _LegacyErrorPrefixTool(Tool):
return ToolResult.error(result)
return result
async def close(self) -> None:
await self._wrapped.close()
def __getattr__(self, name: str) -> Any:
return getattr(self._wrapped, name)
+13
View File
@@ -200,6 +200,19 @@ class ToolRegistry:
except Exception as e:
return ToolResult.error(f"Error executing {name}: {str(e)}" + hint)
async def close(self) -> None:
"""Close every registered tool, attempting all cleanups."""
errors: list[BaseException] = []
for tool in self._tools.values():
try:
await tool.close()
except BaseException as exc:
errors.append(exc)
if len(errors) == 1:
raise errors[0]
if errors:
raise BaseExceptionGroup("failed to close tools", errors)
@property
def tool_names(self) -> list[str]:
"""Get list of registered tool names."""
+16
View File
@@ -12,7 +12,9 @@ from nanobot.config_base import Base
from nanobot.cron.types import CronSchedule
if TYPE_CHECKING:
from nanobot.agent.tools.browser_tool import BrowserToolConfig
from nanobot.agent.tools.cli_apps import CliAppsToolConfig
from nanobot.agent.tools.computer_use import ComputerUseToolConfig
from nanobot.agent.tools.filesystem import FileToolsConfig
from nanobot.agent.tools.image_generation import ImageGenerationToolConfig
from nanobot.agent.tools.self import MyToolConfig
@@ -399,6 +401,16 @@ class ToolsConfig(Base):
"""
web: WebToolsConfig = Field(default_factory=lambda: _lazy_default("nanobot.agent.tools.web", "WebToolsConfig"))
browser: BrowserToolConfig = Field(
default_factory=lambda: _lazy_default(
"nanobot.agent.tools.browser_tool", "BrowserToolConfig"
)
)
computer_use: ComputerUseToolConfig = Field(
default_factory=lambda: _lazy_default(
"nanobot.agent.tools.computer_use", "ComputerUseToolConfig"
)
)
exec: ExecToolConfig = Field(default_factory=lambda: _lazy_default("nanobot.agent.tools.shell", "ExecToolConfig"))
file: FileToolsConfig = Field(default_factory=lambda: _lazy_default("nanobot.agent.tools.filesystem", "FileToolsConfig"))
cli_apps: CliAppsToolConfig = Field(default_factory=lambda: _lazy_default("nanobot.agent.tools.cli_apps", "CliAppsToolConfig"))
@@ -670,7 +682,9 @@ def _resolve_tool_config_refs() -> None:
"""
import sys
from nanobot.agent.tools.browser_tool import BrowserToolConfig
from nanobot.agent.tools.cli_apps import CliAppsToolConfig
from nanobot.agent.tools.computer_use import ComputerUseToolConfig
from nanobot.agent.tools.filesystem import FileToolsConfig
from nanobot.agent.tools.image_generation import ImageGenerationToolConfig
from nanobot.agent.tools.self import MyToolConfig
@@ -680,6 +694,8 @@ def _resolve_tool_config_refs() -> None:
# Re-export into this module's namespace
mod = sys.modules[__name__]
mod.ExecToolConfig = ExecToolConfig # type: ignore[attr-defined]
mod.BrowserToolConfig = BrowserToolConfig # type: ignore[attr-defined]
mod.ComputerUseToolConfig = ComputerUseToolConfig # type: ignore[attr-defined]
mod.FileToolsConfig = FileToolsConfig # type: ignore[attr-defined]
mod.CliAppsToolConfig = CliAppsToolConfig # type: ignore[attr-defined]
mod.WebToolsConfig = WebToolsConfig # type: ignore[attr-defined]
+69 -1
View File
@@ -671,6 +671,73 @@ class OpenAICompatProvider(LLMProvider):
dumped = str(content)
return dumped or "(empty)"
@classmethod
def _move_tool_images_to_user(
cls,
messages: list[dict[str, Any]],
) -> list[dict[str, Any]]:
"""Adapt multimodal tool results to Chat Completions' text-only tool role."""
updated: list[dict[str, Any]] = []
pending_images: list[dict[str, Any]] = []
def flush_images(next_message: dict[str, Any] | None = None) -> None:
if not pending_images:
if next_message is not None:
updated.append(next_message)
return
content: list[dict[str, Any]] = [
*pending_images,
{"type": "text", "text": "Images returned by the preceding tool call(s)."},
]
pending_images.clear()
if next_message is not None and next_message.get("role") == "user":
existing = next_message.get("content")
if isinstance(existing, str):
content.append({"type": "text", "text": existing})
elif isinstance(existing, list):
content.extend(cast(list[dict[str, Any]], existing))
updated.append({**next_message, "content": content})
else:
updated.append({"role": "user", "content": content})
if next_message is not None:
updated.append(next_message)
for message in messages:
content = message.get("content")
if message.get("role") == "tool" and isinstance(content, list):
blocks = cast(list[object], content)
images: list[dict[str, Any]] = []
text_blocks: list[object] = []
for block in blocks:
if isinstance(block, dict):
block_data = cast(dict[str, Any], block)
image_url = block_data.get("image_url")
if block_data.get("type") == "image_url" and isinstance(
image_url, dict
):
images.append({"type": "image_url", "image_url": image_url})
continue
text_blocks.append(block_data)
else:
text_blocks.append(block)
if images:
updated.append({
**message,
"content": (
cls._coerce_content_to_string(text_blocks)
if text_blocks
else "(image returned)"
),
})
pending_images.extend(images)
continue
if message.get("role") != "tool":
flush_images(message)
else:
updated.append(message)
flush_images()
return updated
def _sanitize_messages(self, messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
"""Strip non-standard keys, normalize tool_call IDs."""
sanitized = LLMProvider._sanitize_request_messages(messages, _ALLOWED_MSG_KEYS)
@@ -824,9 +891,10 @@ class OpenAICompatProvider(LLMProvider):
model_name = self._request_model_name(model_name)
sanitized_messages = self._sanitize_messages(self._sanitize_empty_content(messages))
kwargs: dict[str, Any] = {
"model": model_name,
"messages": self._sanitize_messages(self._sanitize_empty_content(messages)),
"messages": self._move_tool_images_to_user(sanitized_messages),
}
# GPT-5 and reasoning models (o1/o3/o4) reject temperature when
+12 -5
View File
@@ -356,6 +356,7 @@ _TOOL_RESULT_PREVIEW_CHARS = 1200
_TOOL_RESULTS_DIR = ".nanobot/tool-results"
_TOOL_RESULT_RETENTION_SECS = 7 * 24 * 60 * 60
_TOOL_RESULT_MAX_BUCKETS = 32
_IMAGE_TOKEN_ESTIMATE = 2048
_TRUNCATED_SUFFIX = "\n... (truncated)"
@@ -676,6 +677,7 @@ def _estimate_prompt_tokens_with_source(
reasoning_content, tool_call_id, name, plus per-message framing overhead.
"""
parts: list[str] = []
image_tokens = 0
for msg in messages:
content = msg.get("content")
if isinstance(content, str):
@@ -687,6 +689,8 @@ def _estimate_prompt_tokens_with_source(
text = part.get("text", "")
if isinstance(text, str) and text:
parts.append(text)
elif part is not None and part.get("type") in {"image_url", "input_image"}:
image_tokens += _IMAGE_TOKEN_ESTIMATE
tc = msg.get("tool_calls")
if tc:
@@ -709,7 +713,7 @@ def _estimate_prompt_tokens_with_source(
_estimate_tools_tokens(enc, tools, leading_separator=bool(parts)) if tools else 0
)
message_tokens = len(enc.encode(message_payload)) if message_payload else 0
return message_tokens + tool_tokens + per_message_overhead, "tiktoken"
return message_tokens + image_tokens + tool_tokens + per_message_overhead, "tiktoken"
except Exception:
tool_payload = (
("\n" if message_payload else "") + json.dumps(tools, ensure_ascii=False)
@@ -718,7 +722,7 @@ def _estimate_prompt_tokens_with_source(
)
payload = message_payload + tool_payload
estimated = len(payload.encode("utf-8"))
return estimated + per_message_overhead, "heuristic"
return estimated + image_tokens + per_message_overhead, "heuristic"
def estimate_prompt_tokens(
@@ -734,6 +738,7 @@ def estimate_message_tokens(message: dict[str, Any]) -> int:
"""Estimate prompt tokens contributed by one persisted message."""
content = message.get("content")
parts: list[str] = []
image_tokens = 0
if isinstance(content, str):
parts.append(content)
elif isinstance(content, list):
@@ -743,6 +748,8 @@ def estimate_message_tokens(message: dict[str, Any]) -> int:
text = part.get("text", "")
if isinstance(text, str) and text:
parts.append(text)
elif part is not None and part.get("type") in {"image_url", "input_image"}:
image_tokens += _IMAGE_TOKEN_ESTIMATE
else:
parts.append(json.dumps(raw_part, ensure_ascii=False))
elif content is not None:
@@ -760,13 +767,13 @@ def estimate_message_tokens(message: dict[str, Any]) -> int:
parts.append(rc)
payload = "\n".join(parts)
if not payload:
if not payload and not image_tokens:
return 4
try:
enc = _get_token_encoding()
return max(4, len(enc.encode(payload)) + 4)
return max(4, len(enc.encode(payload)) + image_tokens + 4)
except Exception:
return max(4, len(payload.encode("utf-8")) + 4)
return max(4, len(payload.encode("utf-8")) + image_tokens + 4)
def estimate_prompt_tokens_chain(
+29
View File
@@ -1261,6 +1261,11 @@ def settings_payload(
"use_jina_reader": config.tools.web.fetch.use_jina_reader,
},
},
"computer_use": {
"browser_enabled": config.tools.browser.enable,
"enabled": config.tools.computer_use.enable,
"backend": config.tools.computer_use.backend,
},
"api": {
"host": config.api.host,
"port": config.api.port,
@@ -2045,6 +2050,30 @@ def update_network_safety_settings(query: QueryParams) -> dict[str, Any]:
return settings_payload(requires_restart=changed)
def update_computer_use_settings(query: QueryParams) -> dict[str, Any]:
raw_browser = _query_first_alias(query, "browser_enabled", "browserEnabled")
raw_computer = _query_first_alias(query, "enabled", "computerEnabled")
if raw_browser is None and raw_computer is None:
raise WebUISettingsError("browser_enabled or enabled is required")
config = load_config()
changed = False
if raw_browser is not None:
browser_enabled = _parse_bool(raw_browser, "browser_enabled")
if config.tools.browser.enable != browser_enabled:
config.tools.browser.enable = browser_enabled
changed = True
if raw_computer is not None:
computer_enabled = _parse_bool(raw_computer, "enabled")
if config.tools.computer_use.enable != computer_enabled:
config.tools.computer_use.enable = computer_enabled
changed = True
if changed:
save_config(config)
return settings_payload(requires_restart=changed)
def update_web_search_settings(query: QueryParams) -> dict[str, Any]:
provider_name = (_query_first(query, "provider") or "").strip().lower()
provider_option = _WEB_SEARCH_PROVIDER_BY_NAME.get(provider_name)
+12
View File
@@ -64,6 +64,7 @@ from nanobot.webui.settings_api import (
settings_usage_payload,
update_agent_settings,
update_api_settings,
update_computer_use_settings,
update_image_generation_settings,
update_model_call_order,
update_model_configuration,
@@ -174,6 +175,8 @@ class WebUISettingsRouter:
return await self._handle_settings_provider_oauth(request, "logout")
if path == "/api/settings/web-search/update":
return self._handle_settings_web_search_update(request)
if path == "/api/settings/computer-use/update":
return self._handle_settings_computer_use_update(request)
if path == "/api/settings/api-service":
return self._handle_settings_api_service(request)
if path == "/api/settings/api-service/start":
@@ -506,6 +509,15 @@ class WebUISettingsRouter:
return self._error_response(e.status, e.message)
return self._json_response(self._with_restart_state(payload, section="browser"))
def _handle_settings_computer_use_update(self, request: WsRequest) -> Response:
if not self._authorized(request):
return self._unauthorized()
try:
payload = update_computer_use_settings(self._query(request))
except WebUISettingsError as e:
return self._error_response(e.status, e.message)
return self._json_response(self._with_restart_state(payload, section="runtime"))
def _handle_settings_api_service(self, request: WsRequest) -> Response:
if not self._authorized(request):
return self._unauthorized()
+5
View File
@@ -88,6 +88,11 @@ pdf = [
olostep = [
"olostep>=0.1.0; python_version < '3.14'",
]
computer-use = [
"pyautogui>=0.9.54",
"pillow>=10.0.0",
"playwright>=1.48.0",
]
dev = [
"pytest>=9.0.0,<10.0.0",
"pytest-asyncio>=1.3.0,<2.0.0",
+25
View File
@@ -1,6 +1,13 @@
from nanobot.agent.context_governance import ContextGovernor
def _image_result(label: str) -> list[dict]:
return [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{label}"}},
{"type": "text", "text": label},
]
def _assistant_tool_call(call_id: str) -> dict:
return {
"role": "assistant",
@@ -37,3 +44,21 @@ def test_drop_orphan_tool_results_drops_duplicate_tool_result() -> None:
tool_results = [m for m in result if m.get("role") == "tool"]
assert len(tool_results) == 1
assert tool_results[0]["content"] == "first"
def test_drop_stale_visual_tool_images_keeps_latest_per_tool() -> None:
messages = [
{"role": "tool", "name": "computer_use", "content": _image_result("history")},
{"role": "tool", "name": "computer_use", "content": _image_result("old")},
{"role": "tool", "name": "browser", "content": _image_result("browser")},
{"role": "tool", "name": "computer_use", "content": _image_result("latest")},
]
result = ContextGovernor.drop_stale_visual_tool_images(messages, start_index=1)
assert result is not messages
assert result[0]["content"] == messages[0]["content"]
assert [block["type"] for block in result[1]["content"]] == ["text", "text"]
assert result[2]["content"] == messages[2]["content"]
assert result[3]["content"] == messages[3]["content"]
assert messages[1]["content"][0]["type"] == "image_url"
+2
View File
@@ -76,11 +76,13 @@ class TestHandleStop:
loop.subagents.close = close_subagents
loop._exec_session_manager.close_all = AsyncMock()
loop.tools.close = AsyncMock()
with patch("nanobot.agent.loop.agent_context.close_mcp", AsyncMock()):
await loop.close_mcp()
assert events == ["turn_cancelled", "resources_closed"]
assert task.cancelled()
loop.tools.close.assert_awaited_once()
@pytest.mark.asyncio
async def test_close_mcp_serializes_duplicate_cleanup(self):
@@ -0,0 +1,50 @@
from nanobot.providers.openai_compat_provider import OpenAICompatProvider
def test_chat_completions_moves_tool_images_after_parallel_results():
image = {
"type": "image_url",
"image_url": {"url": "data:image/png;base64,AAAA"},
"_meta": {"path": "screen.png"},
}
messages = [
{"role": "assistant", "content": None, "tool_calls": [{"id": "a"}, {"id": "b"}]},
{
"role": "tool",
"tool_call_id": "a",
"content": [image, {"type": "text", "text": "clicked"}],
},
{"role": "tool", "tool_call_id": "b", "content": "other result"},
{"role": "assistant", "content": "done"},
]
result = OpenAICompatProvider._move_tool_images_to_user(messages)
assert result[1]["content"] == "clicked"
assert result[2] == messages[2]
assert result[3]["role"] == "user"
assert result[3]["content"][0] == {
"type": "image_url",
"image_url": {"url": "data:image/png;base64,AAAA"},
}
assert result[4] == messages[3]
def test_chat_completions_merges_tool_images_into_following_user_message():
messages = [
{
"role": "tool",
"tool_call_id": "a",
"content": [
{"type": "image_url", "image_url": {"url": "data:image/png;base64,AAAA"}},
],
},
{"role": "user", "content": "continue"},
]
result = OpenAICompatProvider._move_tool_images_to_user(messages)
assert len(result) == 2
assert result[0]["content"] == "(image returned)"
assert result[1]["role"] == "user"
assert result[1]["content"][-1] == {"type": "text", "text": "continue"}
+316
View File
@@ -0,0 +1,316 @@
"""Tests for DOM-based browser control."""
from __future__ import annotations
import asyncio
import io
from types import SimpleNamespace
from unittest.mock import AsyncMock, MagicMock
import pytest
from nanobot.agent.tools.browser_tool import BrowserTool, BrowserToolConfig
from nanobot.agent.tools.computer_use_backends import browser_playwright
from nanobot.agent.tools.computer_use_backends.browser_playwright import BrowserBackend
from nanobot.agent.tools.context import RequestContext, request_context
class _FakeDomBackend:
environment = "browser"
def __init__(self):
self.calls: list[tuple] = []
self.elements = [
{"ref": 1, "tag": "button", "role": "", "type": "", "name": "Submit", "href": ""},
{"ref": 2, "tag": "input", "role": "", "type": "text", "name": "your name", "href": ""},
]
async def navigate(self, url):
self.calls.append(("navigate", url))
async def dom_snapshot(self, max_elements=200):
return self.elements
async def click_ref(self, ref):
self.calls.append(("click", ref))
async def fill_ref(self, ref, text, submit=False):
self.calls.append(("fill", ref, text, submit))
async def select_ref(self, ref, value):
self.calls.append(("select", ref, value))
async def scroll_page(self, direction, amount):
self.calls.append(("scroll", direction, amount))
async def key(self, combo):
self.calls.append(("key", combo))
async def go_back(self):
self.calls.append(("back",))
async def read_text(self, max_chars=4000):
return "the number is 42"
async def current_url(self):
return "http://test.local/page"
async def screenshot(self):
from PIL import Image
buf = io.BytesIO()
Image.new("RGB", (1280, 800), (0, 0, 0)).save(buf, format="PNG")
return buf.getvalue()
async def close(self):
self.calls.append(("close",))
def _tool(**kw):
fb = _FakeDomBackend()
return BrowserTool(BrowserToolConfig(**kw), backend_impl=fb), fb
def _route(url: str, *, navigation: bool):
return SimpleNamespace(
request=SimpleNamespace(
url=url,
is_navigation_request=MagicMock(return_value=navigation),
),
abort=AsyncMock(),
continue_=AsyncMock(),
)
class TestConfigAndMetadata:
def test_defaults_off(self):
cfg = BrowserToolConfig()
assert cfg.enable is False
assert cfg.headless is True
assert cfg.include_screenshot is False
assert cfg.max_elements == 200
assert cfg.max_sessions == 8
def test_enabled_reads_config(self):
ctx = MagicMock()
ctx.config.browser.enable = True
assert BrowserTool.enabled(ctx) is True
ctx.config.browser.enable = False
assert BrowserTool.enabled(ctx) is False
def test_create_from_ctx(self):
ctx = MagicMock()
ctx.config.browser = BrowserToolConfig(enable=True, allowed_domains=["example.com"])
tool = BrowserTool.create(ctx)
assert isinstance(tool, BrowserTool)
assert tool.config.allowed_domains == ["example.com"]
def test_metadata(self):
tool, _ = _tool()
assert tool.name == "browser"
assert tool.exclusive is True
assert tool.read_only is False
assert "subagent" not in tool._scopes
def test_schema_actions(self):
tool, _ = _tool()
enum = tool.parameters["properties"]["action"]["enum"]
for a in ("navigate", "snapshot", "click", "type", "read_text"):
assert a in enum
class TestDispatch:
@pytest.mark.asyncio
async def test_navigate_returns_snapshot(self):
tool, fb = _tool()
result = await tool.execute(action="navigate", url="https://example.com")
assert ("navigate", "https://example.com") in fb.calls
assert isinstance(result, str)
assert "Navigated to https://example.com" in result
# snapshot of interactive elements is appended
assert '[1] button "Submit"' in result
assert '[2] input[text] "your name"' in result
@pytest.mark.asyncio
@pytest.mark.parametrize(
("action", "kwargs", "expected"),
[
("click", {"ref": 1}, ("click", 1)),
("type", {"ref": 2, "text": "Ada", "submit": True}, ("fill", 2, "Ada", True)),
("select", {"ref": 2, "value": "opt1"}, ("select", 2, "opt1")),
],
)
@pytest.mark.asyncio
async def test_element_actions(self, action, kwargs, expected):
tool, fb = _tool()
result = await tool.execute(action=action, **kwargs)
assert expected in fb.calls
assert "Interactive elements" in result
@pytest.mark.asyncio
async def test_scroll_and_key_and_back(self):
tool, fb = _tool()
await tool.execute(action="scroll", scroll_direction="down", scroll_amount=4)
await tool.execute(action="key", text="Enter")
await tool.execute(action="back")
assert ("scroll", "down", 4) in fb.calls
assert ("key", "Enter") in fb.calls
assert ("back",) in fb.calls
@pytest.mark.asyncio
async def test_read_text_returns_text_no_snapshot(self):
tool, _ = _tool()
result = await tool.execute(action="read_text")
assert isinstance(result, str)
assert "the number is 42" in result
assert "Interactive elements" not in result
@pytest.mark.asyncio
async def test_include_screenshot_returns_blocks(self):
tool, _ = _tool(include_screenshot=True)
result = await tool.execute(action="click", ref=1)
assert isinstance(result, list)
imgs = [b for b in result if b.get("type") == "image_url"]
texts = [b for b in result if b.get("type") == "text"]
assert imgs and texts
assert "Clicked element [1]" in texts[-1]["text"]
@pytest.mark.asyncio
async def test_calls_are_serialized_across_sessions(self):
class SlowBackend(_FakeDomBackend):
active = 0
max_active = 0
async def dom_snapshot(self, max_elements=200):
self.active += 1
self.max_active = max(self.max_active, self.active)
await asyncio.sleep(0.01)
self.active -= 1
return await super().dom_snapshot(max_elements)
backend = SlowBackend()
tool = BrowserTool(backend_impl=backend)
async def snapshot(session: str):
with request_context(
RequestContext(channel="test", chat_id=session, session_key=session)
):
return await tool.execute(action="snapshot")
await asyncio.gather(snapshot("a"), snapshot("b"))
assert backend.max_active == 1
class TestErrorsAndPolicy:
@pytest.mark.parametrize(
("kwargs", "error"),
[
({"action": "teleport"}, "unknown action"),
({"action": "click"}, "requires an element 'ref'"),
],
)
@pytest.mark.asyncio
async def test_tool_errors_are_returned_to_model(self, kwargs, error):
tool, _ = _tool()
result = await tool.execute(**kwargs)
assert isinstance(result, str) and error in result
@pytest.mark.asyncio
async def test_backend_blocks_disallowed_navigation(self):
backend = BrowserBackend(allowed_domains=["example.com"])
page = SimpleNamespace(goto=AsyncMock())
backend._page = page
with pytest.raises(ValueError, match="allowed_domains"):
await backend.navigate("https://evil.test/")
page.goto.assert_not_awaited()
@pytest.mark.asyncio
async def test_backend_allows_subdomain_navigation(self, monkeypatch: pytest.MonkeyPatch):
check = MagicMock(return_value=(True, ""))
monkeypatch.setattr(browser_playwright, "validate_url_target", check)
backend = BrowserBackend(allowed_domains=["example.com"])
page = SimpleNamespace(goto=AsyncMock())
backend._page = page
await backend.navigate("https://app.example.com/x")
page.goto.assert_awaited_once_with("https://app.example.com/x")
check.assert_called_once()
@pytest.mark.parametrize(
"url",
[
"file:///etc/passwd",
"http://127.0.0.1/",
"http://169.254.169.254/latest/meta-data/",
"ws://localhost/socket",
],
)
@pytest.mark.asyncio
async def test_browser_network_policy_blocks_local_targets(self, url: str):
backend = BrowserBackend()
with pytest.raises(ValueError, match="blocked"):
await backend.navigate(url)
@pytest.mark.asyncio
async def test_backend_intercepts_blocked_navigation(self):
backend = BrowserBackend(allowed_domains=["example.com"])
route = _route("https://evil.test/", navigation=True)
await backend._route_request(route)
route.abort.assert_awaited_once_with("blockedbyclient")
route.continue_.assert_not_awaited()
assert "allowed_domains" in (backend.pop_blocked_navigation() or "")
@pytest.mark.asyncio
async def test_backend_intercepts_private_subresource(self):
backend = BrowserBackend()
route = _route(
"http://169.254.169.254/latest/meta-data/",
navigation=False,
)
await backend._route_request(route)
route.abort.assert_awaited_once_with("blockedbyclient")
assert backend.pop_blocked_navigation() is None
@pytest.mark.asyncio
async def test_backend_does_not_apply_navigation_allowlist_to_subresources(
self, monkeypatch: pytest.MonkeyPatch
):
monkeypatch.setattr(
browser_playwright,
"validate_url_target",
MagicMock(return_value=(True, "")),
)
backend = BrowserBackend(allowed_domains=["example.com"])
route = _route("https://cdn.other.test/app.js", navigation=False)
await backend._route_request(route)
route.continue_.assert_awaited_once()
route.abort.assert_not_awaited()
@pytest.mark.asyncio
async def test_backend_intercepts_private_websocket(self):
backend = BrowserBackend()
web_socket = SimpleNamespace(
url="ws://127.0.0.1/socket",
close=AsyncMock(),
connect_to_server=AsyncMock(),
)
await backend._route_web_socket(web_socket)
web_socket.close.assert_awaited_once()
web_socket.connect_to_server.assert_not_awaited()
@pytest.mark.asyncio
async def test_backend_rejects_file_start_url_before_launch(self):
backend = BrowserBackend(start_url="file:///etc/passwd")
with pytest.raises(ValueError, match="start_url is blocked"):
await backend.dimensions()
+330
View File
@@ -0,0 +1,330 @@
"""Tests for screenshot-based computer control."""
from __future__ import annotations
import asyncio
import io
import sys
from types import SimpleNamespace
from unittest.mock import MagicMock
import pytest
from nanobot.agent.tools.computer_use import ComputerUseTool, ComputerUseToolConfig
from nanobot.agent.tools.computer_use_backends.base import ComputerBackend, SessionBackendPool
from nanobot.agent.tools.computer_use_backends.desktop_pyautogui import DesktopBackend
from nanobot.agent.tools.context import RequestContext, request_context
from nanobot.config.schema import ToolsConfig
class _FakeBackend(ComputerBackend):
"""Records actuation calls and serves a solid-colour PNG of a fixed size."""
environment = "desktop"
def __init__(self, width: int = 2560, height: int = 1600):
self.calls: list[tuple] = []
self._w, self._h = width, height
self.closed = False
async def dimensions(self) -> tuple[int, int]:
return (self._w, self._h)
async def screenshot(self) -> bytes:
from PIL import Image
img = Image.new("RGB", (self._w, self._h), (10, 20, 30))
buf = io.BytesIO()
img.save(buf, format="PNG")
return buf.getvalue()
async def click(self, x, y, button="left", count=1):
self.calls.append(("click", x, y, button, count))
async def move(self, x, y):
self.calls.append(("move", x, y))
async def drag(self, x, y):
self.calls.append(("drag", x, y))
async def scroll(self, x, y, direction, amount):
self.calls.append(("scroll", x, y, direction, amount))
async def type_text(self, text):
self.calls.append(("type", text))
async def key(self, combo):
self.calls.append(("key", combo))
async def close(self):
self.closed = True
# navigate() inherited -> raises NotImplementedError (desktop has no navigate)
def _split(result):
assert isinstance(result, list), f"expected content blocks, got {result!r}"
images = [b for b in result if isinstance(b, dict) and b.get("type") == "image_url"]
texts = [b for b in result if isinstance(b, dict) and b.get("type") == "text"]
return images, texts
def _tool(**kw):
fb = _FakeBackend(width=kw.pop("w", 2560), height=kw.pop("h", 1600))
config = ComputerUseToolConfig(target_width=1280, target_height=800, **kw)
tool = ComputerUseTool(config, backend_impl=fb)
return tool, fb
# --------------------------- config + metadata ---------------------------
class TestConfigAndMetadata:
def test_defaults_off(self):
cfg = ComputerUseToolConfig()
assert cfg.enable is False
assert cfg.backend == "desktop"
assert (cfg.target_width, cfg.target_height) == (1280, 800)
assert cfg.max_sessions == 8
assert "require_approval" not in type(cfg).model_fields
def test_tools_config_accepts_camel_case(self):
cfg = ToolsConfig.model_validate({
"browser": {"enable": True, "maxSessions": 4},
"computerUse": {"enable": True, "backend": "browser", "maxSessions": 6},
})
assert cfg.browser.enable is True
assert cfg.browser.max_sessions == 4
assert cfg.computer_use.enable is True
assert cfg.computer_use.backend == "browser"
assert cfg.computer_use.max_sessions == 6
dumped = cfg.model_dump(by_alias=True)
assert "computerUse" in dumped
assert dumped["computerUse"]["maxSessions"] == 6
def test_enabled_reads_config(self):
ctx = MagicMock()
ctx.config.computer_use.enable = True
assert ComputerUseTool.enabled(ctx) is True
ctx.config.computer_use.enable = False
assert ComputerUseTool.enabled(ctx) is False
def test_create_from_ctx(self):
ctx = MagicMock()
ctx.config.computer_use = ComputerUseToolConfig(
enable=True, backend="browser", target_width=1024, target_height=768
)
tool = ComputerUseTool.create(ctx)
assert isinstance(tool, ComputerUseTool)
assert tool.config.backend == "browser"
assert (tool.config.target_width, tool.config.target_height) == (1024, 768)
def test_tool_metadata(self):
tool, _ = _tool()
assert tool.name == "computer_use"
assert tool.exclusive is True
assert tool.read_only is False
assert tool.concurrency_safe is False
# not exposed to subagents
assert "subagent" not in tool._scopes
def test_schema_has_action_enum(self):
tool, _ = _tool()
action = tool.parameters["properties"]["action"]
assert "screenshot" in action["enum"]
assert "left_click" in action["enum"]
assert tool.parameters["required"] == ["action"]
# --------------------------- execute dispatch ---------------------------
class TestExecute:
@pytest.mark.asyncio
async def test_screenshot_returns_image_blocks(self):
tool, fb = _tool()
result = await tool.execute(action="screenshot")
images, texts = _split(result)
assert len(images) == 1
assert images[0]["image_url"]["url"].startswith("data:image/png;base64,")
assert "1280x800" in texts[-1]["text"]
assert fb.calls == [] # screenshot performs no actuation
@pytest.mark.asyncio
async def test_left_click_scales_coordinates(self):
tool, fb = _tool() # real 2560x1600 -> target 1280x800 (2x)
result = await tool.execute(action="left_click", x=100, y=50)
assert fb.calls == [("click", 200, 100, "left", 1)]
_, texts = _split(result)
assert "left_click at (200, 100)" in texts[-1]["text"]
@pytest.mark.asyncio
async def test_click_clamps_coordinates_to_screen(self):
tool, fb = _tool()
await tool.execute(action="left_click", x=5000, y=-10)
assert fb.calls == [("click", 2559, 0, "left", 1)]
@pytest.mark.asyncio
@pytest.mark.parametrize(
("action", "kwargs", "expected"),
[
("double_click", {"x": 10, "y": 10}, ("click", 20, 20, "left", 2)),
("triple_click", {"x": 10, "y": 10}, ("click", 20, 20, "left", 3)),
("right_click", {"x": 5, "y": 5}, ("click", 10, 10, "right", 1)),
("middle_click", {"x": 5, "y": 5}, ("click", 10, 10, "middle", 1)),
(
"scroll",
{"x": 100, "y": 100, "scroll_direction": "down", "scroll_amount": 5},
("scroll", 200, 200, "down", 5),
),
("type", {"text": "hello"}, ("type", "hello")),
("key", {"text": "ctrl+s"}, ("key", "ctrl+s")),
("mouse_move", {"x": 10, "y": 10}, ("move", 20, 20)),
("left_click_drag", {"x": 20, "y": 30}, ("drag", 40, 60)),
],
)
@pytest.mark.asyncio
async def test_actions_dispatch_to_backend(self, action, kwargs, expected):
tool, fb = _tool()
await tool.execute(action=action, **kwargs)
assert fb.calls == [expected]
@pytest.mark.asyncio
async def test_wait(self):
tool, fb = _tool()
result = await tool.execute(action="wait", duration=0.0)
_, texts = _split(result)
assert "Waited" in texts[-1]["text"]
@pytest.mark.parametrize(
("kwargs", "error"),
[
({"action": "frobnicate"}, "unknown action"),
({"action": "left_click"}, "requires"),
({"action": "navigate", "url": "https://example.com"}, "Error"),
],
)
@pytest.mark.asyncio
async def test_errors_are_returned_to_model(self, kwargs, error):
tool, _ = _tool()
result = await tool.execute(**kwargs)
assert isinstance(result, str) and error in result
@pytest.mark.asyncio
async def test_backend_pool_isolates_sessions_and_closes_all():
created: list[_FakeBackend] = []
finalized: list[bool] = []
def factory():
backend = _FakeBackend()
created.append(backend)
return backend
async def finalize():
finalized.append(all(backend.closed for backend in created))
pool = SessionBackendPool(factory, finalizer=finalize)
with request_context(RequestContext(channel="test", chat_id="a", session_key="test:a")):
first = await pool.get()
assert await pool.get() is first
with request_context(RequestContext(channel="test", chat_id="b", session_key="test:b")):
second = await pool.get()
assert first is not second
await pool.close()
assert len(created) == 2
assert all(backend.closed for backend in created)
assert finalized == [True]
await pool.close()
assert finalized == [True]
@pytest.mark.asyncio
async def test_backend_pool_evicts_least_recently_used_session():
created: list[_FakeBackend] = []
def factory():
backend = _FakeBackend()
created.append(backend)
return backend
pool = SessionBackendPool(factory, max_backends=2)
contexts = [
RequestContext(channel="test", chat_id=key, session_key=f"test:{key}")
for key in ("a", "b", "c")
]
with request_context(contexts[0]):
first = await pool.get()
with request_context(contexts[1]):
second = await pool.get()
with request_context(contexts[0]):
assert await pool.get() is first
with request_context(contexts[2]):
await pool.get()
assert first.closed is False
assert second.closed is True
await pool.close()
@pytest.mark.asyncio
async def test_desktop_tool_serializes_calls_across_sessions():
class SlowBackend(_FakeBackend):
active = 0
max_active = 0
async def dimensions(self):
self.active += 1
self.max_active = max(self.max_active, self.active)
await asyncio.sleep(0.01)
self.active -= 1
return await super().dimensions()
backend = SlowBackend(width=1280, height=800)
tool = ComputerUseTool(backend_impl=backend)
async def screenshot(session: str):
with request_context(RequestContext(channel="test", chat_id=session, session_key=session)):
return await tool.execute(action="screenshot")
await asyncio.gather(screenshot("a"), screenshot("b"))
assert backend.max_active == 1
@pytest.mark.asyncio
async def test_desktop_backend_uses_safe_pyautogui_calls():
pg = MagicMock()
pg.easeInOutQuad = object()
backend = DesktopBackend()
backend._pg = pg
await backend.drag(10, 20)
await backend.scroll(10, 20, "down", 3)
assert pg.dragTo.call_args.kwargs == {
"duration": 0.3,
"tween": pg.easeInOutQuad,
"button": "left",
}
pg.scroll.assert_called_once_with(-3)
@pytest.mark.asyncio
async def test_desktop_backend_rejects_unicode_instead_of_typing_incorrect_keys():
pg = MagicMock()
backend = DesktopBackend()
backend._pg = pg
with pytest.raises(ValueError, match="ASCII"):
await backend.type_text("你好")
pg.typewrite.assert_not_called()
def test_desktop_backend_preserves_pyautogui_failsafe(monkeypatch):
pg = SimpleNamespace(FAILSAFE=True)
monkeypatch.setitem(sys.modules, "pyautogui", pg)
backend = DesktopBackend()
assert backend._ensure() is pg
assert pg.FAILSAFE is True
+30
View File
@@ -28,6 +28,18 @@ class _FakeTool(Tool):
async def execute(self, **kwargs: Any) -> Any:
return kwargs
class _ClosableTool(_FakeTool):
def __init__(self, name: str, *, error: BaseException | None = None):
super().__init__(name)
self.closed = False
self.error = error
async def close(self) -> None:
self.closed = True
if self.error is not None:
raise self.error
def _tool_names(definitions: list[dict[str, Any]]) -> list[str]:
names: list[str] = []
for definition in definitions:
@@ -58,6 +70,24 @@ def test_get_definitions_orders_builtins_then_mcp_tools() -> None:
]
async def test_close_attempts_every_registered_tool() -> None:
registry = ToolRegistry()
broken = _ClosableTool("broken", error=RuntimeError("close failed"))
healthy = _ClosableTool("healthy")
registry.register(broken)
registry.register(healthy)
try:
await registry.close()
except RuntimeError as exc:
assert str(exc) == "close failed"
else:
raise AssertionError("expected close failure")
assert broken.closed is True
assert healthy.closed is True
def test_prepare_call_rejects_near_miss_tool_name_with_suggestion() -> None:
registry = ToolRegistry()
registry.register(_FakeTool("read_file"))
+30
View File
@@ -29,6 +29,36 @@ def test_estimate_prompt_tokens_chain_falls_back_without_provider_counter() -> N
assert source == "tiktoken"
def test_image_blocks_have_bounded_token_cost() -> None:
text = [{"role": "tool", "content": [{"type": "text", "text": "screen"}]}]
small_image = [{
"role": "tool",
"content": [
{"type": "image_url", "image_url": {"url": "data:image/png;base64,A"}},
{"type": "text", "text": "screen"},
],
}]
large_image = [{
"role": "tool",
"content": [
{
"type": "image_url",
"image_url": {"url": "data:image/png;base64," + "A" * 100_000},
},
{"type": "text", "text": "screen"},
],
}]
text_tokens = estimate_prompt_tokens(text)
small_tokens = estimate_prompt_tokens(small_image)
large_tokens = estimate_prompt_tokens(large_image)
assert small_tokens >= text_tokens + 2_000
assert large_tokens == small_tokens
assert estimate_message_tokens(large_image[0]) >= text_tokens + 2_000
assert estimate_message_tokens({"role": "user", "content": small_image[0]["content"][:1]}) > 2_000
def test_estimate_prompt_tokens_chain_falls_back_when_provider_counter_fails() -> None:
tokens, source = estimate_prompt_tokens_chain(
_BrokenCounterProvider(),
+48
View File
@@ -29,6 +29,7 @@ from nanobot.webui.settings_api import (
settings_usage_payload,
update_agent_settings,
update_api_settings,
update_computer_use_settings,
update_model_call_order,
update_model_configuration,
update_network_safety_settings,
@@ -1007,6 +1008,27 @@ def test_settings_payload_includes_network_safety_fields(
assert payload["advanced"]["ssrf_whitelist_count"] == 1
def test_settings_payload_includes_computer_use_tools(
tmp_path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
config_path = tmp_path / "config.json"
config = Config()
config.tools.browser.enable = True
config.tools.computer_use.enable = True
config.tools.computer_use.backend = "browser"
save_config(config, config_path)
monkeypatch.setattr("nanobot.config.loader._current_config_path", config_path)
payload = settings_payload()
assert payload["computer_use"] == {
"browser_enabled": True,
"enabled": True,
"backend": "browser",
}
def test_settings_payload_includes_exec_path_flags(
tmp_path,
monkeypatch: pytest.MonkeyPatch,
@@ -1374,6 +1396,32 @@ def test_update_network_safety_settings_writes_local_service_flag(
assert payload["requires_restart"] is True
def test_update_computer_use_settings_writes_only_requested_switches(
tmp_path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
config_path = tmp_path / "config.json"
save_config(Config(), config_path)
monkeypatch.setattr("nanobot.config.loader._current_config_path", config_path)
browser_payload = update_computer_use_settings({"browser_enabled": ["true"]})
saved = load_config(config_path)
assert saved.tools.browser.enable is True
assert saved.tools.computer_use.enable is False
assert browser_payload["requires_restart"] is True
computer_payload = update_computer_use_settings({"computerEnabled": ["true"]})
saved = load_config(config_path)
assert saved.tools.browser.enable is True
assert saved.tools.computer_use.enable is True
assert computer_payload["computer_use"]["enabled"] is True
def test_update_computer_use_settings_requires_a_switch() -> None:
with pytest.raises(WebUISettingsError, match="browser_enabled or enabled"):
update_computer_use_settings({})
def test_update_network_safety_settings_accepts_legacy_restricted_default_access(
tmp_path,
monkeypatch: pytest.MonkeyPatch,
+25
View File
@@ -138,3 +138,28 @@ async def test_model_preset_mutation_routes(
assert response.status_code == 200
assert json.loads(response.body)["routed"] == function_name
assert captured["query"] == expected_query
@pytest.mark.asyncio
async def test_computer_use_update_route(monkeypatch) -> None:
captured: dict[str, object] = {}
def update(query):
captured["query"] = query
return {"requires_restart": True}
monkeypatch.setattr("nanobot.webui.settings_routes.update_computer_use_settings", update)
request = SimpleNamespace(
path="/api/settings/computer-use/update?browser_enabled=true",
headers=Headers(),
)
response = await _router().dispatch(
None,
request,
"/api/settings/computer-use/update",
)
assert response is not None
assert response.status_code == 200
assert captured["query"] == {"browser_enabled": ["true"]}
+147 -1
View File
@@ -139,6 +139,7 @@ import {
startApiService,
stopApiService,
updateAutomation,
updateComputerUseSettings,
updateImageGenerationSettings,
updateMcpServerTools,
updateModelCallOrder,
@@ -179,6 +180,7 @@ import type {
AutomationUpdatePayload,
CliAppInfo,
CliAppsPayload,
ComputerUseSettingsUpdate,
ImageGenerationSettingsUpdate,
McpPresetInfo,
McpPresetsPayload,
@@ -762,6 +764,7 @@ export function SettingsView({
const [imageGenerationSaving, setImageGenerationSaving] = useState(false);
const [transcriptionSaving, setTranscriptionSaving] = useState(false);
const [networkSafetySaving, setNetworkSafetySaving] = useState(false);
const [computerUseSaving, setComputerUseSaving] = useState<"browser" | "computer" | null>(null);
const [apiService, setApiService] = useState<ApiServicePayload | null>(null);
const [apiServiceLoading, setApiServiceLoading] = useState(false);
const [apiServiceAction, setApiServiceAction] = useState<"start" | "stop" | null>(null);
@@ -1002,7 +1005,7 @@ export function SettingsView({
useEffect(() => {
if (
!pageVisible
|| !["channels", "models", "browser", "runtime"].includes(activeSection)
|| !["channels", "models", "browser", "runtime", "advanced"].includes(activeSection)
) {
return;
}
@@ -1559,6 +1562,34 @@ export function SettingsView({
}
};
const setComputerUseEnabled = async (
target: "browser" | "computer",
enabled: boolean,
) => {
if (!settings || computerUseSaving) return;
setComputerUseSaving(target);
try {
if (enabled && !(await installCapabilities(["computer-use"]))) return;
const update: ComputerUseSettingsUpdate = target === "browser"
? { browserEnabled: enabled }
: { computerEnabled: enabled };
const payload = await updateComputerUseSettings(token, update);
applyPayload(payload);
if (payload.requires_restart) {
setPendingRestartSections((prev) => ({
...prev,
[target === "browser" ? "browser" : "runtime"]: true,
}));
}
await maybeRestartHostEngine(payload);
setError(null);
} catch (err) {
setError((err as Error).message);
} finally {
setComputerUseSaving(null);
}
};
const handleApiServiceAction = async (
action: "start" | "stop",
values?: { host: string; port: number; timeout: number; apiKey?: string },
@@ -2238,6 +2269,11 @@ export function SettingsView({
olostepFeature={featureCatalog.find((feature) => feature.name === "olostep")}
olostepInstalling={nanobotFeatureAction === "enable:olostep"}
capabilityError={nanobotFeaturesError}
browserAutomationEnabled={settings.computer_use?.browser_enabled ?? false}
computerUseFeature={featureCatalog.find((feature) => feature.name === "computer-use")}
computerUseSaving={computerUseSaving === "browser"}
computerUseInstalling={nanobotFeatureAction === "enable:computer-use"}
onToggleBrowserAutomation={(enabled) => void setComputerUseEnabled("browser", enabled)}
/>
);
case "channels":
@@ -2364,6 +2400,13 @@ export function SettingsView({
onRestart={restartViaSettingsSurface}
isRestarting={isRestarting || hostEngineApplying}
requiresRestartPending={pendingRestartSections.runtime}
computerControlEnabled={settings.computer_use?.enabled ?? false}
computerUseBackend={settings.computer_use?.backend ?? "desktop"}
computerUseFeature={featureCatalog.find((feature) => feature.name === "computer-use")}
computerUseSaving={computerUseSaving === "computer"}
computerUseInstalling={nanobotFeatureAction === "enable:computer-use"}
capabilityError={nanobotFeaturesError}
onToggleComputerControl={(enabled) => void setComputerUseEnabled("computer", enabled)}
/>
);
default:
@@ -5255,6 +5298,11 @@ function WebSettings({
olostepFeature,
olostepInstalling,
capabilityError,
browserAutomationEnabled,
computerUseFeature,
computerUseSaving,
computerUseInstalling,
onToggleBrowserAutomation,
}: {
settings: SettingsPayload;
form: WebSearchSettingsUpdate;
@@ -5274,6 +5322,11 @@ function WebSettings({
olostepFeature?: NanobotFeatureInfo;
olostepInstalling: boolean;
capabilityError: string | null;
browserAutomationEnabled: boolean;
computerUseFeature?: NanobotFeatureInfo;
computerUseSaving: boolean;
computerUseInstalling: boolean;
onToggleBrowserAutomation: (enabled: boolean) => void;
}) {
const { t } = useTranslation();
const tx = (key: string, fallback: string) => t(key, { defaultValue: fallback });
@@ -5302,9 +5355,46 @@ function WebSettings({
: selectedProvider?.credential === "base_url"
? !baseUrl
: false;
const computerUseInstalled = computerUseFeature?.installed ?? true;
const browserAutomationDescription = computerUseInstalling
? tx("settings.help.computerUseInstalling", "Installing computer-use support...")
: computerUseInstalled
? tx(
"settings.help.browserAutomation",
"Let nanobot navigate and act on web pages using structured page elements. Requires Playwright Chromium.",
)
: tx(
"settings.help.computerUseInstall",
"Required Python support will be installed when you turn this on.",
);
return (
<div className="space-y-7">
<section>
<SettingsSectionTitle>
{tx("settings.sections.browserAutomation", "Browser automation")}
</SettingsSectionTitle>
<SettingsGroup>
<SettingsRow
title={tx("settings.rows.browserAutomation", "Browser automation")}
description={browserAutomationDescription}
>
<ToggleButton
checked={browserAutomationEnabled}
disabled={computerUseSaving || computerUseInstalling}
onChange={onToggleBrowserAutomation}
ariaLabel={tx("settings.rows.browserAutomation", "Browser automation")}
label={browserAutomationEnabled
? tx("settings.values.on", "On")
: tx("settings.values.off", "Off")}
/>
</SettingsRow>
</SettingsGroup>
{capabilityError ? (
<p className="mt-2 px-1 text-[12px] text-destructive">{capabilityError}</p>
) : null}
</section>
<section>
<SettingsSectionTitle>{tx("settings.sections.webSearch", "Web search")}</SettingsSectionTitle>
{form.provider === "olostep" && olostepFeature && !olostepFeature.installed ? (
@@ -8765,6 +8855,13 @@ function AdvancedSettings({
onSave,
onRestart,
isRestarting,
computerControlEnabled,
computerUseBackend,
computerUseFeature,
computerUseSaving,
computerUseInstalling,
capabilityError,
onToggleComputerControl,
}: {
form: NetworkSafetySettingsUpdate;
dirty: boolean;
@@ -8775,11 +8872,60 @@ function AdvancedSettings({
onSave: () => void;
onRestart?: () => void;
isRestarting?: boolean;
computerControlEnabled: boolean;
computerUseBackend: "desktop" | "browser";
computerUseFeature?: NanobotFeatureInfo;
computerUseSaving: boolean;
computerUseInstalling: boolean;
capabilityError: string | null;
onToggleComputerControl: (enabled: boolean) => void;
}) {
const { t } = useTranslation();
const tx = (key: string, fallback: string) => t(key, { defaultValue: fallback });
const computerUseInstalled = computerUseFeature?.installed ?? true;
const computerControlDescription = computerUseInstalling
? tx("settings.help.computerUseInstalling", "Installing computer-use support...")
: !computerUseInstalled
? tx(
"settings.help.computerUseInstall",
"Required Python support will be installed when you turn this on.",
)
: computerUseBackend === "browser"
? tx(
"settings.help.computerControlBrowser",
"Pixel-based control currently targets an isolated browser, as configured in config.json.",
)
: tx(
"settings.help.computerControl",
"Let nanobot see and control the computer running its engine. macOS requires Screen Recording and Accessibility access.",
);
return (
<div className="space-y-7">
<section>
<SettingsSectionTitle>
{tx("settings.sections.computerControl", "Computer control")}
</SettingsSectionTitle>
<SettingsGroup>
<SettingsRow
title={tx("settings.rows.computerControl", "Computer control")}
description={computerControlDescription}
>
<ToggleButton
checked={computerControlEnabled}
disabled={computerUseSaving || computerUseInstalling}
onChange={onToggleComputerControl}
ariaLabel={tx("settings.rows.computerControl", "Computer control")}
label={computerControlEnabled
? tx("settings.values.on", "On")
: tx("settings.values.off", "Off")}
/>
</SettingsRow>
</SettingsGroup>
{capabilityError ? (
<p className="mt-2 px-1 text-[12px] text-destructive">{capabilityError}</p>
) : null}
</section>
<section>
<SettingsSectionTitle>
{isNativeHostSurface
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "Defaults",
"webSearch": "Web search",
"webBehavior": "Behavior",
"browserAutomation": "Browser automation",
"computerControl": "Computer control",
"cliApps": "CLI apps",
"mcp": "MCP services",
"regional": "Regional",
@@ -199,6 +201,8 @@
"maxResults": "Max results",
"timeout": "Timeout",
"jinaReader": "Jina reader",
"browserAutomation": "Browser automation",
"computerControl": "Computer control",
"imageGeneration": "Image generation",
"imageProvider": "Image provider",
"imageProviderStatus": "Provider status",
@@ -244,6 +248,11 @@
"maxResults": "Results returned by each web_search call.",
"timeout": "Seconds before a search provider request times out.",
"jinaReader": "Use Jina Reader for web_fetch when available.",
"browserAutomation": "Let nanobot navigate and act on web pages using structured page elements. Requires Playwright Chromium.",
"computerControl": "Let nanobot see and control the computer running its engine. macOS requires Screen Recording and Accessibility access.",
"computerControlBrowser": "Pixel-based control currently targets an isolated browser, as configured in config.json.",
"computerUseInstall": "Required Python support will be installed when you turn this on.",
"computerUseInstalling": "Installing computer-use support...",
"imageGeneration": "Expose generate_image in chats when a configured image provider is available.",
"imageProvider": "Choose the registry provider used by generate_image.",
"imageProviderStatus": "Image generation reuses provider credentials from Providers.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "Valores predeterminados",
"webSearch": "Búsqueda web",
"webBehavior": "Comportamiento",
"browserAutomation": "Automatización del navegador",
"computerControl": "Control del ordenador",
"regional": "Configuración regional",
"webuiSafety": "Seguridad de WebUI",
"capabilities": "Capacidades",
@@ -145,6 +147,8 @@
"maxResults": "Resultados máximos",
"timeout": "Tiempo de espera",
"jinaReader": "Lector Jina",
"browserAutomation": "Automatización del navegador",
"computerControl": "Control del ordenador",
"imageGeneration": "Generación de imágenes",
"imageProvider": "Proveedor de imágenes",
"imageProviderStatus": "Estado del proveedor",
@@ -188,6 +192,11 @@
"maxResults": "Resultados devueltos por cada llamada web_search.",
"timeout": "Segundos antes de que una solicitud de búsqueda expire.",
"jinaReader": "Usa Jina Reader para web_fetch cuando esté disponible.",
"browserAutomation": "Permite que nanobot navegue y actúe en páginas web mediante elementos estructurados. Requiere Chromium de Playwright.",
"computerControl": "Permite que nanobot vea y controle el ordenador que ejecuta su motor. macOS requiere permisos de grabación de pantalla y accesibilidad.",
"computerControlBrowser": "El control por píxeles apunta actualmente a un navegador aislado, según config.json.",
"computerUseInstall": "El soporte de Python necesario se instalará al activarlo.",
"computerUseInstalling": "Instalando componentes de control...",
"imageGeneration": "Expone generate_image en chats cuando hay un proveedor de imagen configurado.",
"imageProvider": "Elige el proveedor registrado usado por generate_image.",
"imageProviderStatus": "La generación de imágenes reutiliza las credenciales de los proveedores.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "Valeurs par défaut",
"webSearch": "Recherche web",
"webBehavior": "Comportement",
"browserAutomation": "Automatisation du navigateur",
"computerControl": "Contrôle de lordinateur",
"regional": "Paramètres régionaux",
"webuiSafety": "Sécurité WebUI",
"capabilities": "Capacités",
@@ -145,6 +147,8 @@
"maxResults": "Résultats max.",
"timeout": "Délai dattente",
"jinaReader": "Lecteur Jina",
"browserAutomation": "Automatisation du navigateur",
"computerControl": "Contrôle de lordinateur",
"imageGeneration": "Génération dimages",
"imageProvider": "Fournisseur dimages",
"imageProviderStatus": "État du fournisseur",
@@ -188,6 +192,11 @@
"maxResults": "Résultats renvoyés par chaque appel web_search.",
"timeout": "Nombre de secondes avant lexpiration dune requête de recherche.",
"jinaReader": "Utilise Jina Reader pour web_fetch lorsque disponible.",
"browserAutomation": "Permet à nanobot de parcourir et manipuler les pages web à partir de leurs éléments structurés. Nécessite Chromium de Playwright.",
"computerControl": "Permet à nanobot de voir et contrôler lordinateur qui exécute son moteur. macOS exige les autorisations Enregistrement de l’écran et Accessibilité.",
"computerControlBrowser": "Le contrôle par pixels cible actuellement un navigateur isolé, conformément à config.json.",
"computerUseInstall": "Les composants Python requis seront installés lors de lactivation.",
"computerUseInstalling": "Installation des composants de contrôle...",
"imageGeneration": "Expose generate_image dans les chats lorsquun fournisseur dimage configuré est disponible.",
"imageProvider": "Choisissez le fournisseur inscrit utilisé par generate_image.",
"imageProviderStatus": "La génération dimages réutilise les identifiants des fournisseurs.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "Bawaan",
"webSearch": "Pencarian web",
"webBehavior": "Perilaku",
"browserAutomation": "Otomatisasi browser",
"computerControl": "Kontrol komputer",
"regional": "Regional",
"webuiSafety": "Keamanan WebUI",
"capabilities": "Kemampuan",
@@ -145,6 +147,8 @@
"maxResults": "Hasil maksimum",
"timeout": "Batas waktu",
"jinaReader": "Pembaca Jina",
"browserAutomation": "Otomatisasi browser",
"computerControl": "Kontrol komputer",
"imageGeneration": "Pembuatan gambar",
"imageProvider": "Penyedia gambar",
"imageProviderStatus": "Status penyedia",
@@ -188,6 +192,11 @@
"maxResults": "Hasil yang dikembalikan oleh setiap panggilan web_search.",
"timeout": "Detik sebelum permintaan penyedia pencarian mencapai batas waktu.",
"jinaReader": "Gunakan Jina Reader untuk web_fetch jika tersedia.",
"browserAutomation": "Izinkan nanobot menjelajah dan bertindak pada halaman web melalui elemen terstruktur. Memerlukan Chromium dari Playwright.",
"computerControl": "Izinkan nanobot melihat dan mengontrol komputer yang menjalankan mesinnya. macOS memerlukan izin Perekaman Layar dan Aksesibilitas.",
"computerControlBrowser": "Kontrol berbasis piksel saat ini menargetkan browser terisolasi sesuai config.json.",
"computerUseInstall": "Dukungan Python yang diperlukan akan dipasang saat diaktifkan.",
"computerUseInstalling": "Memasang komponen kontrol komputer...",
"imageGeneration": "Tampilkan generate_image di chat saat penyedia gambar yang dikonfigurasi tersedia.",
"imageProvider": "Pilih penyedia registry yang digunakan oleh generate_image.",
"imageProviderStatus": "Pembuatan gambar menggunakan kembali kredensial penyedia dari bagian Penyedia.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "既定値",
"webSearch": "ウェブ検索",
"webBehavior": "動作",
"browserAutomation": "ブラウザ自動操作",
"computerControl": "コンピュータ操作",
"regional": "地域",
"webuiSafety": "WebUI の安全性",
"capabilities": "機能",
@@ -145,6 +147,8 @@
"maxResults": "最大結果数",
"timeout": "タイムアウト",
"jinaReader": "Jina リーダー",
"browserAutomation": "ブラウザ自動操作",
"computerControl": "コンピュータ操作",
"imageGeneration": "画像生成",
"imageProvider": "画像プロバイダー",
"imageProviderStatus": "プロバイダー状態",
@@ -188,6 +192,11 @@
"maxResults": "各 web_search 呼び出しで返す結果数です。",
"timeout": "検索プロバイダーのリクエストがタイムアウトするまでの秒数です。",
"jinaReader": "利用可能な場合、web_fetch に Jina Reader を使います。",
"browserAutomation": "構造化されたページ要素を使って、nanobot がウェブページを閲覧・操作できるようにします。Playwright Chromium が必要です。",
"computerControl": "nanobot がエンジンを実行しているコンピュータを表示・操作できるようにします。macOS では画面収録とアクセシビリティの許可が必要です。",
"computerControlBrowser": "ピクセル操作は現在、config.json の設定に従って分離ブラウザを対象としています。",
"computerUseInstall": "有効にすると必要な Python コンポーネントがインストールされます。",
"computerUseInstalling": "コンピュータ操作コンポーネントをインストール中...",
"imageGeneration": "画像プロバイダーが設定済みのとき、チャットで generate_image を有効にします。",
"imageProvider": "generate_image で使用する登録済みプロバイダーを選択します。",
"imageProviderStatus": "画像生成はプロバイダー設定の認証情報を再利用します。",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "기본값",
"webSearch": "웹 검색",
"webBehavior": "동작",
"browserAutomation": "브라우저 자동화",
"computerControl": "컴퓨터 제어",
"regional": "지역",
"webuiSafety": "WebUI 보안",
"capabilities": "기능",
@@ -145,6 +147,8 @@
"maxResults": "최대 결과 수",
"timeout": "타임아웃",
"jinaReader": "Jina 리더",
"browserAutomation": "브라우저 자동화",
"computerControl": "컴퓨터 제어",
"imageGeneration": "이미지 생성",
"imageProvider": "이미지 제공자",
"imageProviderStatus": "제공자 상태",
@@ -188,6 +192,11 @@
"maxResults": "각 web_search 호출에서 반환되는 결과 수입니다.",
"timeout": "검색 제공자 요청이 타임아웃되기 전의 초입니다.",
"jinaReader": "가능할 때 web_fetch에 Jina Reader를 사용합니다.",
"browserAutomation": "nanobot이 구조화된 페이지 요소를 사용해 웹페이지를 탐색하고 조작할 수 있게 합니다. Playwright Chromium이 필요합니다.",
"computerControl": "nanobot이 엔진을 실행하는 컴퓨터를 보고 제어할 수 있게 합니다. macOS에서는 화면 기록 및 손쉬운 사용 권한이 필요합니다.",
"computerControlBrowser": "픽셀 기반 제어는 현재 config.json 설정에 따라 격리된 브라우저를 대상으로 합니다.",
"computerUseInstall": "켜면 필요한 Python 구성 요소가 설치됩니다.",
"computerUseInstalling": "컴퓨터 제어 구성 요소 설치 중...",
"imageGeneration": "구성된 이미지 제공자가 있을 때 채팅에서 generate_image를 노출합니다.",
"imageProvider": "generate_image에 사용할 등록 제공자를 선택합니다.",
"imageProviderStatus": "이미지 생성은 제공자 자격 증명을 재사용합니다.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "Padrões",
"webSearch": "Busca na web",
"webBehavior": "Comportamento",
"browserAutomation": "Automação do navegador",
"computerControl": "Controle do computador",
"cliApps": "Aplicativos CLI",
"mcp": "Servidores MCP",
"regional": "Regional",
@@ -199,6 +201,8 @@
"maxResults": "Máx. de resultados",
"timeout": "Tempo limite",
"jinaReader": "Leitor Jina",
"browserAutomation": "Automação do navegador",
"computerControl": "Controle do computador",
"imageGeneration": "Geração de imagens",
"imageProvider": "Provedor de imagem",
"imageProviderStatus": "Status do provedor",
@@ -244,6 +248,11 @@
"maxResults": "Resultados retornados por cada chamada de web_search.",
"timeout": "Segundos antes de uma requisição de busca expirar.",
"jinaReader": "Usa o Jina Reader para web_fetch quando disponível.",
"browserAutomation": "Permite que o nanobot navegue e interaja com páginas usando elementos estruturados. Requer o Chromium do Playwright.",
"computerControl": "Permite que o nanobot veja e controle o computador que executa o mecanismo. No macOS, requer acesso à Gravação de Tela e Acessibilidade.",
"computerControlBrowser": "O controle por pixels está direcionado a um navegador isolado, conforme o config.json.",
"computerUseInstall": "O suporte Python necessário será instalado ao ativar.",
"computerUseInstalling": "Instalando componentes de controle...",
"imageGeneration": "Expõe generate_image nas conversas quando há um provedor de imagem configurado.",
"imageProvider": "Escolha o provedor do registro usado por generate_image.",
"imageProviderStatus": "A geração de imagens reaproveita as credenciais dos provedores.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "Mặc định",
"webSearch": "Tìm kiếm web",
"webBehavior": "Hành vi",
"browserAutomation": "Tự động hóa trình duyệt",
"computerControl": "Điều khiển máy tính",
"regional": "Khu vực",
"webuiSafety": "An toàn WebUI",
"capabilities": "Khả năng",
@@ -145,6 +147,8 @@
"maxResults": "Kết quả tối đa",
"timeout": "Thời gian chờ",
"jinaReader": "Trình đọc Jina",
"browserAutomation": "Tự động hóa trình duyệt",
"computerControl": "Điều khiển máy tính",
"imageGeneration": "Tạo hình ảnh",
"imageProvider": "Nhà cung cấp hình ảnh",
"imageProviderStatus": "Trạng thái nhà cung cấp",
@@ -188,6 +192,11 @@
"maxResults": "Số kết quả được trả về sau mỗi lần gọi web_search.",
"timeout": "Số giây trước khi yêu cầu của nhà cung cấp tìm kiếm hết thời gian.",
"jinaReader": "Dùng Jina Reader cho web_fetch khi có thể.",
"browserAutomation": "Cho phép nanobot duyệt và thao tác trên trang web bằng các phần tử có cấu trúc. Cần Playwright Chromium.",
"computerControl": "Cho phép nanobot xem và điều khiển máy tính đang chạy engine. macOS yêu cầu quyền Ghi màn hình và Trợ năng.",
"computerControlBrowser": "Điều khiển theo điểm ảnh hiện nhắm tới trình duyệt cách ly theo config.json.",
"computerUseInstall": "Hỗ trợ Python cần thiết sẽ được cài đặt khi bật.",
"computerUseInstalling": "Đang cài đặt thành phần điều khiển...",
"imageGeneration": "Hiển thị generate_image trong chat khi đã cấu hình nhà cung cấp hình ảnh.",
"imageProvider": "Chọn nhà cung cấp registry được generate_image sử dụng.",
"imageProviderStatus": "Tạo ảnh dùng lại thông tin xác thực từ mục Nhà cung cấp.",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "默认值",
"webSearch": "网络搜索",
"webBehavior": "网络行为",
"browserAutomation": "浏览器自动化",
"computerControl": "电脑控制",
"cliApps": "CLI 应用",
"mcp": "MCP 服务",
"regional": "区域",
@@ -199,6 +201,8 @@
"maxResults": "最大结果数",
"timeout": "超时",
"jinaReader": "Jina 阅读器",
"browserAutomation": "浏览器自动化",
"computerControl": "电脑控制",
"imageGeneration": "图片生成",
"imageProvider": "图片提供商",
"imageProviderStatus": "提供商状态",
@@ -244,6 +248,11 @@
"maxResults": "每次 web_search 调用返回的结果数。",
"timeout": "搜索提供商请求超时前等待的秒数。",
"jinaReader": "可用时为 web_fetch 使用 Jina Reader。",
"browserAutomation": "允许 nanobot 通过结构化页面元素浏览网页并执行操作,需要 Playwright Chromium。",
"computerControl": "允许 nanobot 查看并控制运行引擎的电脑。macOS 需要授予“屏幕录制”和“辅助功能”权限。",
"computerControlBrowser": "像素控制当前按 config.json 配置作用于隔离浏览器。",
"computerUseInstall": "开启时会自动安装所需 Python 组件。",
"computerUseInstalling": "正在安装电脑操作组件...",
"imageGeneration": "配置图片提供商后,即可在聊天中使用 generate_image。",
"imageProvider": "选择 generate_image 使用的注册提供商。",
"imageProviderStatus": "图片生成会复用「提供商」里的凭据。",
+9
View File
@@ -115,6 +115,8 @@
"imageDefaults": "預設值",
"webSearch": "網路搜尋",
"webBehavior": "網路行為",
"browserAutomation": "瀏覽器自動化",
"computerControl": "電腦控制",
"regional": "區域",
"webuiSafety": "WebUI 安全",
"capabilities": "能力",
@@ -145,6 +147,8 @@
"maxResults": "最大結果數",
"timeout": "逾時",
"jinaReader": "Jina 閱讀器",
"browserAutomation": "瀏覽器自動化",
"computerControl": "電腦控制",
"imageGeneration": "圖片生成",
"imageProvider": "圖片供應商",
"imageProviderStatus": "供應商狀態",
@@ -188,6 +192,11 @@
"maxResults": "每次呼叫 web_search 所回傳的結果數。",
"timeout": "搜尋供應商請求逾時前的秒數。",
"jinaReader": "若可用,則讓 web_fetch 使用 Jina Reader。",
"browserAutomation": "允許 nanobot 透過結構化頁面元素瀏覽網頁並執行操作,需要 Playwright Chromium。",
"computerControl": "允許 nanobot 查看並控制執行引擎的電腦。macOS 需要授予螢幕錄製與輔助使用權限。",
"computerControlBrowser": "像素控制目前依 config.json 設定作用於隔離瀏覽器。",
"computerUseInstall": "開啟時會自動安裝所需 Python 元件。",
"computerUseInstalling": "正在安裝電腦操作元件...",
"imageGeneration": "設定圖片供應商後,即可在聊天中使用 generate_image。",
"imageProvider": "選擇 generate_image 使用的註冊供應商。",
"imageProviderStatus": "圖片生成功能會沿用 [供應商] 中的憑證。",
+19
View File
@@ -7,6 +7,7 @@ import type {
ChannelValidationPayload,
ChatSummary,
CliAppsPayload,
ComputerUseSettingsUpdate,
FilePreviewPayload,
ImageGenerationSettingsUpdate,
McpPresetsPayload,
@@ -1077,6 +1078,24 @@ export async function updateNetworkSafetySettings(
);
}
export async function updateComputerUseSettings(
token: string,
update: ComputerUseSettingsUpdate,
base: string = "",
): Promise<SettingsPayload> {
const query = new URLSearchParams();
if (update.browserEnabled !== undefined) {
query.set("browser_enabled", String(update.browserEnabled));
}
if (update.computerEnabled !== undefined) {
query.set("enabled", String(update.computerEnabled));
}
return request<SettingsPayload>(
`${base}/api/settings/computer-use/update?${query}`,
token,
);
}
export async function updateImageGenerationSettings(
token: string,
update: ImageGenerationSettingsUpdate,
+10
View File
@@ -579,6 +579,11 @@ export interface SettingsPayload {
use_jina_reader: boolean;
};
};
computer_use?: {
browser_enabled: boolean;
enabled: boolean;
backend: "desktop" | "browser";
};
api?: {
host: string;
port: number;
@@ -1102,6 +1107,11 @@ export interface NetworkSafetySettingsUpdate {
webuiDefaultAccessMode: WebuiDefaultAccessMode;
}
export interface ComputerUseSettingsUpdate {
browserEnabled?: boolean;
computerEnabled?: boolean;
}
export interface ImageGenerationSettingsUpdate {
enabled: boolean;
provider: string;
+15
View File
@@ -46,6 +46,7 @@ import {
pollChannelConnect,
startChannelConnect,
updateAutomation,
updateComputerUseSettings,
updateSidebarState,
updateImageGenerationSettings,
updateModelCallOrder,
@@ -836,6 +837,20 @@ describe("webui API helpers", () => {
);
});
it("updates computer-use capability switches", async () => {
await updateComputerUseSettings("tok", { browserEnabled: true });
expect(fetch).toHaveBeenCalledWith(
"/api/settings/computer-use/update?browser_enabled=true",
expect.objectContaining({ headers: { Authorization: "Bearer tok" } }),
);
await updateComputerUseSettings("tok", { computerEnabled: false });
expect(fetch).toHaveBeenCalledWith(
"/api/settings/computer-use/update?enabled=false",
expect.objectContaining({ headers: { Authorization: "Bearer tok" } }),
);
});
it("manages the API service capability", async () => {
await fetchApiService("tok");
expect(fetch).toHaveBeenCalledWith(
+9
View File
@@ -69,6 +69,8 @@ const LOCALIZED_SETTINGS_COPY_KEYS = [
"settings.sections.localPreferences",
"settings.sections.webSearch",
"settings.sections.webBehavior",
"settings.sections.browserAutomation",
"settings.sections.computerControl",
"settings.sections.webuiSafety",
"settings.sections.capabilities",
"settings.sections.apps",
@@ -149,6 +151,8 @@ const LOCALIZED_SETTINGS_COPY_KEYS = [
"settings.rows.fileEditDisplay",
"settings.rows.codeWrap",
"settings.rows.brandLogos",
"settings.rows.browserAutomation",
"settings.rows.computerControl",
"settings.rows.currentModel",
"settings.rows.localServiceAccess",
"settings.rows.webuiDefaultAccess",
@@ -160,6 +164,11 @@ const LOCALIZED_SETTINGS_COPY_KEYS = [
"settings.help.fileEditDisplay",
"settings.help.codeWrap",
"settings.help.brandLogos",
"settings.help.browserAutomation",
"settings.help.computerControl",
"settings.help.computerControlBrowser",
"settings.help.computerUseInstall",
"settings.help.computerUseInstalling",
"settings.help.currentModel",
"settings.help.localServiceAccess",
"settings.help.webuiDefaultAccess",
+97
View File
@@ -64,6 +64,11 @@ function settingsPayload(): SettingsPayload {
search: { max_results: 5, timeout: 30 },
fetch: { use_jina_reader: true },
},
computer_use: {
browser_enabled: false,
enabled: false,
backend: "desktop",
},
api: {
host: "127.0.0.1",
port: 8900,
@@ -4202,6 +4207,98 @@ describe("SettingsView Apps catalog", () => {
});
});
it("enables browser automation from Web settings", async () => {
const payload = settingsPayload();
const updatedPayload: SettingsPayload = {
...payload,
computer_use: { ...payload.computer_use!, browser_enabled: true },
requires_restart: true,
restart_required_sections: ["runtime"],
};
const fetchMock = vi.fn(async (input: RequestInfo | URL) => {
const url = String(input);
if (url === "/api/settings") return jsonResponse(payload);
if (url === "/api/settings/nanobot-features") {
return jsonResponse({
features: [{
name: "computer-use",
display_name: "Computer Use",
type: "feature",
enabled: true,
installed: true,
ready: true,
status: "enabled",
install_supported: true,
requires_restart: true,
}],
enabled_count: 1,
});
}
if (url === "/api/settings/computer-use/update?browser_enabled=true") {
return jsonResponse(updatedPayload);
}
return { ok: false, status: 404, json: async () => ({}) } as Response;
});
vi.stubGlobal("fetch", fetchMock);
renderSettingsView({ initialSection: "browser" });
fireEvent.click(await screen.findByRole("switch", { name: "Browser automation" }));
await waitFor(() =>
expect(fetchMock).toHaveBeenCalledWith(
"/api/settings/computer-use/update?browser_enabled=true",
expect.objectContaining({ headers: { Authorization: "Bearer tok" } }),
),
);
});
it("enables computer control from Security settings", async () => {
const payload = settingsPayload();
const updatedPayload: SettingsPayload = {
...payload,
computer_use: { ...payload.computer_use!, enabled: true },
requires_restart: true,
restart_required_sections: ["runtime"],
};
const fetchMock = vi.fn(async (input: RequestInfo | URL) => {
const url = String(input);
if (url === "/api/settings") return jsonResponse(payload);
if (url === "/api/settings/nanobot-features") {
return jsonResponse({
features: [{
name: "computer-use",
display_name: "Computer Use",
type: "feature",
enabled: true,
installed: true,
ready: true,
status: "enabled",
install_supported: true,
requires_restart: true,
}],
enabled_count: 1,
});
}
if (url === "/api/settings/computer-use/update?enabled=true") {
return jsonResponse(updatedPayload);
}
return { ok: false, status: 404, json: async () => ({}) } as Response;
});
vi.stubGlobal("fetch", fetchMock);
renderSettingsView({ initialSection: "advanced" });
fireEvent.click(await screen.findByRole("switch", { name: "Computer control" }));
await waitFor(() =>
expect(fetchMock).toHaveBeenCalledWith(
"/api/settings/computer-use/update?enabled=true",
expect.objectContaining({ headers: { Authorization: "Bearer tok" } }),
),
);
});
it("saves network safety without exposing technical SSRF copy", async () => {
const payload = settingsPayload();
const fetchMock = vi.fn(async (input: RequestInfo | URL) => {