nanobot

mirror of https://github.com/HKUDS/nanobot.git synced 2026-05-20 16:42:25 +00:00

History

chengyongru b9346b0d59 feat: generalize multimodal support with audio/video handling

Add comprehensive audio and video support across the agent pipeline:

- Generalize media placeholder system: _strip_image_content → _strip_media_content,
  _media_placeholder with type-specific labels, unified across providers
- Add detect_audio_mime with magic-byte detection and filename fallback
- Add _AUDIO_FORMAT_MAP for correct MIME-to-API-format conversion
- Add InputLimitsConfig with count limits (max_input_audios/videos) and byte limits
- Support input_audio blocks in context builder with OpenAI-compatible format
- Support video_url blocks with base64 inline data
- Add audio/video passthrough in Codex provider, placeholder fallback in Anthropic provider
- Thread supports_vision/audio/video capability flags through AgentLoop
- Unify placeholder format: [audio: path]/[video: path] instead of generic [file: path]
- Optimize file I/O: single read_bytes() instead of header+full double reads
- Extract _STRIP_MEDIA_TYPES as class constant to avoid per-call allocation

2026-04-08 23:14:40 +08:00

test_azure_openai_provider.py

fix(providers): sanitize azure responses input messages

2026-04-02 13:43:34 +08:00

test_cached_tokens.py

fix(providers): normalize anthropic cached token usage

2026-04-02 12:51:45 +08:00

test_custom_provider.py

fix(provider): accept plain text OpenAI-compatible responses

2026-03-25 01:22:21 +00:00

test_litellm_kwargs.py

fix: include byteplus providers, guard None reasoning_effort, merge extra_body