mirror of
https://github.com/HKUDS/nanobot.git
synced 2026-08-04 08:28:36 +00:00
* fix(api): forward real LLM usage in /v1/chat/completions response _chat_completion_response() hardcoded prompt_tokens/completion_tokens to zero. Now reads agent_loop._last_usage (set by process_direct after every LLM call) and forwards the actual prompt/completion counts. Streaming path is unchanged; usage is only surfaced in non-streaming responses for now. Fixes #4309 * fix: use defensive getattr for _last_usage and add it to all test mock agents - Use getattr(agent_loop, '_last_usage', None) in server.py for safety - Add _last_usage = {} to mock agents in test_api_attachment.py and test_api_stream.py - Prevents AttributeError/500 when mock agents don't have the attribute * fix(api): preserve provider total usage --------- Co-authored-by: michaelxer <michaelxer@users.noreply.github.com> Co-authored-by: Xubin Ren <52506698+Re-bin@users.noreply.github.com>