Only reuse the exact managed config/workspace instance so CLI overrides cannot silently attach to another gateway. Revalidate cached release sidecars before execution and recover from corrupted cache entries.
Co-authored-by: Bingxi Zhao <150592536+pancacake@users.noreply.github.com>
Rebuild the terminal client on OpenTUI while keeping the Python gateway as the single agent, session, tool, and memory runtime. Preserve a classic prompt fallback and publish version-matched native sidecars for supported platforms.
Co-authored-by: Bingxi Zhao <150592536+pancacake@users.noreply.github.com>
Co-authored-by: chengyongru <2755839590@qq.com>
The gateway shutdown path never closed agent resources explicitly: it relied
on the agent loop task's own finally to run close_mcp() when that task is
cancelled. When the service stops with an in-flight exec session or MCP
subprocess, that path can be skipped or cut short, leaving asyncio subprocess
transports alive after the event loop closes. They are then finalized by
__del__ against a closed loop, producing "RuntimeError: Event loop is closed"
noise in the shutdown log, and in the worst case orphaned subprocesses with
the stop stalling until systemd's timeout kills the cgroup.
The teardown is now extracted into _close_gateway_runtime() with explicit
ordering and bounds:
- Runtime tasks (including the agent loop and any in-flight turn) are
cancelled and awaited -- bounded -- before exec sessions, subagents, and MCP
servers are closed, so no active turn is using a shared resource when it
closes.
- Channel transports are closed before waiting for their runners to exit, since
some SDKs swallow task cancellation while attempting to reconnect.
- agent.close_mcp() is invoked explicitly, bounded to 15s, and is idempotent:
it is a no-op when the agent loop's own cleanup already ran, and the
guaranteed final close otherwise.
- A coroutine that swallows cancellation (e.g. an SDK reconnect loop) can no
longer hold the stop open until systemd's timeout kills the cgroup; cleanup
failures are logged instead of blocking shutdown.
Symptom
-------
LLM requests intermittently fail with:
'utf-8' codec can't encode characters in position N-N+1: surrogates not allowed
when messages contain emoji-heavy content (e.g. HTML with mixed emoji + JSON round-trips).
This blocks the affected session until the session file is quarantined.
Root cause
----------
Surrogate sanitization was only applied at the CLI entry point
(nanobot/cli/commands.py: _sanitize_surrogates). Requests entering
the LLM provider layer through other channels (Feishu, cron, webui,
tool results, memory injection) had no defensive cleaning, so any
message that happened to carry unpaired UTF-16 surrogates (from an
upstream JSON round-trip with ensure_ascii=True on ill-formed input,
memory rehydration, or third-party content) would blow up at
json.dumps -> HTTP encode time inside the provider client.
Fix
---
1. Extract sanitize_surrogates() and sanitize_surrogates_deep() into
nanobot/utils/helpers.py as the single source of truth. Both use
utf-16-le round-tripping with errors='surrogatepass' / 'replace',
so paired surrogates reconstruct back into their real code point
and lone surrogates collapse to U+FFFD.
2. Make nanobot/cli/commands.py:_sanitize_surrogates a thin wrapper
that re-exports the shared helper (backward compatible).
3. Add defense-in-depth at the LLM provider boundary in
nanobot/providers/base.py:_sanitize_empty_content by running
sanitize_surrogates_deep over each message and its content blocks
right before requests are serialized to JSON.
Non-goals
---------
- truncate_text() is intentionally left untouched. Python str slicing
cannot split a single code point into surrogate halves, so it is
not the source of lone surrogates.
- session/manager storage layer is untouched. Archived sessions
reproduced the failure only through the request path, not through
storage.
Verification
------------
- New regression suite tests/providers/test_sanitize_surrogates.py
covers: paired surrogate reconstruction, lone surrogate replacement,
identity return on clean input (zero allocation), deep recursion on
dict/list/tuple, provider _sanitize_empty_content integration, and
full utf-8 encodability of the sanitized request body.
- 14/14 new tests pass; full existing test module also green.
- Replayed 58 archived real session messages plus adversarial
lone-surrogate injection through the provider path with no encode
errors after the fix.
Impact
------
- No behaviour change for clean inputs (sanitize_surrogates_deep is
an identity return when no surrogate is present).
- Fails-safe: unpaired surrogates degrade to U+FFFD instead of
aborting the entire request.