fix(providers): sanitize UTF-16 surrogates at provider request boundary
Symptom ------- LLM requests intermittently fail with: 'utf-8' codec can't encode characters in position N-N+1: surrogates not allowed when messages contain emoji-heavy content (e.g. HTML with mixed emoji + JSON round-trips). This blocks the affected session until the session file is quarantined. Root cause ---------- Surrogate sanitization was only applied at the CLI entry point (nanobot/cli/commands.py: _sanitize_surrogates). Requests entering the LLM provider layer through other channels (Feishu, cron, webui, tool results, memory injection) had no defensive cleaning, so any message that happened to carry unpaired UTF-16 surrogates (from an upstream JSON round-trip with ensure_ascii=True on ill-formed input, memory rehydration, or third-party content) would blow up at json.dumps -> HTTP encode time inside the provider client. Fix --- 1. Extract sanitize_surrogates() and sanitize_surrogates_deep() into nanobot/utils/helpers.py as the single source of truth. Both use utf-16-le round-tripping with errors='surrogatepass' / 'replace', so paired surrogates reconstruct back into their real code point and lone surrogates collapse to U+FFFD. 2. Make nanobot/cli/commands.py:_sanitize_surrogates a thin wrapper that re-exports the shared helper (backward compatible). 3. Add defense-in-depth at the LLM provider boundary in nanobot/providers/base.py:_sanitize_empty_content by running sanitize_surrogates_deep over each message and its content blocks right before requests are serialized to JSON. Non-goals --------- - truncate_text() is intentionally left untouched. Python str slicing cannot split a single code point into surrogate halves, so it is not the source of lone surrogates. - session/manager storage layer is untouched. Archived sessions reproduced the failure only through the request path, not through storage. Verification ------------ - New regression suite tests/providers/test_sanitize_surrogates.py covers: paired surrogate reconstruction, lone surrogate replacement, identity return on clean input (zero allocation), deep recursion on dict/list/tuple, provider _sanitize_empty_content integration, and full utf-8 encodability of the sanitized request body. - 14/14 new tests pass; full existing test module also green. - Replayed 58 archived real session messages plus adversarial lone-surrogate injection through the provider path with no encode errors after the fix. Impact ------ - No behaviour change for clean inputs (sanitize_surrogates_deep is an identity return when no surrogate is present). - Fails-safe: unpaired surrogates degrade to U+FFFD instead of aborting the entire request.
This commit is contained in:
+6
-12
@@ -78,7 +78,12 @@ from nanobot.config.paths import get_workspace_path, is_default_workspace # noq
|
||||
from nanobot.config.schema import Config # noqa: E402
|
||||
from nanobot.security.network import is_loopback_host # noqa: E402
|
||||
from nanobot.utils.evaluator import evaluate_response, resolve_evaluator_prompt # noqa: E402
|
||||
from nanobot.utils.helpers import sync_workspace_templates # noqa: E402
|
||||
from nanobot.utils.helpers import ( # noqa: E402
|
||||
sanitize_surrogates as _sanitize_surrogates,
|
||||
)
|
||||
from nanobot.utils.helpers import ( # noqa: E402
|
||||
sync_workspace_templates,
|
||||
)
|
||||
from nanobot.utils.restart import ( # noqa: E402
|
||||
consume_restart_notice_from_env,
|
||||
format_restart_completed_message,
|
||||
@@ -92,17 +97,6 @@ from nanobot.webui.build import ( # noqa: E402
|
||||
from nanobot.webui.sidebar_state import read_webui_sidebar_state # noqa: E402
|
||||
|
||||
|
||||
def _sanitize_surrogates(text: str) -> str:
|
||||
"""Reconstruct surrogate pairs into real characters; replace lone surrogates.
|
||||
|
||||
On Windows, console input may produce lone surrogate code points (e.g.
|
||||
``\\ud83d\\udc08`` for U+1F408). Round-tripping through UTF-16 reconstructs
|
||||
paired surrogates into their actual characters and replaces unpaired ones
|
||||
with U+FFFD.
|
||||
"""
|
||||
return text.encode("utf-16-le", errors="surrogatepass").decode("utf-16-le", errors="replace")
|
||||
|
||||
|
||||
def _signal_name(signum: int) -> str:
|
||||
with suppress(ValueError):
|
||||
return signal.Signals(signum).name
|
||||
|
||||
Reference in New Issue
Block a user