fix: two bugs in document extraction pipeline
Bug 1: _drain_pending did not call extract_documents on follow-up messages arriving mid-turn. Documents attached to queued messages were silently dropped because _build_user_content only handles images. Fix: call extract_documents before _build_user_content in _drain_pending. Bug 2: extract_documents read the entire file into memory (up to 50 MB) just to check 16 bytes of magic header for MIME detection. Fix: read only the first 16 bytes via open()+read(16) instead of Path.read_bytes(). Added regression tests for both bugs. Made-with: Cursor
This commit is contained in:
@@ -251,8 +251,9 @@ def extract_documents(
|
||||
)
|
||||
continue
|
||||
|
||||
raw = p.read_bytes()
|
||||
mime = detect_image_mime(raw) or mimetypes.guess_type(path_str)[0]
|
||||
with open(p, "rb") as f:
|
||||
header = f.read(16)
|
||||
mime = detect_image_mime(header) or mimetypes.guess_type(path_str)[0]
|
||||
if mime and mime.startswith("image/"):
|
||||
image_paths.append(path_str)
|
||||
else:
|
||||
|
||||
Reference in New Issue
Block a user