mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 22:13:08 +00:00
Review follow-ups. A BOM told the sniff which encoding to read, but was also taken as the verdict: three prepended bytes let any binary through, including as notes.txt. A BOM now only selects the test — UTF-8 falls through to the byte rules on the remainder, UTF-16/32 decode and judge the characters (NUL, unprintable, or replacement chars from bytes the decoder could not read). Real Notepad-Unicode text still passes, mp4-behind-a-BOM does not, in either language. The gate treated the full parser table as a given, but without docling the fallback extractor has no .tif/.tiff/.bmp/.webp/.vtt/.xml handler, so those suffixes skipped the content check and reached the plain-text fallthrough — the original bug, one install away. The worker now passes the keys of the extractor it actually built, making the second gate stricter than the route's static one rather than a copy of it. Cache bins: extraction coerced a missing value to 0 and only non-zero bins were recorded, so a provider reporting cached_tokens=0 persisted as NULL — indistinguishable from "not reported", and OpenAI reports exactly that on every uncached request. Bins are now carried as Optional and recorded when not None, which is what the nullable columns and the NULL-means-unknown comment already assumed. Anthropic's cache_read/cache_creation bins had the same shape and are fixed alongside; the int-or-None coercion is shared in llm/base.py.