mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 16:13:23 +00:00
OpenAI-compatible clients send multimodal user turns as a `content` array of typed parts. translate_request previously assigned the array straight to the question, breaking the string-only retrieval / token-budgeting / history paths (HTTP 500). Now: - content_to_text() extracts text from content arrays for the question, history and system prompt, so the string paths work unchanged. - The full content array (text + image_url parts) is preserved as `multimodal_content`, threaded to the agent and emitted as the final user message so images reach the model. Token budgeting uses the text only. The content array (incl. image_url) now reaches the LLM call intact; images render for vision-capable models. A text-only upstream model will reject the image_url variant, as expected.