Files
DocsGPT/application/agents
Alex 9eb7878ad3 feat(v1): support multimodal content arrays (text + image_url)
OpenAI-compatible clients send multimodal user turns as a `content` array of
typed parts. translate_request previously assigned the array straight to the
question, breaking the string-only retrieval / token-budgeting / history paths
(HTTP 500). Now:

- content_to_text() extracts text from content arrays for the question,
  history and system prompt, so the string paths work unchanged.
- The full content array (text + image_url parts) is preserved as
  `multimodal_content`, threaded to the agent and emitted as the final user
  message so images reach the model. Token budgeting uses the text only.

The content array (incl. image_url) now reaches the LLM call intact; images
render for vision-capable models. A text-only upstream model will reject the
image_url variant, as expected.
2026-06-04 12:23:08 +01:00
..
2026-04-27 22:09:33 +01:00
2025-04-01 12:33:43 +05:30
2026-03-25 15:16:18 +00:00
2026-03-25 22:34:25 +00:00
2026-05-22 16:05:03 +01:00
2026-05-22 16:05:03 +01:00
2026-04-27 22:09:33 +01:00
2026-05-22 16:05:03 +01:00
2026-04-18 13:13:57 +01:00