mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 02:12:21 +00:00
Route opted-in models to /v1/responses while keeping Chat Completions as the default. The flavor is selected per-model in the catalog YAML and threaded through ModelCapabilities (with a new reasoning_effort field). - Translate chat-shaped messages/tools/structured-output into the Responses input / text.format / flattened-tool shapes at the provider boundary; stored history and intra-turn tool round-trips need no migration. - Normalize the typed SSE event stream back into the existing handler contract (content / thought / tool-call chunks) so the streaming accumulator and OpenAI handler stay unchanged. - Carry encrypted reasoning items across the in-turn tool loop (provider-local, keyed by call_id) for stateless reasoning continuity. - OPENAI_RESPONSES_STORE opt-in enables server-side state and previous_response_id chaining across turns, persisted in message metadata. - Set gpt-5.5 to api_flavor: responses, reasoning_effort: medium. - Remove AzureOpenAILLM and the azure_openai provider/catalog.
21 lines
723 B
YAML
21 lines
723 B
YAML
provider: openai
|
|
defaults:
|
|
supports_tools: true
|
|
supports_structured_output: true
|
|
attachments: [image]
|
|
context_window: 400000
|
|
|
|
models:
|
|
- id: gpt-5.5
|
|
display_name: GPT-5.5
|
|
description: Flagship frontier model for complex reasoning, coding, and agentic work with a 1M-token context window
|
|
context_window: 1050000
|
|
api_flavor: responses
|
|
reasoning_effort: medium
|
|
- id: gpt-5.4-mini
|
|
display_name: GPT-5.4 Mini
|
|
description: Cost-efficient GPT-5.4-class model for high-volume coding, computer use, and subagent workloads
|
|
- id: gpt-5.4-nano
|
|
display_name: GPT-5.4 Nano
|
|
description: Cheapest GPT-5.4-class model, optimized for simple high-volume tasks where speed and cost matter most
|