Commit Graph
36424 Commits
Author SHA1 Message Date
Sameer Kankute 76176f2a64 fix(file_search): restore should_use_emulated helper, fix dedup, extract DB helper, clean docstring
- Re-add should_use_emulated_file_search() to emulated_handler.py so H5/H6/H7/H13 tests don't fail with ImportError
- Remove per-file-id deduplication from _build_search_results_for_include so all chunks are returned (matching OpenAI native file_search behaviour); update test_H14 to assert 2 results
- Extract raw prisma DB query in check_vector_store_ids_access into a static _fetch_managed_vector_stores_by_uuids helper so the hot request path uses a named, testable function instead of an inline prisma_client.db.* call
- Remove developer-local path from test module docstring

Made-with: Cursor
2026-03-18 11:26:27 +05:30
Sameer KankuteGitHubgreptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
694cf22c9e Update tests/test_litellm/llms/vertex_ai/test_vertex_ai_batch_transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-18 11:09:20 +05:30
Krish DholakiaandGitHub cec3e9e7d4 Merge pull request #23808 from voidborne-d/fix/shared-aiohttp-session-auto-recovery
fix: auto-recover shared aiohttp session when closed
2026-03-17 22:23:01 -07:00
joereyna 8a4ef0bd05 revert: restore full changelog base to v1.82.0-stable 2026-03-17 22:17:12 -07:00
joereyna 19f82c229b fix: update full changelog base from v1.82.0-stable to v1.82.0 2026-03-17 22:11:43 -07:00
Sameer Kankute 1181adbaf3 address greptile review feedback (greploop iteration 1)
- Remove dead elif branch in retrieve_api_base derivation
- Replace unreachable try/except httpx.HTTPStatusError around GET
  calls with logging inside the status_code check (HTTPHandler.get()
  does not call raise_for_status())
- Add comments noting HTTPHandler.get()/AsyncHTTPHandler.get() do not
  accept a timeout parameter

Made-with: Cursor
2026-03-18 10:41:31 +05:30
Sameer Kankute 0dbed192e9 Add test for reasoning effort none 2026-03-18 10:37:40 +05:30
Sameer Kankute c4d27cb239 fix(vertex-ai): address greptile review – proxy retrieve URL, timeout forwarding, sync logging
- Fix retrieve_api_base derivation to handle custom proxies with
  path-based routing (not just :cancel suffix)
- Forward timeout to POST calls in cancel_batch (sync + async)
- Add try/except error logging to sync cancel path (parity with async)
- Add tests for timeout forwarding and custom proxy retrieve URL

Made-with: Cursor
2026-03-18 10:30:05 +05:30
yuneng-jiangandClaude Opus 4.6 41a7747e8c fix: document org scope behavior, fix test mocks, add org admin tests
- Document intentional legacy-matching behavior: when user_id is
  provided to an org admin, no org filter is applied (returns all of
  that user's teams across all orgs, same as legacy endpoint)
- Fix two existing security tests to properly patch user_api_key_cache,
  proxy_logging_obj, and get_user_object instead of relying on
  incidental error handling
- Add three new org admin test cases:
  - Org admin sees org-scoped teams (200 with correct where clause)
  - Org admin rejected when filtering by other org (403)
  - Org admin with user_id filter returns target user's teams

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 21:56:57 -07:00
Sameer Kankute ecb8c05d37 Add test for reasoning effort none 2026-03-18 10:24:21 +05:30
Sameer Kankute 1ff7c70011 fix(file_search): serialize first_response output items to dicts for follow-up input
Pydantic model instances (ResponseFunctionToolCall, etc.) from first_response.output
were included raw in follow_up_input; the transformation layer expects plain dicts and
called .get() on them, raising AttributeError. Serialize via model_dump(exclude_none=True).

Made-with: Cursor
2026-03-18 10:12:13 +05:30
Sameer Kankute dc7b7f852d fix(file_search): address greptile review — dead code, follow-up context, cost tracking
- Remove dead `should_use_emulated_file_search` (main.py uses its own inline guard)
- Remove dead `fallback_vector_store_ids` param from `_run_vector_searches`
- Include all first_response.output items in follow_up_input so text blocks/reasoning
  from providers like Anthropic aren't dropped from conversation context
- Accumulate first provider call's response_cost into synthesized _hidden_params so
  billing callbacks see the total cost of both emulated-flow LLM calls
- Remove broad tools=[] filter from transformation.py (backward-incompatible); the
  follow-up call already passes tools=None which is filtered by the v is not None guard

Made-with: Cursor
2026-03-18 10:10:29 +05:30
Sameer Kankute 547db8f5d1 Fix greptile comments 2026-03-18 10:02:13 +05:30
Sameer Kankute e46dd949f2 Add test for reasoning effort none 2026-03-18 09:58:20 +05:30
yuneng-jiangandClaude Opus 4.6 0485a1859a fix: use get_user_object helper, preserve caller org_id filter
- Replace raw find_unique with get_user_object in
  _build_team_list_where_conditions for cache/metrics consistency
- Remove over-complex OR clause for org admin + user_id: when user_id
  is provided, filter by that user's direct team memberships (same as
  regular users) since the access control gate already verified the
  org admin's authority
- Preserve caller-supplied organization_id instead of overwriting with
  org_admin_org_ids
- Update test mock to match get_user_object call path

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 21:19:42 -07:00
Sameer KankuteGitHubgreptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
d0d593beb8 Update tests/test_litellm/llms/vertex_ai/test_vertex_ai_batch_transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-18 09:48:24 +05:30
Cesar GarciaandGitHub 6a5b0058e3 Merge pull request #23926 from Chesars/fix/azure-gpt5-4-responses-api-routing
fix(azure): auto-route gpt-5.4+ tools+reasoning to Responses API
2026-03-18 01:14:50 -03:00
Sameer Kankute 74ae17d153 greptile comments 2026-03-18 09:41:46 +05:30
Sameer Kankute 5dd89f16f5 address greptile review: remove unused import, normalize model lookup, add xhigh tests
- Remove unused _get_model_info_helper import
- Normalize model via get_llm_provider in _is_reasoning_effort_level_explicitly_disabled
  so provider-prefixed names (openai/gpt-5.4-mini) resolve correctly
- Add test_gpt5_4_mini_allows_reasoning_effort_xhigh
- Add test_gpt5_4_nano_allows_reasoning_effort_xhigh
- Add test_gpt5_4_mini_provider_prefixed_rejects_minimal
- Extend test_gpt5_minimal_explicitly_disabled_check for openai/gpt-5.4-mini
2026-03-18 09:37:19 +05:30
yuneng-jiangandClaude Opus 4.6 1998571d94 fix: address second review round on v2/team/list
- _get_org_admin_org_ids: catch only ValueError (user not found) instead
  of bare Exception — DB errors now propagate as 500s instead of silently
  demoting org admins to regular users
- _build_team_list_where_conditions: return None (not a sentinel string)
  when user has no team memberships; list_team_v2 short-circuits to empty
  response without hitting the DB
- Org admin + team_id + user_id: use exact team_id match with org scope
  instead of OR expansion that effectively ignored the team_id filter
- Org admin + user_id (no team_id): OR(org teams, direct memberships)
  now matches legacy _authorize_and_filter_teams behaviour

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 21:03:35 -07:00
Sameer Kankute 74382f1c89 fix(vertex-ai): address greptile review feedback on batch cancel
- Replace misleading endpoint extraction with explicit endpoint = "cancel"
- Compute retrieve_api_base from URL components directly instead of
  stripping ":cancel" from the post-proxy URL, removing the hard ValueError
  that broke any custom Vertex AI proxy configuration
- Align cancel_batch provider priority in proxy endpoints to match
  create_batch order: body field → request headers → query params → default

Made-with: Cursor
2026-03-18 09:28:19 +05:30
Chesars aaf860c19b docs: add Azure custom deployment name guidance for auto-routing 2026-03-18 00:56:11 -03:00
Sameer Kankute a41239cb96 greptile comments 2026-03-18 09:24:34 +05:30
yuneng-jiangandClaude Opus 4.6 5e2fb72f42 fix: address review feedback on v2/team/list
- Fix org admin own-query regression: always check org admin status
  before the standard route check so own-queries see all org teams
- Clear user_id when org admin is detected so org scope replaces
  user-membership scope
- Remove dead isinstance(organization_id, list) branch
- Remove unused datetime import
- Remove orphaned _convert_teams_to_response helper

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 20:50:58 -07:00
Sameer Kankute 52bf372319 fix(gpt5): treat missing supports_minimal_reasoning_effort as supported
Add _is_reasoning_effort_level_explicitly_disabled to use opt-out semantics
for minimal effort: unknown/unlisted models pass through, only blocked when
the model map explicitly sets supports_minimal_reasoning_effort=false.

xhigh keeps opt-in semantics (must be explicitly supported).

Adds test for unknown-model passthrough and explicit-disabled detection.

Made-with: Cursor
2026-03-18 09:17:57 +05:30
Cesar GarciaandGitHub 3c7e37799a Merge pull request #23928 from Chesars/fix/gemini-context-caching-custom-api-base
fix(gemini): pass model to context caching URL builder for custom api_base
2026-03-18 00:43:16 -03:00
Sameer Kankute 018ccff23f fix(vertex-ai): address greptile review feedback on batch cancel
- Add try/except httpx.HTTPStatusError blocks in _async_cancel_batch for
  both POST cancel and GET retrieve calls, with verbose_logger error logging
- Fix endpoint extraction inconsistency: compute endpoint from URL without
  :cancel suffix so it matches behaviour of create_batch/retrieve_batch
- Add explicit validation that api_base ends with ':cancel' before
  stripping it, raising a descriptive error for unsupported custom proxy
  URL rewriting scenarios
- Use string-based patch() in test instead of patch.object() for robustness
  against import order changes

Made-with: Cursor
2026-03-18 09:11:20 +05:30
Sameer KankuteGitHubgreptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
6514446dcb Update litellm/llms/azure_ai/agents/handler.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-18 09:09:30 +05:30
Chesars ff536e664a fix(gemini): propagate model to check_cache/async_check_cache for custom api_base
check_and_create_cache calls check_cache first (to avoid duplicates),
which also needs model for the URL when api_base is set. Without this,
the full flow still raises ValueError before reaching the create step.
2026-03-18 00:38:35 -03:00
yuneng-jiangandClaude Opus 4.6 bd2502eeaf [Feature] /v2/team/list: Add org admin access control, members_count, and indexes
Add org admin support to /v2/team/list so org admins can list teams
within their organizations instead of getting 401. Also enrich the
response with members_count and add missing indexes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 20:34:15 -07:00
Sameer Kankute 6fe3188af0 fix(azure-ai-agents): accumulate annotations from multiple text items in streaming
- Fix bug where only last text item's annotations were preserved when
  thread.message.completed contained multiple text content items
- Accumulate annotations via extend() instead of overwriting
- Add test_azure_ai_agents_streaming_annotations_from_completed_message
- Add test_azure_ai_agents_streaming_accumulates_annotations_from_multiple_text_items

Addresses Greptile review on PR #23849

Made-with: Cursor
2026-03-18 09:04:00 +05:30
Cesar GarciaandGitHub f059ba55a9 Merge pull request #23925 from Chesars/fix/mistral-diarize-segments-response
fix(mistral): preserve diarization segments in transcription response
2026-03-18 00:32:53 -03:00
Cesar GarciaandGitHub a46b88c237 Merge pull request #23906 from Chesars/fix/anthropic-file-block-cache-control
fix(anthropic): preserve cache directive on file-type content blocks
2026-03-18 00:32:35 -03:00
Cesar GarciaandGitHub 4947074aac Merge pull request #23907 from Chesars/fix/vertex-count-tokens-location-override
fix(vertex): respect vertex_count_tokens_location for Claude count_tokens
2026-03-18 00:30:55 -03:00
Cesar GarciaandGitHub 501ddb428c Merge pull request #23618 from gambletan/fix/file-to-input-file-mapping
fix: map Chat Completion file type to Responses API input_file
2026-03-18 00:29:06 -03:00
Sameer Kankute c20c465a02 greptile comments 2026-03-18 08:55:40 +05:30
Sameer Kankute 0564e9547b Fix greptile comments 2026-03-18 08:37:31 +05:30
Chesars 8828f002be fix(gemini): pass model to context caching URL builder for custom api_base
_get_token_and_url_context_caching() was hardcoding model=None when
calling _check_custom_proxy(), which raises ValueError when api_base
is set because Gemini proxy URLs need the model name:
{api_base}/models/{model}:cachedContents

Fixes #23846
2026-03-17 23:21:24 -03:00
Chesars cb15296693 fix(azure): auto-route gpt-5.4+ tools+reasoning to Responses API
Azure GPT-5.4+ models now get the same auto-routing treatment as OpenAI
when both `reasoning_effort` and `tools` are used in `litellm.completion()`.
Previously, `reasoning_effort` was silently dropped for Azure; now the
request is bridged to the Responses API which supports both parameters.

Fixes #23914
2026-03-17 23:06:39 -03:00
Chesars 9afc469725 fix(mistral): preserve diarization segments in transcription response
Fixes #23890 — Mistral's Voxtral transcription with `diarize=true` returns
`segments` (with speaker_id, timestamps) and `language`, but these fields
were dropped when mapping the response to TranscriptionResponse.
2026-03-17 23:04:51 -03:00
yuneng-jiangandGitHub cfeafbe388 Merge pull request #23921 from BerriAI/litellm_mar17_extras
[Infra] Security and Proxy Extras for Nightly

Only known flaky tests failing. The fix for security and proxy extras worked
v1.82.4-nightly
2026-03-17 18:01:19 -07:00
Krish DholakiaandGitHub 5e570b3a66 Merge pull request #23911 from kelvin-tran/fix/cache-control-params-anthropic-document-file-message-blocks 2026-03-17 18:00:05 -07:00
Krish DholakiaandGitHub 3bd4422a97 Merge pull request #23881 from xianzongxie-stripe/xianzong-upstream-changes 2026-03-17 17:58:36 -07:00
d 🔹 88f59e1465 fix: use AsyncMock for concurrent test consistency
Address review feedback from greptile — use new_callable=AsyncMock
on the concurrent test's patch.object to ensure the mock is properly
typed as async, even though side_effect already handles the coroutine.
2026-03-18 00:54:23 +00:00
Ishaan Jaffer bae2eddd73 docs fix sidebar 2026-03-17 17:50:58 -07:00
Ishaan JaffandGitHub fc315ab4af docs(mcp_zero_trust): add MCP zero trust auth guide (#23918)
* docs(mcp_zero_trust): add MCP zero trust auth guide with hero image

* fix(docs): move hero image to static/img/ for Docusaurus build
2026-03-17 17:45:16 -07:00
yuneng-jiang 62835ff03d adding package-lock 2026-03-17 17:44:01 -07:00
yuneng-jiang 3e2845181c bumping next version 2026-03-17 17:38:09 -07:00
yuneng-jiang cc37bf5934 adding build 2026-03-17 17:37:25 -07:00
yuneng-jiang 9fa1809c30 bump: version 0.4.56 → 0.4.57 2026-03-17 17:37:04 -07:00