Commit Graph
26633 Commits
Author SHA1 Message Date
1ebef5c95d fix: bedrock-pricing-geo-inregion-cross-region / add Global Cross-Region Inference (#15685)
* fix: bedrock-pricing-geo-inregion-cross-region

* Add Global Cross-Region Inference global.anthropic.claude-haiku-4-5-20251001

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
2025-10-17 19:37:12 -07:00
Ishaan Jaffer 8780b1af6f test fix v1.78.4-nightly 2025-10-17 18:55:05 -07:00
Ishaan Jaffer 0768edeaf6 fix vercel_ai_gateway/glm-4.6 2025-10-17 18:45:15 -07:00
Ishaan Jaffer 7c76202b33 add vercel_ai_gateway/glm-4.6 2025-10-17 18:24:26 -07:00
CopilotGitHubcopilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>ishaan-jaff
40076516dc Add glm-4.6 model to pricing configuration (#15679)
* Initial plan

* Add glm-4.6 model to pricing configuration

Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>
2025-10-17 18:19:34 -07:00
Ishaan JaffandGitHub cea318330e [Feat] Add Guardrails for /v1/messages and /v1/responses API (#15686)
* add get_guardrails_messages_for_call_type

* fix call type for /messages

* add anthropic endpoints

* fix bedrock guardrails

* fix config.yaml

* fix types

* fix async_pre_call_hook

* ruff fix

* fix guard

* fix test bedrock guardrail

* fix linting

* fix linting

* docs guardrails

* fix mypy linting
2025-10-17 18:09:00 -07:00
Ishaan Jaffer c72bbd0f06 bump: version 1.78.3 → 1.78.4 2025-10-17 18:08:16 -07:00
Ishaan Jaffer 6784326ec0 fix DYNAMIC_RATE_LIMIT_ERROR_THRESHOLD_PER_MINUTE 2025-10-17 18:07:12 -07:00
Ishaan Jaffer fa8c123f0d test fix 2025-10-17 18:07:04 -07:00
Ishaan Jaffer 91cd66b8dc UI new build 2025-10-17 18:03:07 -07:00
Ishaan Jaffer c69a691fb5 fix merge issue 2025-10-17 18:02:02 -07:00
Ishaan Jaffer 0144881b90 fix python3.8 2025-10-17 17:59:34 -07:00
Ishaan Jaffer b19d9435cf fix _supports_penalty_parameters 2025-10-17 17:57:24 -07:00
Ishaan Jaffer 9c6de6972f fix imports 2025-10-17 17:54:12 -07:00
3852fc96c1 [Oct Staging Branch] (#15460)
* Implement fix for thinking_blocks and converse API calls

This fixes Claude's models via the Converse API, which should also fix
Claude Code.

* Add thinking literal

* Fix mypy issues

* Type fix for redacted thinking

* Add voyage model integration in sagemaker

* Add config file logic

* Use already exiting voyage transformation

* refactor code as per comments

* fix merge error

* refactor code as per comments

* refactor code as per comments

* UI new build

* [Fix] router - regression when adding/removing models  (#15451)

* fix(router): update model_name_to_deployment_indices on deployment removal

When a deployment is deleted, the model_name_to_deployment_indices map
was not being updated, causing stale index references. This could lead
to incorrect routing behavior when deployments with the same model_name
were dynamically removed.

Changes:
- Update _update_deployment_indices_after_removal to maintain
  model_name_to_deployment_indices mapping
- Remove deleted indices and decrement indices greater than removed index
- Clean up empty entries when no deployments remain for a model name
- Update test to verify proper index shifting and cleanup behavior

* fix(router): remove redundant index building during initialization

Remove duplicate index building operations that were causing unnecessary
work during router initialization:

1. Removed redundant `_build_model_id_to_deployment_index_map` call in
   __init__ - `set_model_list` already builds all indices from scratch

2. Removed redundant `_build_model_name_index` call at end of
   `set_model_list` - the index is already built incrementally via
   `_create_deployment` -> `_add_model_to_list_and_index_map`

Both indices (model_id_to_deployment_index_map and
model_name_to_deployment_indices) are properly maintained as lookup
indexes through existing helper methods. This change eliminates O(N)
duplicate work during initialization without any behavioral changes.

The indices continue to be correctly synchronized with model_list on
all operations (add/remove/upsert).

* fix(prometheus): Fix Prometheus metric collection in a multi-workers environment (#14929)

Co-authored-by: sotazhang <sotazhang@tencent.com>

* Add tiered pricing and cost calculation for xai

* Use generic cost calculator

* Resolve conflicts in generated HTML files

* Remove penalty params as supported params for gemini preview model (#15503)

* fix conversion of thinking block

* add application level encryption in SQS (#15512)

* docs: fix doc

* docs(index.md): bump rc

* [Fix] GEMINI - CLI -  add google_routes to llm_api_routes (#15500)

* fix: add google_routes to llm_api_routes

* test: test_virtual_key_llm_api_routes_allows_google_routes

* build: bump version

* bump: version 1.78.0 → 1.78.1

* add application level encryption in SQS

* add application level encryption in SQS

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: deepanshu <deepanshu.lulla@hq.bill.com>

* [Feat] Bedrock Knowledgebase - return search_response when using /chat/completions API with LiteLLM (#15509)

* docs: fix doc

* docs(index.md): bump rc

* [Fix] GEMINI - CLI -  add google_routes to llm_api_routes (#15500)

* fix: add google_routes to llm_api_routes

* test: test_virtual_key_llm_api_routes_allows_google_routes

* add AnthropicCitation

* fix async_post_call_success_deployment_hook

* fix add vector_store_custom_logger to global callbacks

* test_e2e_bedrock_knowledgebase_retrieval_with_llm_api_call

* async_post_call_success_deployment_hook

* add async_post_call_streaming_deployment_hook

* async def test_e2e_bedrock_knowledgebase_retrieval_with_llm_api_call_streaming(setup_vector_store_registry):

* fix _call_post_streaming_deployment_hook

* fix async_post_call_streaming_deployment_hook

* test update

* docs: Accessing Search Results

* docs KB

* fix chatUI

* fix searchResults

* fix onSearchResults

* fix kb

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com>

* [Feat] Add dynamic rate limits on LiteLLM Gateway  (#15518)

* docs: fix doc

* docs(index.md): bump rc

* [Fix] GEMINI - CLI -  add google_routes to llm_api_routes (#15500)

* fix: add google_routes to llm_api_routes

* test: test_virtual_key_llm_api_routes_allows_google_routes

* build: bump version

* bump: version 1.78.0 → 1.78.1

* fix: KeyRequestBase

* fix rpm_limit_type

* fix dynamic rate limits

* fix use dynamic limits here

* fix _should_enforce_rate_limit

* fix _should_enforce_rate_limit

* fix counter

* test_dynamic_rate_limiting_v3

* use _create_rate_limit_descriptors

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com>

* Add google rerank endpoint

* Add docs

* fix mypy error

* fix mypy and lint errors

* Add haiku 4.5 integration

* Add haiku 4.5 integration for other regions as well

* Handle citation field correctly

* Fix filtering headers for signature calcs

* Add haiku 4.5 integration (#15650)

---------

Co-authored-by: Leslie Cheng <leslie.cheng5@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>
Co-authored-by: Alexsander Hamir <alexsanderhamirgomesbaptista@gmail.com>
Co-authored-by: Lucas <10226902+LoadingZhang@users.noreply.github.com>
Co-authored-by: sotazhang <sotazhang@tencent.com>
Co-authored-by: Deepanshu Lulla <deepanshu.lulla@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: deepanshu <deepanshu.lulla@hq.bill.com>
2025-10-17 17:52:25 -07:00
Ishaan JaffandGitHub ef28b593e4 Fix: Separate OAuth M2M authentication from UI SSO + Handle Introspection endpoint for Oauth2 (#15667)
* add patch for Oauth2 token

* fix Oauth2 checks
2025-10-17 17:51:43 -07:00
Alexsander HamirandGitHub eee28a7af8 fix: add missing context (#15688) 2025-10-17 17:39:21 -07:00
Ishaan JaffandGitHub a8be3ae412 [Feat] Add Cost Tracking for /ocr endpoints (#15678)
* mistral/mistral-ocr-latest

* fix: add _hidden_params to OCRResponses

* test: hidden params exists

* feat: add mistral/mistral-ocr-2505-completion

* fix test

* add ModelInfoBase fields

* fix get OCR cost

* check response cost from OCR

* add handling for OCR costs

* add mistral-document-ai-2505

* docs OCR

* ruff check fix
2025-10-17 15:54:10 -07:00
Ishaan JaffandGitHub 6c26971cd4 [Bug Fix] Tags as metadata dicts were raising exceptions (#15625)
* fix get_tags_from_request_body

* Revert "fix get_tags_from_request_body"

This reverts commit 1c044dad999500b5c11ec96f528212c248f72f61.

* fix get_tags_from_request_body

* test_get_tags_from_request_body_with_dict_tags
2025-10-17 13:20:07 -07:00
Ishaan JaffandGitHub 92bfc3e12c Fix: Support us-gov prefix for AWS GovCloud Bedrock models (#15626)
* fix _supported_cross_region_inference_region

* test_govcloud_cross_region_inference_prefix
2025-10-17 13:19:51 -07:00
Ishaan JaffandGitHub 845c43a24e Merge pull request #15642 from jlan-nl/litellm-gemini-flash-2.5-image-web-search
Fix: Gemini 2.5 Flash Image should not have supports_web_search=true
2025-10-17 13:19:08 -07:00
Ishaan JaffandGitHub 9bbf1e0c96 Merge pull request #15670 from BerriAI/litellm_fix_watsonx_pricing
[Fix (pricing)] - Fix pricing for watsonx model family for various models
2025-10-17 13:18:27 -07:00
Ishaan JaffandGitHub fa1b6feccd Merge pull request #15668 from BerriAI/litellm_fix_bedrock_gpt_oss
[Fix] GPT-OSS in Bedrock now supports streaming. Revert fake streaming
2025-10-17 13:16:48 -07:00
Ishaan JaffandGitHub 8f710abe05 Merge pull request #15672 from BerriAI/litellm_key_max_budget_error_fix
[Fix] UI - Key Max Budget Removal Error Fix
2025-10-17 12:15:20 -07:00
yuneng-jiang 4d070abbce Key Max Budget Removal Error Fix 2025-10-17 12:04:58 -07:00
Ishaan Jaffer 3bc8f76d12 fixes for various models with incorrect pricing 2025-10-17 11:51:22 -07:00
Ishaan Jaffer 8b179f0a89 pricing: fix watsonx/openai/gpt-oss-120b 2025-10-17 11:40:34 -07:00
Ishaan Jaffer f342ef3842 test fix 2025-10-17 11:18:29 -07:00
Ishaan Jaffer b2aa24f938 fix: should_fake_stream 2025-10-17 11:17:42 -07:00
Ishaan Jaffer 3df0fb0b45 test fix v1.78.3-nightly 2025-10-17 10:46:42 -07:00
Ishaan JaffandGitHub 3b1fdb56b9 Merge pull request #15658 from BerriAI/litellm_response_mode_add_health
Add support for responses mode in health check
2025-10-17 09:44:34 -07:00
Ishaan Jaffer 09a2e32663 test_completion_cohere_command_r_plus_function_call 2025-10-17 09:40:14 -07:00
Ishaan Jaffer 80af02a1c3 test fix: tests.test_litellm.proxy.guardrails.guardrail_hooks.guardrails_ai.test_guardrails_ai::test_guardrails_ai_process_input 2025-10-17 09:39:32 -07:00
Sameer Kankute 282eedb72b Add responses mode to health check 2025-10-17 22:06:29 +05:30
IQHL (Hans Jacob Landelius) d49c5dfd18 supports_web_search=false for gemini 2.5 flash image 2025-10-17 10:58:23 +02:00
YutaSaitoandGitHub bb3b77cbeb feat: add guardrail for image generation (#15619)
* feat: add guardrail for image generation

* fix: use get_str_from_messages
2025-10-16 21:50:30 -07:00
Lucas SugiandGitHub be8854f62e fix: Convert object to a correct type (#15634) 2025-10-16 21:47:32 -07:00
Nicholas CoutureandGitHub 8032e73872 [Fix] Ensure guardrail memory sync after database updates (#15633)
* chore: Consistency in install-test-deps using poetry run

* feat: update in-memory guardrails after database CRUD operations

* test: add parameterized tests for guardrail CRUD with memory sync
2025-10-16 21:46:49 -07:00
Krrish Dholakia 814aeb5fad fix(pass_through_endpoints.py): ensure passthrough endpoints are loaded in from db correctly on restarts 2025-10-16 21:39:48 -07:00
3ff073c811 UI - add arize on ui, LLMs - clarifai refactor to openai compatible route, added azure ai/grok-4 model family
* added oauth mcp to docs

* added azure ai/grok-4 model family

* Revert "added oauth mcp to docs"

This reverts commit 950b7cef44f14b2db1429f6fbd32548a7c95d325.

* fix: arize ui integration

* need to remove a file

This reverts commit d6c877b73ac763464f204b77135f3786342373b7.

* fix: add arize from ui

* updated clarifai functions to openai compatible (#15615)

* fix: npm build errors

* Snowflake provider support: added embeddings, PAT, account_id (#15372)

* snowflake support PAT, account_id and embeddings

* format

* test embeddings

* format

* complete test

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>

* Revert "Snowflake provider support: added embeddings, PAT, account_id (#15372)" (#15632)

This reverts commit c6d58e5b4af8493c020fa519d72ec6ebc90c896b.

---------

Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Mubashir Osmani <ilikewafflesomcuh@gmail.com>
Co-authored-by: mogith-pn <143642606+mogith-pn@users.noreply.github.com>
Co-authored-by: Andrey <elkin.andr@gmail.com>
2025-10-16 20:39:15 -07:00
Ishaan Jaffer 39b0e76a33 test_completion_cohere_command_r_plus_function_call 2025-10-16 20:33:41 -07:00
Ishaan Jaffer 64876634ef ui new build 2025-10-16 20:32:51 -07:00
Ishaan Jaffer 7a1f654da6 ui new build 2025-10-16 18:37:42 -07:00
Ishaan Jaffer 13e895b9b0 test fix 2025-10-16 18:00:46 -07:00
Ishaan Jaffer 1c21b2e285 fixes: AI/ML API 2025-10-16 17:52:37 -07:00
Ishaan Jaffer 1c53a686d5 fix e2eui 2025-10-16 17:48:27 -07:00
Ishaan Jaffer bf20ecab80 Revert "fix get_tags_from_request_body"
This reverts commit 1c044dad999500b5c11ec96f528212c248f72f61.
2025-10-16 17:40:08 -07:00
Ishaan Jaffer 96449b319c fix get_tags_from_request_body 2025-10-16 17:40:08 -07:00
Ishaan JaffandGitHub 02758ebc28 Merge pull request #15553 from BerriAI/litellm_oct_staging2
[Staging] Oct Staging branch 2
2025-10-16 17:38:20 -07:00
Ishaan Jaffer ea69f4547d Merge branch 'main' into litellm_oct_staging2 2025-10-16 17:06:29 -07:00