Commit Graph
28783 Commits
Author SHA1 Message Date
Ishaan Jaffer 3cac6b0a40 vllm batch 2025-12-14 12:36:07 -08:00
Ishaan Jaffer c79e0c165f add vllm batches 2025-12-14 12:36:07 -08:00
Emerson GomesandGitHub d8fc5c3a37 Add Azure Cohere 4 reranking models (#17961) 2025-12-14 12:34:46 -08:00
Ishaan Jaffer 869368ea31 docs fix 2025-12-14 11:52:08 -08:00
Ishaan Jaffer 338dbaa6bc docs a2a gateway 2025-12-14 11:51:35 -08:00
Ishaan Jaffer 11187024b3 fix title 2025-12-14 11:44:42 -08:00
dongbin-lunarkandGitHub 0be6dd57c3 fix: pass credentials to PredictionServiceClient for Vertex AI custom endpoints (#17757)
When calling Vertex AI Model Garden custom endpoints with service account
credentials (instead of ADC), the credentials were not passed to the
PredictionServiceClient, resulting in "default credentials not found" error.

- Move credential loading before cache check to ensure availability
- Pass credentials to sync/async PredictionServiceClient
- Add vertex_credentials param to async_completion and async_streaming

Closes #8597
2025-12-14 08:41:43 +05:30
Cesar GarciaandGitHub c892c2c83d fix(anthropic): use dynamic max_tokens based on model (#17900)
* fix(anthropic): use dynamic max_tokens based on model

When users don't specify max_tokens in requests to Anthropic models,
LiteLLM now uses the correct max_output_tokens value from the model
pricing JSON instead of a hardcoded 4096.

This fixes truncated responses for Claude 3.5+ models which support
higher output limits (8192 for Claude 3.5, 128k for Claude 3.7, etc.)

Fixes #8835

* fix(anthropic): restore env var support for backwards compatibility

Keep DEFAULT_ANTHROPIC_CHAT_MAX_TOKENS as fallback when model is not
found in JSON, allowing users to configure via environment variable.
2025-12-14 08:31:27 +05:30
Cesar GarciaandGitHub bd1a075a89 feat(stability): add Stability AI image generation support (#17894)
Add direct Stability AI REST API support for image generation endpoints.
This enables using Stability's SD3, SD3.5, and Stable Image models via
LiteLLM's OpenAI-compatible interface.

Changes:
- Add STABILITY provider to LlmProviders enum
- Create StabilityImageGenerationConfig with multipart/form-data support
- Add OpenAI size to Stability aspect_ratio mapping
- Register provider in ProviderConfigManager
- Add 9 Stability models to model_prices_and_context_window.json
- Add documentation at docs/providers/stability.md
- Add 25 unit tests

Supported models:
- stability/sd3, sd3-large, sd3-large-turbo, sd3-medium
- stability/sd3.5-large, sd3.5-large-turbo, sd3.5-medium
- stability/stable-image-ultra, stable-image-core
2025-12-14 08:29:45 +05:30
Cesar GarciaandGitHub 5262896d62 fix(perplexity): use API-provided cost instead of manual calculation (#17887)
Fixes #15337

Perplexity API returns pre-calculated costs in `usage.cost.total_cost`
that include the `request_cost` (fixed per-request fee). LiteLLM was
ignoring this and calculating costs manually, resulting in ~27x
underreporting (e.g., $0.0002 vs actual $0.006).

Changes:
- Use `usage.cost.total_cost` from Perplexity response when available
- Fall back to manual calculation if cost object not present
- Add tests for both behaviors
2025-12-14 08:24:44 +05:30
Kerem TurgutluandGitHub 1da0bdd33d fix gemini web search requests count (#17921)
* fix gemini web search requests count

* filter queries
2025-12-14 08:18:26 +05:30
Ishaan Jaffer 88d5efb2b7 docs v1.80.10.rc.1 2025-12-13 18:02:06 -08:00
Ishaan Jaffer bd0ae49c74 ui new build v1.80.10.rc.1 v1.80.10-nightly 2025-12-13 17:28:04 -08:00
Ishaan Jaffer efa9f69991 TestRunwaymlImageGeneration 2025-12-13 17:21:20 -08:00
Ishaan JaffandGitHub 2641f58be5 Litellm 1 80 10 (#17945)
* update providers

* v0

* docs fix

* docs fix

* docs fix
2025-12-13 17:19:58 -08:00
Ishaan Jaffer accebc49a2 TestNvidiaNim 2025-12-13 16:38:11 -08:00
Ishaan Jaffer f6c4ad92e4 async def test_update_team_guardrails_with_org_id(): 2025-12-13 16:24:58 -08:00
Ishaan Jaffer 95e818fdc4 bump litellm-proxy-extras 2025-12-13 16:22:26 -08:00
Ishaan Jaffer d9794f0811 fix bedrock embeddings - validate env 2025-12-13 16:16:30 -08:00
Ishaan Jaffer 92b72fc759 test_runwayml_tts_async 2025-12-13 16:10:48 -08:00
Ishaan Jaffer 2fd8621b38 test recraft 2025-12-13 16:10:34 -08:00
Ishaan Jaffer 050264f7d7 test_recraft_image_edit_api 2025-12-13 16:09:52 -08:00
Ishaan JaffandGitHub 14eed8aff7 [Fixes] A2a Gateway - ensure azure foundry agents work (#17943)
* add agents  v2 fixes azure

* fix auth

* get_azure_ad_token fix

* docs foundry
2025-12-13 16:08:03 -08:00
yuneng-jiangandGitHub 232bb33fec Merge pull request #17942 from BerriAI/litellm_ui_notification
[Feature] Show progress and pause on hover for Notifications
2025-12-13 15:37:55 -08:00
yuneng-jiangandGitHub 6567e43560 Merge pull request #17940 from BerriAI/litellm_ui_mcp_headers
[Fix] Add extra_headers and allowed_tools to UpdateMCPServerRequest
2025-12-13 15:20:55 -08:00
yuneng-jiang db7454c6ee Show progress and pause on hover for notifications 2025-12-13 15:16:13 -08:00
Alexsander HamirandGitHub 6635325629 fix: filter internal params in fallback code and fix test issues (#17941)
- Filter skip_mcp_handler and other internal params in fallback_utils.py before calling acompletion
  Fixes issue where internal parameters were being passed to provider APIs causing errors
- Remove deployment field from GCS bucket logger test metadata
  Fixes model name mismatch where deployment field was overriding the model in logging
- Update Bedrock Titan test to use non-deprecated model (titan-text-express-v1)
  Fixes test failure due to deprecated amazon.titan-text-lite-v1 model
2025-12-13 15:05:26 -08:00
Ishaan JaffandGitHub ed356fdfc0 [Docs] Cursor Integration (#17939)
* docs cursor

* remove bloat

* stash changes

* docs fix

* simpler docs

* docs

* docs cursor

* add cursor/chat/completions
2025-12-13 14:44:40 -08:00
yuneng-jiang 61767779f8 Adding tests 2025-12-13 14:44:26 -08:00
yuneng-jiang 5ee2338f9e Adding extra_headers and allowed_tools in UpdateMCPServerRequest 2025-12-13 14:36:37 -08:00
Alexsander HamirandGitHub fab1b81b7f fix: add agent_id field to GCS PubSub spend_logs_payload.json test expectation (#17938)
- Add agent_id: null to expected JSON to match actual payload structure
- Fixes test_async_gcs_pub_sub_v1 test failure
- agent_id is an optional field in SpendLogsPayload that is always included (as null when not provided)
2025-12-13 13:35:20 -08:00
yuneng-jiangandGitHub b681729e0c Merge pull request #17937 from BerriAI/litellm_a2a_doc_fix
[Docs] Add Import Image to A2A Docs
2025-12-13 13:30:06 -08:00
yuneng-jiang 8b0dd58a64 Importing Image in A2A Doc 2025-12-13 13:29:01 -08:00
Alexsander HamirandGitHub 425b8400b0 fix: add storage_backend and storage_url columns to schema.prisma files (#17936)
- Updated litellm/proxy/schema.prisma to include storage_backend and storage_url columns
- Updated litellm-proxy-extras/litellm_proxy_extras/schema.prisma to include storage_backend and storage_url columns
- Fixes database schema mismatch causing 'storage_backend column does not exist' errors
- Keeps all schema files in sync with root schema.prisma
2025-12-13 13:28:34 -08:00
yuneng-jiangandGitHub 1b28ea7728 Merge pull request #17935 from BerriAI/litellm_fix_agent_docs_2
[Docs] Fixing links
2025-12-13 13:10:25 -08:00
yuneng-jiang d18ed525b8 Fixing links 2025-12-13 13:09:39 -08:00
yuneng-jiangandGitHub af7f50d42b Merge pull request #17934 from BerriAI/litellm_merge_agent_docs
[Docs] Merge Agent Usage with A2A Cost Tracking
2025-12-13 13:05:23 -08:00
yuneng-jiang 773c4d08b4 Merge Agent Usage with A2A Cost Tracking 2025-12-13 13:04:28 -08:00
yuneng-jiangandGitHub 6745c81800 Merge pull request #17932 from BerriAI/litellm_doc_fix
[Docs] Agent Usage doc fix
2025-12-13 12:54:50 -08:00
yuneng-jiang b692e87836 Agent doc fix 2025-12-13 12:54:00 -08:00
Ishaan JaffandGitHub 24d6ec67c7 [QA] Cursor Integration x LiteLLM (#17855)
* fix utils.py

* ValidUserMessageContentTypesLiteral

* add _transform_tool_choice

* _transform_responses_api_content_to_chat_completion_content

* TestContentTypeTransformation

* test_map_tool_choice_string_auto

* fix validate_chat_completion_user_messages

* fix _is_input_item_tool_call_output

* fix LiteLLMCompletionResponsesConfig
2025-12-13 12:49:45 -08:00
yuneng-jiangandGitHub 2d75875ea6 Merge pull request #17931 from BerriAI/litellm_agent_usage_md
[Docs] Agent Usage Doc
2025-12-13 12:49:06 -08:00
yuneng-jiang 39b8acae55 Agent Usage Doc 2025-12-13 12:46:56 -08:00
Alexsander HamirandGitHub 892d7e8d70 [Fix] CI/CD - Fix Bedrock tool calling test failures with non-serializable objects and internal parameters (#17930)
* fix(bedrock): filter non-serializable objects from request params

- Enhanced filter_exceptions_from_params() to filter callable objects (functions) and Logging objects
- Applied filtering in Bedrock's _prepare_request_params() before deepcopy
- Applied filtering to additional_request_params before JSON serialization
- Prevents TypeError during deepcopy (APIConnectionError objects) and JSON serialization (functions, Logging objects)
- Fixes test_bedrock_tool_calling test failures

Root cause: MCP-related functions (handle_chat_completion_with_mcp, completion_callable) and litellm_logging_obj were incorrectly added to optional_params via add_provider_specific_params_to_optional_params(), which then ended up in additional_request_params. These objects should be in litellm_params, not optional_params.

* fix(bedrock): filter internal MCP parameters from API requests

Filter out LiteLLM internal/MCP-related parameters (skip_mcp_handler,
mcp_handler_context, _skip_mcp_handler) from additional_request_params
before sending to Bedrock API to prevent 'extraneous key' errors.

- Added filter_internal_params() helper function in core_helpers.py
- Applied filtering in Bedrock's _prepare_request_params() method
- Fixes test_bedrock_completion.py::test_bedrock_tool_calling

* fix: mypy type error

* fix: add filter_exceptions_from_params to recursive function ignore list

- Add filter_exceptions_from_params to IGNORE_FUNCTIONS in recursive_detector.py
- Function is safe: has max_depth parameter (default 20) to prevent infinite recursion
2025-12-13 12:38:07 -08:00
yuneng-jiangandGitHub 5a3cf8b171 Merge pull request #17929 from BerriAI/litellm_v18010_doc
[Docs] v1.80.10 draft
2025-12-13 12:05:47 -08:00
yuneng-jiang e8bbb561b9 Adding image 2025-12-13 12:03:24 -08:00
yuneng-jiang 261437623a Agent usage docs WIP 2025-12-13 11:35:04 -08:00
yoshi-p27andGitHub a154671320 Regex guardrails update (#17915)
* Update patterns.json

* regex filtering update
2025-12-13 11:12:24 -08:00
yuneng-jiangandGitHub 952e555ed3 Merge pull request #17928 from BerriAI/litellm_ui_logs_fix
[Fix] UI - Request and Response in Logs
2025-12-13 10:52:02 -08:00
yuneng-jiang 67a9d61e74 Fixing regression on logs 2025-12-13 10:43:04 -08:00