mirror of
https://github.com/tiennm99/litellm.git
synced 2026-08-21 08:26:34 +00:00
docs(index.md): cleanup doc
This commit is contained in:
@@ -1,5 +1,5 @@
|
||||
---
|
||||
title: "v1.76.0-stable - RPS Improvements"
|
||||
title: "[PRE-RELEASE]v1.76.0-stable - RPS Improvements"
|
||||
slug: "v1-76-0"
|
||||
date: 2025-08-23T10:00:00
|
||||
authors:
|
||||
@@ -28,175 +28,151 @@ This release is not live yet.
|
||||
|
||||
|
||||
---
|
||||
1. New Models / Updated Models
|
||||
1. Bugs
|
||||
1. OpenAI
|
||||
1. Gpt-5 chat: clarify does not support function calling https://github.com/BerriAI/litellm/pull/13612, s/o @superpoussin22
|
||||
2. VertexAI
|
||||
1. fix vertexai batch file format by @thiagosalvatore in https://github.com/BerriAI/litellm/pull/13576
|
||||
3. LiteLLM Proxy (`litellm_proxy/`)
|
||||
1. Add support for calling image_edits + image_generations via SDK to Proxy - https://github.com/BerriAI/litellm/pull/13735
|
||||
4. OpenRouter
|
||||
1. Fix max_output_tokens value for anthropic Claude 4 - https://github.com/BerriAI/litellm/pull/13526
|
||||
5. Gemini
|
||||
1. Fix prompt caching cost calculation - https://github.com/BerriAI/litellm/pull/13742
|
||||
6. Azure
|
||||
1. Support `../openai/v1/respones` api base - https://github.com/BerriAI/litellm/pull/13526
|
||||
2. Fix azure/gpt-5-chat max_input_tokens - https://github.com/BerriAI/litellm/pull/13660
|
||||
7. Groq
|
||||
1. streaming ASCII encoding issue - https://github.com/BerriAI/litellm/pull/13675
|
||||
8. Baseten
|
||||
1. Refactored integration to use new openai-compatible endpoints - https://github.com/BerriAI/litellm/pull/13783
|
||||
9. Bedrock
|
||||
1. fix application inference profile for pass-through endpoints for bedrock - https://github.com/BerriAI/litellm/pull/13881
|
||||
10. DataRobot
|
||||
1. Updated URL handling for DataRobot provider URL - https://github.com/BerriAI/litellm/pull/13880
|
||||
2. Features
|
||||
1. Together AI
|
||||
1. Added Qwen3, Deepseek R1 0528 Throughput, GLM 4.5 and GPT-OSS models cost tracking - https://github.com/BerriAI/litellm/pull/13637 s/o Tasmay-Tibrewal
|
||||
2. Fireworks AI
|
||||
1. add fireworks_ai/accounts/fireworks/models/deepseek-v3-0324 - https://github.com/BerriAI/litellm/pull/13821
|
||||
3. VertexAI
|
||||
1. Add VertexAI qwen API Service - https://github.com/BerriAI/litellm/pull/13828
|
||||
2. Add new VertexAI image models vertex_ai/imagen-4.0-generate-001, vertex_ai/imagen-4.0-ultra-generate-001, vertex_ai/imagen-4.0-fast-generate-001 - https://github.com/BerriAI/litellm/pull/13874
|
||||
4. Anthropic
|
||||
1. Add long context support w/ cost tracking - https://github.com/BerriAI/litellm/pull/13759
|
||||
5. DeepInfra
|
||||
1. Add rerank endpoint support for deepinfra - https://github.com/BerriAI/litellm/pull/13820
|
||||
2. Add new models for cost tracking - https://github.com/BerriAI/litellm/pull/13883 s/o @Toy-97
|
||||
6. Bedrock
|
||||
1. Add tool prompt caching on async calls - https://github.com/BerriAI/litellm/pull/13803 s/o UlookEE
|
||||
2. role chaining and session name with webauthentication for aws bedrock - https://github.com/BerriAI/litellm/pull/13753 s/o RichardoC
|
||||
7. Ollama
|
||||
1. Handle Ollama null response when using tool calling with non-tool trained models - https://github.com/BerriAI/litellm/pull/13902
|
||||
8. OpenRouter
|
||||
1. Add deepseek/deepseek-chat-v3.1 support - https://github.com/BerriAI/litellm/pull/13897
|
||||
9. Mistral
|
||||
1. Add support for calling mistral files via chat completions - https://github.com/BerriAI/litellm/pull/13866 s/o @jinskjoy
|
||||
2. Handle empty assistant content - https://github.com/BerriAI/litellm/pull/13671
|
||||
3. Support new ‘thinking’ response block - https://github.com/BerriAI/litellm/pull/13671
|
||||
10. Databricks
|
||||
1. remove deprecated dbrx models (dbrx-instruct, llama 3.1) - https://github.com/BerriAI/litellm/pull/13843
|
||||
11. AI/ML API
|
||||
1. Image gen api support - https://github.com/BerriAI/litellm/pull/13893
|
||||
|
||||
Final Count = 27
|
||||
## New Models / Updated Models
|
||||
|
||||
#### Bugs
|
||||
- **[OpenAI](../../docs/providers/openai)**
|
||||
- Gpt-5 chat: clarify does not support function calling [PR #13612](https://github.com/BerriAI/litellm/pull/13612), s/o @[superpoussin22](https://github.com/superpoussin22)
|
||||
- **[VertexAI](../../docs/providers/vertex)**
|
||||
- fix vertexai batch file format by @[thiagosalvatore](https://github.com/thiagosalvatore) in [PR #13576](https://github.com/BerriAI/litellm/pull/13576)
|
||||
- **[LiteLLM Proxy](../../docs/providers/litellm_proxy)**
|
||||
- Add support for calling image_edits + image_generations via SDK to Proxy - [PR #13735](https://github.com/BerriAI/litellm/pull/13735)
|
||||
- **[OpenRouter](../../docs/providers/openrouter)**
|
||||
- Fix max_output_tokens value for anthropic Claude 4 - [PR #13526](https://github.com/BerriAI/litellm/pull/13526)
|
||||
- **[Gemini](../../docs/providers/gemini)**
|
||||
- Fix prompt caching cost calculation - [PR #13742](https://github.com/BerriAI/litellm/pull/13742)
|
||||
- **[Azure](../../docs/providers/azure)**
|
||||
- Support `../openai/v1/respones` api base - [PR #13526](https://github.com/BerriAI/litellm/pull/13526)
|
||||
- Fix azure/gpt-5-chat max_input_tokens - [PR #13660](https://github.com/BerriAI/litellm/pull/13660)
|
||||
- **[Groq](../../docs/providers/groq)**
|
||||
- streaming ASCII encoding issue - [PR #13675](https://github.com/BerriAI/litellm/pull/13675)
|
||||
- **[Baseten](../../docs/providers/baseten)**
|
||||
- Refactored integration to use new openai-compatible endpoints - [PR #13783](https://github.com/BerriAI/litellm/pull/13783)
|
||||
- **[Bedrock](../../docs/providers/bedrock)**
|
||||
- fix application inference profile for pass-through endpoints for bedrock - [PR #13881](https://github.com/BerriAI/litellm/pull/13881)
|
||||
- **[DataRobot](../../docs/providers/datarobot)**
|
||||
- Updated URL handling for DataRobot provider URL - [PR #13880](https://github.com/BerriAI/litellm/pull/13880)
|
||||
|
||||
#### Features
|
||||
- **[Together AI](../../docs/providers/together)**
|
||||
- Added Qwen3, Deepseek R1 0528 Throughput, GLM 4.5 and GPT-OSS models cost tracking - [PR #13637](https://github.com/BerriAI/litellm/pull/13637), s/o @[Tasmay-Tibrewal](https://github.com/Tasmay-Tibrewal)
|
||||
- **[Fireworks AI](../../docs/providers/fireworks_ai)**
|
||||
- add fireworks_ai/accounts/fireworks/models/deepseek-v3-0324 - [PR #13821](https://github.com/BerriAI/litellm/pull/13821)
|
||||
- **[VertexAI](../../docs/providers/vertex)**
|
||||
- Add VertexAI qwen API Service - [PR #13828](https://github.com/BerriAI/litellm/pull/13828)
|
||||
- Add new VertexAI image models vertex_ai/imagen-4.0-generate-001, vertex_ai/imagen-4.0-ultra-generate-001, vertex_ai/imagen-4.0-fast-generate-001 - [PR #13874](https://github.com/BerriAI/litellm/pull/13874)
|
||||
- **[Anthropic](../../docs/providers/anthropic)**
|
||||
- Add long context support w/ cost tracking - [PR #13759](https://github.com/BerriAI/litellm/pull/13759)
|
||||
- **[DeepInfra](../../docs/providers/deepinfra)**
|
||||
- Add rerank endpoint support for deepinfra - [PR #13820](https://github.com/BerriAI/litellm/pull/13820)
|
||||
- Add new models for cost tracking - [PR #13883](https://github.com/BerriAI/litellm/pull/13883), s/o @[Toy-97](https://github.com/Toy-97)
|
||||
- **[Bedrock](../../docs/providers/bedrock)**
|
||||
- Add tool prompt caching on async calls - [PR #13803](https://github.com/BerriAI/litellm/pull/13803), s/o @[UlookEE](https://github.com/UlookEE)
|
||||
- role chaining and session name with webauthentication for aws bedrock - [PR #13753](https://github.com/BerriAI/litellm/pull/13753), s/o @[RichardoC](https://github.com/RichardoC)
|
||||
- **[Ollama](../../docs/providers/ollama)**
|
||||
- Handle Ollama null response when using tool calling with non-tool trained models - [PR #13902](https://github.com/BerriAI/litellm/pull/13902)
|
||||
- **[OpenRouter](../../docs/providers/openrouter)**
|
||||
- Add deepseek/deepseek-chat-v3.1 support - [PR #13897](https://github.com/BerriAI/litellm/pull/13897)
|
||||
- **[Mistral](../../docs/providers/mistral)**
|
||||
- Add support for calling mistral files via chat completions - [PR #13866](https://github.com/BerriAI/litellm/pull/13866), s/o @[jinskjoy](https://github.com/jinskjoy)
|
||||
- Handle empty assistant content - [PR #13671](https://github.com/BerriAI/litellm/pull/13671)
|
||||
- Support new ‘thinking’ response block - [PR #13671](https://github.com/BerriAI/litellm/pull/13671)
|
||||
- **[Databricks](../../docs/providers/databricks)**
|
||||
- remove deprecated dbrx models (dbrx-instruct, llama 3.1) - [PR #13843](https://github.com/BerriAI/litellm/pull/13843)
|
||||
- **[AI/ML API](../../docs/providers/ai_ml_api)**
|
||||
- Image gen api support - [PR #13893](https://github.com/BerriAI/litellm/pull/13893)
|
||||
|
||||
|
||||
|
||||
1. LLM API Endpoints
|
||||
1. Bugs
|
||||
1. Responses API
|
||||
1. add default api version for openai responses api calls - https://github.com/BerriAI/litellm/pull/13526
|
||||
2. support allowed_openai_params - https://github.com/BerriAI/litellm/pull/13671
|
||||
2. Features
|
||||
1.
|
||||
|
||||
Final Count = 2
|
||||
## LLM API Endpoints
|
||||
#### Bugs
|
||||
- **[Responses API](../../docs/response_api)**
|
||||
- add default api version for openai responses api calls - [PR #13526](https://github.com/BerriAI/litellm/pull/13526)
|
||||
- support allowed_openai_params - [PR #13671](https://github.com/BerriAI/litellm/pull/13671)
|
||||
|
||||
|
||||
1. MCP Gateway
|
||||
1. Bugs
|
||||
1. fix StreamableHTTPSessionManager .run() error - https://github.com/BerriAI/litellm/pull/13666
|
||||
2. Features
|
||||
1.
|
||||
## MCP Gateway
|
||||
#### Bugs
|
||||
- fix StreamableHTTPSessionManager .run() error - https://github.com/BerriAI/litellm/pull/13666
|
||||
|
||||
Final count = 1
|
||||
## Vector Stores
|
||||
#### Bugs
|
||||
- **[Bedrock](../../docs/providers/bedrock)**
|
||||
- Using LiteLLM Managed Credentials for Query - [PR #13787](https://github.com/BerriAI/litellm/pull/13787)
|
||||
|
||||
## Management Endpoints / UI
|
||||
#### Bugs
|
||||
- **[Passthrough](../../docs/pass_through/intro)**
|
||||
- Fix query passthrough deletion - [PR #13622](https://github.com/BerriAI/litellm/pull/13622)
|
||||
|
||||
#### Features
|
||||
- **Models**
|
||||
- Add Search Functionality for Public Model Names in Model Dashboard - [PR #13687](https://github.com/BerriAI/litellm/pull/13687)
|
||||
- Auto-Add `azure/` to deployment Name in UI - [PR #13685](https://github.com/BerriAI/litellm/pull/13685)
|
||||
- Models page row UI restructure - [PR #13771](https://github.com/BerriAI/litellm/pull/13771)
|
||||
- **Notifications**
|
||||
- Add new notifications toast UI everywhere - [PR #13813](https://github.com/BerriAI/litellm/pull/13813)
|
||||
- **Keys**
|
||||
- Fix key edit settings after regenerating a key - [PR #13815](https://github.com/BerriAI/litellm/pull/13815)
|
||||
- Require team_id when creating service account keys - [PR #13873](https://github.com/BerriAI/litellm/pull/13873)
|
||||
- Filter - show all options on filter option click - [PR #13858](https://github.com/BerriAI/litellm/pull/13858)
|
||||
- **Usage**
|
||||
- Fix ‘Cannot read properties of undefined’ exception on user agent activity tab - [PR #13892](https://github.com/BerriAI/litellm/pull/13892)
|
||||
- **SSO**
|
||||
- Free SSO usage for up to 5 users - [PR #13843](https://github.com/BerriAI/litellm/pull/13843)
|
||||
|
||||
## Logging / Guardrail Integrations
|
||||
#### Bugs
|
||||
- **[Bedrock Guardrails](../../docs/proxy/guardrails/bedrock)**
|
||||
- Add bedrock api key support - [PR #13835](https://github.com/BerriAI/litellm/pull/13835)
|
||||
#### Features
|
||||
- **[Datadog LLM Observability](../../docs/integrations/datadog)**
|
||||
- Add support for Failure Logging [PR #13726](https://github.com/BerriAI/litellm/pull/13726)
|
||||
- Add time to first token, litellm overhead, guardrail overhead latency metrics - [PR #13734](https://github.com/BerriAI/litellm/pull/13734)
|
||||
- Add support for tracing guardrail input/output - [PR #13767](https://github.com/BerriAI/litellm/pull/13767)
|
||||
- **[Langfuse OTEL](../../docs/integrations/langfuse)**
|
||||
- Allow using Key/Team Based Logging - [PR #13791](https://github.com/BerriAI/litellm/pull/13791)
|
||||
- **[AIM](../../docs/integrations/aim)**
|
||||
- Migrate to new firewall API - [PR #13748](https://github.com/BerriAI/litellm/pull/13748)
|
||||
- **[OTEL](../../docs/observability/opentelemetry_integration)**
|
||||
- Add OTEL tracing for actual LLM API call - [PR #13836](https://github.com/BerriAI/litellm/pull/13836)
|
||||
- **[MLFlow](../../docs/observability/mlflow_integration)**
|
||||
- Include predicted output in MLflow tracing - [PR #13795](https://github.com/BerriAI/litellm/pull/13795), s/o @TomeHirata
|
||||
|
||||
|
||||
1. Vector Stores
|
||||
1. Bugs
|
||||
1. Bedrock
|
||||
1. Using LiteLLM Managed Credentials for Query - https://github.com/BerriAI/litellm/pull/13787
|
||||
2.
|
||||
2. Features
|
||||
## Performance / Loadbalancing / Reliability improvements
|
||||
#### Bugs
|
||||
- **[Cooldowns](../../docs/routing#how-cooldowns-work)**
|
||||
- don't return raw Azure Exceptions to client (can contain prompt leakage) - [PR #13529](https://github.com/BerriAI/litellm/pull/13529)
|
||||
- **[Auto-router](../../docs/proxy/auto_routing)**
|
||||
- Ensures the relevant dependencies for auto router existing on LiteLLM Docker - [PR #13788](https://github.com/BerriAI/litellm/pull/13788)
|
||||
- **Model Alias**
|
||||
- Fix calling key with access to model alias - [PR #13830](https://github.com/BerriAI/litellm/pull/13830)
|
||||
|
||||
Final Count = 1
|
||||
|
||||
1. Management Endpoints / UI
|
||||
1. Bugs
|
||||
1. Passthrough
|
||||
1. Fix query passthrough deletion - https://github.com/BerriAI/litellm/pull/13622
|
||||
2. Features
|
||||
1. Models
|
||||
1. Add Search Functionality for Public Model Names in Model Dashboard - https://github.com/BerriAI/litellm/pull/13687
|
||||
2. Auto-Add `azure/` to deployment Name in UI - https://github.com/BerriAI/litellm/pull/13685
|
||||
3. Models page row UI restructure - https://github.com/BerriAI/litellm/pull/13771
|
||||
2. Notifications
|
||||
1. Add new notifications toast UI everywhere - https://github.com/BerriAI/litellm/pull/13813
|
||||
3. Keys
|
||||
1. Fix key edit settings after regenerating a key - https://github.com/BerriAI/litellm/pull/13815
|
||||
2. Require team_id when creating service account keys - https://github.com/BerriAI/litellm/pull/13873
|
||||
3. Filter - show all options on filter option click - https://github.com/BerriAI/litellm/pull/13858
|
||||
4. Usage
|
||||
1. Fix ‘Cannot read properties of undefined’ exception on user agent activity tab - https://github.com/BerriAI/litellm/pull/13892
|
||||
5. SSO
|
||||
1. Free SSO usage for up to 5 users - https://github.com/BerriAI/litellm/pull/13843
|
||||
2.
|
||||
|
||||
Final Count = 10
|
||||
|
||||
1. Logging / Guardrail Integrations
|
||||
1. Bugs
|
||||
1. Bedrock Guardrails
|
||||
1. Add bedrock api key support - https://github.com/BerriAI/litellm/pull/13835
|
||||
2. Features
|
||||
1. Datadog LLM Observability
|
||||
1. Add support for Failure Logging https://github.com/BerriAI/litellm/pull/13726
|
||||
2. Add time to first token, litellm overhead, guardrail overhead latency metrics - https://github.com/BerriAI/litellm/pull/13734
|
||||
3. Add support for tracing guardrail input/output - https://github.com/BerriAI/litellm/pull/13767
|
||||
2. Langfuse OTEL
|
||||
1. Allow using Key/Team Based Logging - https://github.com/BerriAI/litellm/pull/13791
|
||||
3. AIM
|
||||
1. Migrate to new firewall API - https://github.com/BerriAI/litellm/pull/13748
|
||||
4. OTEL
|
||||
1. Add OTEL tracing for actual LLM API call - https://github.com/BerriAI/litellm/pull/13836
|
||||
5. MLFlow
|
||||
1. Include predicted output in MLflow tracing - https://github.com/BerriAI/litellm/pull/13795 s/o @TomeHirata
|
||||
|
||||
Final Count = 8
|
||||
|
||||
1. Performance / Loadbalancing / Reliability improvements
|
||||
1. Bugs
|
||||
1. Cooldowns
|
||||
1. don't return raw Azure Exceptions to client (can contain prompt leakage) - https://github.com/BerriAI/litellm/pull/13529
|
||||
2. Auto-router
|
||||
1. Ensures the relevant dependencies for auto router existing on LiteLLM Docker - https://github.com/BerriAI/litellm/pull/13788
|
||||
3. Model alias
|
||||
1. Fix calling key with access to model alias - https://github.com/BerriAI/litellm/pull/13830
|
||||
4.
|
||||
2. Features
|
||||
1. S3 Caching [doc link]
|
||||
1. Use namespace as prefix for s3 cache - https://github.com/BerriAI/litellm/pull/13704
|
||||
2. Async S3 Caching support (4x RPS improvement) - https://github.com/BerriAI/litellm/pull/13852 s/o @michal-otmianowski
|
||||
2. Model Group header forwarding [doc link]
|
||||
1. reuse same logic as global header forwarding - https://github.com/BerriAI/litellm/pull/13741
|
||||
2. add support for hosted_vllm on UI - https://github.com/BerriAI/litellm/pull/13885
|
||||
3. Performance
|
||||
1. Improve LiteLLM Python SDK RPS by +200 RPS (braintrust import + aiohttp transport fixes) - https://github.com/BerriAI/litellm/pull/13839
|
||||
2. Use O(1) Set lookups for model routing - https://github.com/BerriAI/litellm/pull/13879
|
||||
3. Reduce Significant CPU overhead from litellm_logging.py - https://github.com/BerriAI/litellm/pull/13895
|
||||
4. Improvements for Async Success Handler (Logging Callbacks) - Approx +130 RPS - https://github.com/BerriAI/litellm/pull/13905
|
||||
#### Features
|
||||
- **[S3 Caching](../../docs/proxy/caching)**
|
||||
- Use namespace as prefix for s3 cache - [PR #13704](https://github.com/BerriAI/litellm/pull/13704)
|
||||
- Async S3 Caching support (4x RPS improvement) - [PR #13852](https://github.com/BerriAI/litellm/pull/13852), s/o @[michal-otmianowski](https://github.com/michal-otmianowski)
|
||||
- **Model Group header forwarding**
|
||||
- reuse same logic as global header forwarding - [PR #13741](https://github.com/BerriAI/litellm/pull/13741)
|
||||
- add support for hosted_vllm on UI - [PR #13885](https://github.com/BerriAI/litellm/pull/13885)
|
||||
- **Performance**
|
||||
- Improve LiteLLM Python SDK RPS by +200 RPS (braintrust import + aiohttp transport fixes) - [PR #13839](https://github.com/BerriAI/litellm/pull/13839)
|
||||
- Use O(1) Set lookups for model routing - [PR #13879](https://github.com/BerriAI/litellm/pull/13879)
|
||||
- Reduce Significant CPU overhead from litellm_logging.py - [PR #13895](https://github.com/BerriAI/litellm/pull/13895)
|
||||
- Improvements for Async Success Handler (Logging Callbacks) - Approx +130 RPS - [PR #13905](https://github.com/BerriAI/litellm/pull/13905)
|
||||
|
||||
|
||||
Final Count = 11
|
||||
## General Proxy Improvements
|
||||
#### Bugs
|
||||
|
||||
1. General Proxy Improvements
|
||||
1. Bugs
|
||||
1. Fix litellm compatibility with newest release of openAI (>v1.100.0) - https://github.com/BerriAI/litellm/pull/13728
|
||||
2. Helm
|
||||
1. Add possibility to configure resources for migrations-job - https://github.com/BerriAI/litellm/pull/13617
|
||||
2. Ensure Helm chart auto generated master keys follow sk-xxxx format - https://github.com/BerriAI/litellm/pull/13871
|
||||
3. Enhance database configuration: add support for optional endpointKey - https://github.com/BerriAI/litellm/pull/13763
|
||||
3. Rate Limits
|
||||
1. fixing descriptor/response size mismatch on parallel_request_limiter_v3 - https://github.com/BerriAI/litellm/pull/13863 s/o luizrennocosta
|
||||
4. Non-root
|
||||
1. fix permission access on prisma migrate in non-root image - https://github.com/BerriAI/litellm/pull/13848 s/o @Ithanil
|
||||
2.
|
||||
2. Features
|
||||
|
||||
|
||||
Final Count = 6
|
||||
|
||||
|
||||
Total = 66
|
||||
- **SDK**
|
||||
- Fix litellm compatibility with newest release of openAI (>v1.100.0) - [PR #13728](https://github.com/BerriAI/litellm/pull/13728)
|
||||
- **Helm**
|
||||
- Add possibility to configure resources for migrations-job - [PR #13617](https://github.com/BerriAI/litellm/pull/13617)
|
||||
- Ensure Helm chart auto generated master keys follow sk-xxxx format - [PR #13871](https://github.com/BerriAI/litellm/pull/13871)
|
||||
- Enhance database configuration: add support for optional endpointKey - [PR #13763](https://github.com/BerriAI/litellm/pull/13763)
|
||||
- **Rate Limits**
|
||||
- fixing descriptor/response size mismatch on parallel_request_limiter_v3 - [PR #13863](https://github.com/BerriAI/litellm/pull/13863), s/o @[luizrennocosta](https://github.com/luizrennocosta)
|
||||
- **Non-root**
|
||||
- fix permission access on prisma migrate in non-root image - [PR #13848](https://github.com/BerriAI/litellm/pull/13848), s/o @[Ithanil](https://github.com/Ithanil)
|
||||
Reference in New Issue
Block a user