* Add new model provider Novita AI (#7582)
* feat: add new model provider Novita AI
* feat: use deepseek r1 model for examples in Novita AI docs
* fix: fix tests
* fix: fix tests for novita
* fix: fix novita transformation
* ci: fix ci yaml
* fix: fix novita transformation and test (#10056)
---------
Co-authored-by: Jason <ggbbddjm@gmail.com>
* fix(factory.py): Add reasoning content handling for missing assistant content
* fix(factory.py): Improve handling of thinking blocks for assistant content
* test(factory.py): Add test for Bedrock processing of thinking blocks with None content
* Fixed Json.dumps in JSON Schema Validation Error
* Added Response Schema to Ollama chat for structured response
* Added Test cases
* refactor(ollama): remove redundant response_format check
The response_format parameter conversion is already handled in utils.py's
get_optional_params function, making the duplicate check in ollama_chat.py
unnecessary. This change removes the redundant code while maintaining the
same functionality.
* Support pdf url's to openai (#10640)
* fix(gpt_transformation.py): support pdf url input to openai
pass as base64 as openai doesn't support image url's
* fix(openai.py): support async message transformation
allows async get request to convert url to base64
* fix(gpt_transformation.py): fix linting errrors and use common components across sync + async flows
* fix: fix linting errors
* fix(openai.py): pop correct var
* Fix sagemaker chat calls - content length error (#10607)
* fix(sagemaker_chat/): support passing dynamic aws params
previously being ignored
* refactor(sagemaker/chat): more refactoring
* fix(sagemaker_chat/): make sure streaming is correctly handled post-refactor
* refactor: more refactoring to support using signed json str
* fix(sagemaker/chat): working sync streaming post refactor
* fix(sagemaker/chat): support async streaming post refactor
* fix(llm_http_handler.py): await async function
* fix: remove print statements
* test: update test
* test: update test
* fix(llm_http_handler.py): retain passing in data as json str
* test: update test
* fix(base_model_iterator.py): fix linting error
* test: test auth
* fix: fix linting error
* test: update test
* test: update translation test
* fix(gpt_transformation.py): handle awaitable/non-awaitable object
* fix: handle async flow for message transformation on openai compatible api's
* test: cleanup testing
* test: update test
* test(test_router.py): use model with higher quota
* test: simplify test
* test: update test
* fix(auth_checks.py): enforce auth checks on target model names
ensures user has access to models they are trying to call
* test(test_auth_utils.py): add unit tests for auth check
* fix(exception_mapping_utils.py): handle mistral 429 exception
* fix: fix linting error
* fix(auth_checks.py): add max fallback depth
* feat(router.py): translate the model in jsonl for create file deployment to use the deployment model name
* test: add unit test for replace model in jsonl
* test(test_router.py): add unit tests
* test: add unit tests
* fix(router.py): write file to all deployments
allows unified file id to work across multiple deployments
* fix(view_logs/index.tsx): show call type in request logs
* fix(router.py): pass a deep copy of kwargs to avoid conflict across multiple runs
* fix(batch_utils.py): broaden check
* fix(router_utils.py): handle null type for function name
* fix(proxy_track_cost_callback.py): fix ruff check error
* fix(router.py): handle healthy_deployments as a dict
* feat(managed_files.py): support encoding / decoding unified batch id … (#10711)
* feat(managed_files.py): support encoding / decoding unified batch id when using managed files
allows routing retrieve batch to the right model id
* fix: fix linting error
* feat(managed_files.py): support unified output file id
enables batch output file id to be used to retrieve the actual file
* fix(managed_files.py): attempt to fix ci/cd linting error
* fix: fix ruff check
* fix(router.py): write file to all deployments
allows unified file id to work across multiple deployments
* fix(view_logs/index.tsx): show call type in request logs
* fix(router.py): pass a deep copy of kwargs to avoid conflict across multiple runs
* fix(batch_utils.py): broaden check
* fix(router_utils.py): handle null type for function name
* fix(proxy_track_cost_callback.py): fix ruff check error
* fix(router.py): handle healthy_deployments as a dict
* feat(managed_files.py): support encoding / decoding unified batch id … (#10711)
* feat(managed_files.py): support encoding / decoding unified batch id when using managed files
allows routing retrieve batch to the right model id
* fix: fix linting error
* test: add unit tests
* fix: fix ruff check
* Azure LLM: fix passing through of azure_ad_token_provider parameter
* add test
---------
Co-authored-by: Clara Luise Pohland <clara-luise.pohland@telekom.de>
* fix(caching_handler.py): fix embedding str caching result
Fixes issue where str caching results were not being correctly assembled on str input
* feat(azure/image_generation): Support dropping response_format for azure gpt-image-1
Fixes LIT-118
* test(test_utils.py): add unit testing
* test: rename file to avoid testing conflict
* Add --version flag to litellm-proxy CLI
```shell
$ litellm-proxy --version
litellm-proxy version: 1.68.1
```
* Return both client and server version
* Update docs
* Add a test for the version command
* Add litellm/proxy/client/health.py
* fix support for python 3.11-
3.11 introduced datetime.UTC, this provides a fallback for 3.11-
* use litellm.utils.get_utc_datetime
* remove unused timezone import
Co-authored-by: Matthew Farrellee <matt@cs.wisc.edu>
* test(base_llm_unit_tests.py): return '<thinking>' tag in response content
* fix(converse_transformation.py): extract `<thinking>` block from nova tool use response
Fixes https://github.com/BerriAI/litellm/issues/9063
* fix(factory.py): handle non-signature reasoning blocks to bedrock
pass as text input - bedrock raises ""User messages cannot contain reasoning content. Please remove the r
easoning content and try again." otherwise
* fix(main.py): Add drop params support for gpt
Fixes https://github.com/BerriAI/litellm/issues/10501
* fix(converse_transformation.py): fix linting error
* fix(utils.py): fix linting error
* test: cleanup test
* test: skip test until we have bedrock prompt caching permission
* fix(user_api_key_auth.py): add 'headers' to constructed request for websocket
Fix issue on some datastructure versions which require a headers field in scope
* test(test_user_api_key_auth.py): add unit testing for headers in scope change
* fix(router.py): migrate `_arealtime` to generic router endpoint
Fix infinite loop on model name missing for realtime api calls
* test(test_router_helper_utils.py): cleanup test post refactor
* Add support for nscale provider
* Add image generation support and fix unit tests
* Add docs for nscale
* Fix unit test import issues
* Minor doc improvement
* Remove redundant null tokens from model cost map
* Address PR review comments for doc updates
* Revert changes to large text
* Try to add prints
* more print
* fix print
* batch debugging
* Change batch size
* Revert "Change batch size"
This reverts commit af16d8635f17e3928da9000bb9fec5b0af920816.
* Look into periodic task problems
* Fix missing periodic in slack init
* Initialize periodic flush on startup
* Log update_values
* more logging
* Fix startup
* Cleanup change
* Add a unit test for the change
* Renamed and moved to standard
* Merging in with new test
* comment change
* Extend the timeout because normal runs are over 5 min
* refactor(managed_files.py): move enterprise feature into enterprise folder
prevent unexpected surprises
* refactor: safely handle enterprise hooks
* fix: fix ruff check errors
* fix(files_endpoints.py): cleanup enterprise code from OSS
* refactor: complete cleanup
* fix(managed_files.py): complete cleanup
* fix(managed_files.py): instrument to be able to update deployment values post-router selection and just before making llm call
* fix(managed_files.py): instrument to be able to update deployment values post-router selection and just before making llm call
* fix: fix linting error
* fix: fix linting error
* working email integration
* fix get_custom_loggers_for_type
* add SendKeyCreatedEmailEvent type
* bug fix, only send 1 email when creating key for user
* polish for emails for key created
* polish for key created email
* fix test_init_custom_logger_compatible_class_as_callback
* testing resend email integration
* working user invitation email
* working user invite emails
* testing for user invite emails
* testing fixes for email integration
* working email integration
* fix get_custom_loggers_for_type
* add SendKeyCreatedEmailEvent type
* bug fix, only send 1 email when creating key for user
* polish for emails for key created
* polish for key created email
* fix test_init_custom_logger_compatible_class_as_callback
* testing resend email integration
* testing fixes for email integration