4 Commits
Author SHA1 Message Date
viettranx 04a9938f4f feat(tts): wire tenant timeout + fix Gemini text-only 400
- HTTP synthesize + test-connection now read tenant tts.timeout_ms
  (default 120s, was hardcoded 15s/10s). Gemini client default also
  bumped 30s→120s so both layers align when tenant config unset.
- Inline prefix "Speak naturally: " prepended to single-voice text;
  multi-speaker transcripts pass through unchanged.
- ErrTextOnlyResponse sentinel for 400 "text generation" bodies;
  single-voice retries once with stronger prefix. Narrowed needle
  list avoids false positives on unrelated 400s.
- SynthesizeWithFallbackAdapted now returns errors.Join so sentinel
  survives fallback chain; HTTP 422 mapping + locale-translated
  ForLLM in agent tool (EN/VI/ZH catalogs).
- Default Gemini model bumped to gemini-3.1-flash-tts-preview.
2026-04-23 08:31:53 +07:00
viettranx 613b6e38d7 feat(tts): provider capabilities schema + Gemini TTS + dynamic param forms
Baseline groundwork for TTS expansion plan. Introduces a capabilities
system that the follow-up plan (phase-01..04) builds on.

Backend:
- `audio.ParamSchema` / `ProviderCapabilities` types + per-provider
  capabilities.go for edge, elevenlabs, minimax, openai, gemini.
- Gemini TTS provider (client, models, voices, wav encoder,
  multi-speaker via SpeakerVoice, audio-tag aware prompts).
- TTSOptions gains `Params map[string]any` (read-only) + `Speakers`.
- `VoiceListProvider` interface decouples HTTP voice handler from
  provider-specific impls.
- `nested_keys.go` resolves dot-separated param paths for nested
  provider bodies (voice_settings.stability etc.).
- Characterization + defaults-invariant tests per provider lock
  nil-params byte-equivalence before new params land.
- `/v1/tts/capabilities` HTTP endpoint + integration coverage.
- Dual-read tests (PG + SQLite) for tts_config.

Frontend (web + desktop):
- `DynamicParamForm` with depends-on evaluation, split into
  fields/logic modules. Slider primitive added.
- `AudioTagPicker` + `MultiSpeakerEditor` for Gemini.
- `voice-picker` refactored toward portal UX; combobox tightened.
- tts-capabilities API client + typed hooks.
- i18n catalogs (en/vi/zh) expanded; parity tests guard key drift.
- TTS page reorganised (voice-model-section removed; playground +
  credentials + provider-setup split cleanly).

Docs: codebase-summary, project-changelog, tts-provider-capabilities.
2026-04-20 00:05:19 +07:00
viettranx a050079ade fix(test): resolve data race in TTS concurrent test mock
Add stateless flag to mockTTSProvider to skip capturing opts
during concurrent tests. Prevents race on capturedOpts field.
2026-04-16 18:33:02 +07:00
viettranx 6659ba46a8 feat(tts): consolidate TTS config UI with provider-aware pickers and synthesize endpoint
- Add 4-provider catalog (OpenAI, ElevenLabs, Edge, MiniMax) with static
  voice/model data for web and desktop
- Rewrite /tts page as 5 modular sections: provider setup, credentials,
  voice/model picker, test playground, collapsible advanced settings
- Add POST /v1/tts/synthesize endpoint with 15s timeout, 500-char cap,
  RoleOperator auth, ElevenLabs model validation, rate limit reuse
- Fix Edge TTS provider to honor per-request opts.Voice instead of
  always using construction-time default (C1)
- Gate agent detail TTS section on global provider config with
  empty-state CTA for owners, catalog-driven model select, override
  checkbox, and inline voice test button
- Remove TTS from builtin-tools dialog (STT preserved); hide tts tool
  from builtin-tools page list
- Add desktop parity: provider-aware VoicePicker, TTS config hook via
  WS config.get, empty-state component, synced i18n locales (en/vi/zh)
- Add 109 web vitest tests + 12 Go unit tests covering handler
  validation, legacy config regression, and catalog parity
2026-04-16 14:17:48 +07:00