diff --git a/skills/firebase-ai-logic-basics/references/sdk/capabilities/text_to_speech.md b/skills/firebase-ai-logic-basics/references/sdk/capabilities/text_to_speech.md index 5f6d04e..f8e3658 100644 --- a/skills/firebase-ai-logic-basics/references/sdk/capabilities/text_to_speech.md +++ b/skills/firebase-ai-logic-basics/references/sdk/capabilities/text_to_speech.md @@ -1,10 +1,14 @@ # Client-Side Text-to-Speech (TTS) Generation with Gemini +> [!WARNING] **Preview:** Using the Firebase AI Logic SDKs for text-to-speech +> (TTS) generation is in Preview and may change in backwards-incompatible ways. + Firebase AI Logic enables client-side Text-to-Speech (TTS) generation directly -from your mobile and web applications without maintaining custom backend speech -services. Using specialized Gemini TTS models, apps can synthesize natural, -expressive audio with custom voice personas, multi-speaker dialogues, and -in-flight audio streaming. +from your Android, iOS, Flutter, and Web applications without maintaining custom +backend speech services. Using Gemini TTS models, apps can synthesize +controllable speech from exact text transcripts with single- or two-speaker +voices, natural-language style guidance, inline audio tags, and low-latency +streaming. ______________________________________________________________________ @@ -12,169 +16,222 @@ ______________________________________________________________________ Always specify a dedicated Gemini TTS model when generating audio: -| Model ID | Description | -| :----------------------------- | :--------------------------------------------------------------------------------- | -| `gemini-3.1-flash-tts-preview` | Low-latency speech synthesis (preview); supports single- and multi-speaker output. | - -> [!WARNING] Always verify currently supported model availability in the -> [Firebase AI Logic Models documentation](https://firebase.google.com/docs/ai-logic/models.md.txt). +| Model ID | Description | +| :----------------------------- | :------------------------------------------------------------------------------------------------------------ | +| `gemini-3.1-flash-tts-preview` | Low-latency speech synthesis (preview); supports single-speaker, 2-speaker dialogue, streaming, and `[tags]`. | + +> [!NOTE] Firebase AI Logic also supports Gemini 2.x TTS models, but streaming +> (`generateContentStream`), inline audio tags (`[whispers]`, `[laughs]`), and +> expanded auto-detected languages are only supported on Gemini 3.x TTS models +> (`gemini-3.1-flash-tts-preview`). Always check the +> [Firebase AI Logic Models documentation](https://firebase.google.com/docs/ai-logic/models.md.txt) +> for newly released TTS models. + +> [!IMPORTANT] **Model-specific API differences +> (`gemini-3.1-flash-tts-preview`):** +> +> - **Raw PCM output on both unary and streaming calls:** Both `generateContent` +> and `generateContentStream` return **headerless raw 16-bit linear PCM** +> (`audio/pcm`, 24 kHz, mono, little-endian). Do **not** assume unary +> `generateContent` responses include a WAV header — always play via a raw PCM +> API (`AudioTrack`, `AVAudioEngine`, Web Audio `AudioContext`) or prepend a +> 44-byte WAV (RIFF) header before passing bytes to `MediaPlayer`, +> `AVAudioPlayer`, `