Repository navigation
feat(firebase-ai-logic-basics): align TTS guide with gemini-3.1-flash-tts-preview and add Flutter - #215
feat(firebase-ai-logic-basics): align TTS guide with gemini-3.1-flash-tts-preview and add Flutter#215AustinBenoit wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Code Review
This pull request updates the client-side Text-to-Speech (TTS) documentation for Gemini, detailing configurations, model-specific API differences, and providing updated platform code snippets for Android, iOS, Flutter, and Web. The review feedback points out several critical issues in the Flutter (Dart) code snippets, including unsafe null and empty checks on the generative response candidates and parts, as well as the use of BytesBuilder from dart:io which breaks compatibility with Flutter Web. Suggestions are provided to make the Dart examples robust, null-safe, and web-compatible.
| final response = await model.generateContent([Content.text(prompt)]); | ||
|
|
||
| final part = response.candidates.first.content.parts.first; | ||
| if (part is InlineDataPart && part.mimeType.startsWith('audio/')) { | ||
| final Uint8List pcmData = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit) | ||
| return addWavHeader(pcmData); | ||
| } | ||
| return null; |
There was a problem hiding this comment.
In Dart, response.candidates.first.content is nullable (Content?). Accessing .parts directly on it without a null check will cause a compile-time error or runtime crash. Additionally, candidates or parts could be empty, so accessing .first directly is unsafe. Use explicit null and empty checks to ensure robust and safe execution.
| final response = await model.generateContent([Content.text(prompt)]); | |
| final part = response.candidates.first.content.parts.first; | |
| if (part is InlineDataPart && part.mimeType.startsWith('audio/')) { | |
| final Uint8List pcmData = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit) | |
| return addWavHeader(pcmData); | |
| } | |
| return null; | |
| final response = await model.generateContent([Content.text(prompt)]); | |
| if (response.candidates.isNotEmpty) { | |
| final content = response.candidates.first.content; | |
| if (content != null && content.parts.isNotEmpty) { | |
| final part = content.parts.first; | |
| if (part is InlineDataPart && part.mimeType.startsWith('audio/')) { | |
| final Uint8List pcmData = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit) | |
| return addWavHeader(pcmData); | |
| } | |
| } | |
| } | |
| return null; |
| Future<Uint8List> streamSpeech(String prompt) async { | ||
| final model = FirebaseAI.googleAI().generativeModel( | ||
| model: 'gemini-3.1-flash-tts-preview', | ||
| generationConfig: GenerationConfig( | ||
| responseModalities: [ResponseModalities.audio], | ||
| speechConfig: SpeechConfig(voiceName: 'Kore'), | ||
| ), | ||
| ); | ||
|
|
||
| final pcm = BytesBuilder(copy: false); | ||
| final responseStream = model.generateContentStream([Content.text(prompt)]); | ||
|
|
||
| await for (final chunk in responseStream) { | ||
| final part = chunk.candidates.first.content.parts.first; | ||
| if (part is InlineDataPart && part.mimeType.startsWith('audio/')) { | ||
| final Uint8List pcmChunk = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit) | ||
| // Low-latency path: feed `pcmChunk` directly to a raw-PCM audio stream. | ||
| pcm.add(pcmChunk); | ||
| } | ||
| } | ||
|
|
||
| // Buffered path: prepend a 44-byte WAV header for standard audio players. | ||
| return addWavHeader(pcm.takeBytes()); | ||
| } |
There was a problem hiding this comment.
Using BytesBuilder from dart:io prevents this code from compiling or running on Flutter Web, as dart:io is not supported on the web platform. Additionally, chunk.candidates.first.content is nullable and can cause runtime crashes if accessed unsafely. We can accumulate the chunks in a standard List<Uint8List> and concatenate them using standard Uint8List methods, which is highly efficient, robust, and fully web-safe.
Future<Uint8List> streamSpeech(String prompt) async {
final model = FirebaseAI.googleAI().generativeModel(
model: 'gemini-3.1-flash-tts-preview',
generationConfig: GenerationConfig(
responseModalities: [ResponseModalities.audio],
speechConfig: SpeechConfig(voiceName: 'Kore'),
),
);
final chunks = <Uint8List>[];
final responseStream = model.generateContentStream([Content.text(prompt)]);
await for (final chunk in responseStream) {
if (chunk.candidates.isNotEmpty) {
final content = chunk.candidates.first.content;
if (content != null && content.parts.isNotEmpty) {
final part = content.parts.first;
if (part is InlineDataPart && part.mimeType.startsWith('audio/')) {
final Uint8List pcmChunk = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit)
// Low-latency path: feed pcmChunk directly to a raw-PCM audio stream.
chunks.add(pcmChunk);
}
}
}
}
// Buffered path: prepend a 44-byte WAV header for standard audio players.
final totalLength = chunks.fold<int>(0, (sum, chunk) => sum + chunk.length);
final pcm = Uint8List(totalLength);
var offset = 0;
for (final chunk in chunks) {
pcm.setRange(offset, offset + chunk.length, chunk);
offset += chunk.length;
}
return addWavHeader(pcm);
}| return (BytesBuilder(copy: false) | ||
| ..add(headerBytes) | ||
| ..add(pcmBytes)) | ||
| .takeBytes(); | ||
| } |
There was a problem hiding this comment.
Avoid using BytesBuilder from dart:io to ensure compatibility with Flutter Web. Instead, allocate a single Uint8List of the combined size and copy the header and PCM bytes into it using setRange. This is fully web-safe and highly performant.
| return (BytesBuilder(copy: false) | |
| ..add(headerBytes) | |
| ..add(pcmBytes)) | |
| .takeBytes(); | |
| } | |
| final wavBytes = Uint8List(44 + pcmBytes.length); | |
| wavBytes.setRange(0, 44, headerBytes); | |
| wavBytes.setRange(44, wavBytes.length, pcmBytes); | |
| return wavBytes; | |
| } |
2cf6cf6 to
6040941
Compare
6040941 to
67098ef
Compare
67098ef to
51625fa
Compare
51625fa to
af68e9d
Compare
…hConfig with SDKs ### Summary (Stack 6/6) Completes `references/sdk/capabilities/text_to_speech.md` with Flutter (`firebase_ai`) coverage and aligns the Swift and Kotlin `SpeechConfig` constructors with the current SDK signatures. ### Changes - Adds `## Flutter (Dart)` section to `text_to_speech.md` (single-speaker `SpeechConfig`, `SpeechConfig.multiSpeaker`, streaming PCM accumulation, and pure-Dart WAV header helper). - Updates Swift (`SpeechConfig(voiceName:languageCode:)`) and Kotlin (`@OptIn(PublicPreviewAPI::class) SpeechConfig(voice = Voice(...))`) snippets to match `firebase-ios-sdk` and `firebase-android-sdk`.
af68e9d to
575737b
Compare
Summary (Stack 5/5)
Aligns
references/sdk/capabilities/text_to_speech.mdwith the official Firebase AI Logic Speech Generation guide and the Gemini TTS migration rules forgemini-3.1-flash-tts-preview.Changes
generateContentresponses include a 44-byte WAV header (audio/x-wav). Documents thatgemini-3.1-flash-tts-previewreturns headerless 24 kHz, 16-bit, mono linear PCM (audio/pcm) for both unary and streaming requests.SpeechConfig&MultiSpeakerVoiceConfigsnippets: Updates iOS (Swift), Android (Kotlin), and Web (JS/TS) snippets to match the SDKs and adds## Flutter (Dart)coverage.gemini-3.1-flash-tts-previewconstraints: Adds[Sample Context], English-only square-bracket audio tag rules, mixed-language multi-speaker guidance, and the 4 model constraints fromgenerate-speech#model-constraints.