Skip to content

feat(firebase-ai-logic-basics): align TTS guide with gemini-3.1-flash-tts-preview and add Flutter - #215

Open
AustinBenoit wants to merge 1 commit into
ai-logic/5-image-and-hybridfrom
ai-logic/6-tts-flutter
Open

AustinBenoit wants to merge 1 commit into
ai-logic/5-image-and-hybridfrom
ai-logic/6-tts-flutter

Conversation

@AustinBenoit

@AustinBenoit AustinBenoit commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary (Stack 5/5)

Aligns references/sdk/capabilities/text_to_speech.md with the official Firebase AI Logic Speech Generation guide and the Gemini TTS migration rules for gemini-3.1-flash-tts-preview.

Changes

  • Raw PCM output format (unary & streaming): Removes the 3.8-only claim that unary generateContent responses include a 44-byte WAV header (audio/x-wav). Documents that gemini-3.1-flash-tts-preview returns headerless 24 kHz, 16-bit, mono linear PCM (audio/pcm) for both unary and streaming requests.
  • SDK-verified SpeechConfig & MultiSpeakerVoiceConfig snippets: Updates iOS (Swift), Android (Kotlin), and Web (JS/TS) snippets to match the SDKs and adds ## Flutter (Dart) coverage.
  • Prompting & gemini-3.1-flash-tts-preview constraints: Adds [Sample Context], English-only square-bracket audio tag rules, mixed-language multi-speaker guidance, and the 4 model constraints from generate-speech#model-constraints.

@AustinBenoit
AustinBenoit added this pull request to stack #216 October 9, 2026 14:39

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the client-side Text-to-Speech (TTS) documentation for Gemini, detailing configurations, model-specific API differences, and providing updated platform code snippets for Android, iOS, Flutter, and Web. The review feedback points out several critical issues in the Flutter (Dart) code snippets, including unsafe null and empty checks on the generative response candidates and parts, as well as the use of BytesBuilder from dart:io which breaks compatibility with Flutter Web. Suggestions are provided to make the Dart examples robust, null-safe, and web-compatible.

Comment on lines +707 to +714
final response = await model.generateContent([Content.text(prompt)]);

final part = response.candidates.first.content.parts.first;
if (part is InlineDataPart && part.mimeType.startsWith('audio/')) {
final Uint8List pcmData = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit)
return addWavHeader(pcmData);
}
return null;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In Dart, response.candidates.first.content is nullable (Content?). Accessing .parts directly on it without a null check will cause a compile-time error or runtime crash. Additionally, candidates or parts could be empty, so accessing .first directly is unsafe. Use explicit null and empty checks to ensure robust and safe execution.

Suggested change
final response = await model.generateContent([Content.text(prompt)]);
final part = response.candidates.first.content.parts.first;
if (part is InlineDataPart && part.mimeType.startsWith('audio/')) {
final Uint8List pcmData = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit)
return addWavHeader(pcmData);
}
return null;
final response = await model.generateContent([Content.text(prompt)]);
if (response.candidates.isNotEmpty) {
final content = response.candidates.first.content;
if (content != null && content.parts.isNotEmpty) {
final part = content.parts.first;
if (part is InlineDataPart && part.mimeType.startsWith('audio/')) {
final Uint8List pcmData = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit)
return addWavHeader(pcmData);
}
}
}
return null;

Comment on lines +753 to +776
Future<Uint8List> streamSpeech(String prompt) async {
final model = FirebaseAI.googleAI().generativeModel(
model: 'gemini-3.1-flash-tts-preview',
generationConfig: GenerationConfig(
responseModalities: [ResponseModalities.audio],
speechConfig: SpeechConfig(voiceName: 'Kore'),
),
);

final pcm = BytesBuilder(copy: false);
final responseStream = model.generateContentStream([Content.text(prompt)]);

await for (final chunk in responseStream) {
final part = chunk.candidates.first.content.parts.first;
if (part is InlineDataPart && part.mimeType.startsWith('audio/')) {
final Uint8List pcmChunk = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit)
// Low-latency path: feed `pcmChunk` directly to a raw-PCM audio stream.
pcm.add(pcmChunk);
}
}

// Buffered path: prepend a 44-byte WAV header for standard audio players.
return addWavHeader(pcm.takeBytes());
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using BytesBuilder from dart:io prevents this code from compiling or running on Flutter Web, as dart:io is not supported on the web platform. Additionally, chunk.candidates.first.content is nullable and can cause runtime crashes if accessed unsafely. We can accumulate the chunks in a standard List<Uint8List> and concatenate them using standard Uint8List methods, which is highly efficient, robust, and fully web-safe.

Future<Uint8List> streamSpeech(String prompt) async {
  final model = FirebaseAI.googleAI().generativeModel(
    model: 'gemini-3.1-flash-tts-preview',
    generationConfig: GenerationConfig(
      responseModalities: [ResponseModalities.audio],
      speechConfig: SpeechConfig(voiceName: 'Kore'),
    ),
  );

  final chunks = <Uint8List>[];
  final responseStream = model.generateContentStream([Content.text(prompt)]);

  await for (final chunk in responseStream) {
    if (chunk.candidates.isNotEmpty) {
      final content = chunk.candidates.first.content;
      if (content != null && content.parts.isNotEmpty) {
        final part = content.parts.first;
        if (part is InlineDataPart && part.mimeType.startsWith('audio/')) {
          final Uint8List pcmChunk = part.bytes; // Raw PCM bytes (24kHz, 1 channel, 16-bit)
          // Low-latency path: feed pcmChunk directly to a raw-PCM audio stream.
          chunks.add(pcmChunk);
        }
      }
    }
  }

  // Buffered path: prepend a 44-byte WAV header for standard audio players.
  final totalLength = chunks.fold<int>(0, (sum, chunk) => sum + chunk.length);
  final pcm = Uint8List(totalLength);
  var offset = 0;
  for (final chunk in chunks) {
    pcm.setRange(offset, offset + chunk.length, chunk);
    offset += chunk.length;
  }
  return addWavHeader(pcm);
}

Comment on lines +810 to +814
return (BytesBuilder(copy: false)
..add(headerBytes)
..add(pcmBytes))
.takeBytes();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Avoid using BytesBuilder from dart:io to ensure compatibility with Flutter Web. Instead, allocate a single Uint8List of the combined size and copy the header and PCM bytes into it using setRange. This is fully web-safe and highly performant.

Suggested change
return (BytesBuilder(copy: false)
..add(headerBytes)
..add(pcmBytes))
.takeBytes();
}
final wavBytes = Uint8List(44 + pcmBytes.length);
wavBytes.setRange(0, 44, headerBytes);
wavBytes.setRange(44, wavBytes.length, pcmBytes);
return wavBytes;
}

@AustinBenoit
AustinBenoit removed this pull request from stack #216 October 9, 2026 14:48
@AustinBenoit
AustinBenoit added this pull request to stack #218 October 9, 2026 14:48
@AustinBenoit
AustinBenoit force-pushed the ai-logic/6-tts-flutter branch from 2cf6cf6 to 6040941 Compare October 9, 2026 14:48
@AustinBenoit AustinBenoit changed the title feat(firebase-ai-logic-basics): add Flutter TTS guide and align SpeechConfig with SDKs feat(firebase-ai-logic-basics): align TTS guide with gemini-3.1-flash-tts-preview and add Flutter Oct 9, 2026
@AustinBenoit
AustinBenoit force-pushed the ai-logic/6-tts-flutter branch from 6040941 to 67098ef Compare October 9, 2026 15:23
@AustinBenoit
AustinBenoit force-pushed the ai-logic/6-tts-flutter branch from 67098ef to 51625fa Compare October 9, 2026 15:49
@AustinBenoit
AustinBenoit force-pushed the ai-logic/6-tts-flutter branch from 51625fa to af68e9d Compare October 9, 2026 16:13
…hConfig with SDKs

### Summary (Stack 6/6)
Completes `references/sdk/capabilities/text_to_speech.md` with Flutter (`firebase_ai`) coverage and aligns the Swift and Kotlin `SpeechConfig` constructors with the current SDK signatures.

### Changes
- Adds `## Flutter (Dart)` section to `text_to_speech.md` (single-speaker `SpeechConfig`, `SpeechConfig.multiSpeaker`, streaming PCM accumulation, and pure-Dart WAV header helper).
- Updates Swift (`SpeechConfig(voiceName:languageCode:)`) and Kotlin (`@OptIn(PublicPreviewAPI::class) SpeechConfig(voice = Voice(...))`) snippets to match `firebase-ios-sdk` and `firebase-android-sdk`.
@AustinBenoit
AustinBenoit force-pushed the ai-logic/6-tts-flutter branch from af68e9d to 575737b Compare October 9, 2026 16:51
@AustinBenoit
AustinBenoit removed this pull request from stack #218 October 9, 2026 16:53
@AustinBenoit
AustinBenoit added this pull request to stack #223 October 9, 2026 16:53

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant