Skip to content

Fix gemma-trainer docs and dataset utilities - #7

Open
houyongsheng wants to merge 1 commit into
google-gemma:mainfrom
houyongsheng:fix/trainer-docs-scripts
Open

Fix gemma-trainer docs and dataset utilities#7
houyongsheng wants to merge 1 commit into
google-gemma:mainfrom
houyongsheng:fix/trainer-docs-scripts

Conversation

@houyongsheng

Copy link
Copy Markdown

Summary

  • align Gemma SFT examples and distilled records with the standard assistant role
  • fix Hugging Face distillation quantization and temperature handling
  • let dataset validation fall back to a character-count heuristic when no tokenizer is available
  • update quantization, context-window, reward-model, and deployment guidance

Testing

  • Python unittest suite: 5 passed
  • AST parsing: all 5 gemma-trainer asset scripts passed
  • CLI help smoke tests for dataset preparation and distillation

No model weights were downloaded and no GPU training run was performed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant