Repository navigation
Upgrade translation model to NLLB-200-distilled-1.3B - #12
Merged
Merged
Conversation
Replace M2M100_1.2B with facebook/nllb-200-distilled-1.3B: - Similar size (1.2B → 1.3B parameters) - Same VRAM footprint (~5GB) - Better translation quality with more modern architecture (2022) - Supports 200 languages with improved accuracy - Similar speed (~4-5s per segment) This is a safe upgrade that improves quality without requiring more resources. Changes: - Updated TRANSLATION_MODEL default to nllb-200-distilled-1.3B - Added NLLB language code mapping (eng_Latn, deu_Latn, etc.) - Added get_translation_lang_code() method to handle code conversion - Switched from M2M100-specific imports to AutoTokenizer/AutoModel - Updated translate_text() to use mapped language codes Expected VRAM usage: ~13-14GB total (same as before) Expected performance: ~4-5s per translation (similar to m2m100) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
etiennechabert
force-pushed
the
feature/upgrade-nllb-translation-model
branch
from
March 16, 2026 18:21
618e072 to
f69523f
Compare
NLLB tokenizers use lang_code_to_id dictionary instead of get_lang_id() method. Added compatibility check to support both NLLB and M2M100 tokenizers. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Use try-except approach instead of hasattr for more robust tokenizer detection. Also properly handle src_lang parameter for NLLB tokenizers. Changes: - NLLB: Pass src_lang as parameter to tokenizer call - M2M100: Set src_lang as tokenizer attribute - Try lang_code_to_id first, fall back to get_lang_id - Better error message if neither method works Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
NLLB tokenizers use the same approach as M2M100: set src_lang as an attribute, not as a parameter. Both NLLB and M2M100 tokenizers: - Set tokenizer.src_lang = src_code (attribute) - Use lang_code_to_id (NLLB) or get_lang_id() (M2M100) for target Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Frontend changes: - Removed extra languages from admin.html (only EN, DE) - Removed extra languages from viewer.html (only EN, DE) - Updated language flags mapping to match Backend changes: - Added debug logging to show available NLLB language codes - Will help identify correct language code format for NLLB tokenizer Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Print available language codes when loading NLLB tokenizer to help identify the correct code format. Will show at startup which codes are available and whether our expected codes exist. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
- Use convert_tokens_to_ids() instead of lang_code_to_id attribute - Update startup debug to test language codes correctly - Fixes translation errors with eng_Latn and deu_Latn codes Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Upgrade from M2M100_1.2B to the more modern NLLB-200-distilled-1.3B translation model for better quality without requiring additional resources.
What's Changing
Translation Model Upgrade
Old:
facebook/m2m100_1.2B(2020, 1.2B parameters, ~5GB VRAM)New:
facebook/nllb-200-distilled-1.3B(2022, 1.3B parameters, ~5GB VRAM)Benefits
Why distilled-1.3B instead of 3.3B?
Technical Changes
Config (config.py)
TRANSLATION_MODELdefault tofacebook/nllb-200-distilled-1.3BNLLB_LANG_CODE_MAPfor language code conversionget_translation_lang_code()classmethod to handle mappingBackend (app.py)
AutoTokenizer/AutoModelForSeq2SeqLMtranslate_text()to use NLLB language codes (e.g.,eng_Latn,deu_Latn)Language Code Mapping
NLLB uses different codes than M2M100:
en→eng_Latnde→deu_Latnes→spa_Latnfr→fra_LatnThe mapping is handled automatically via
Config.get_translation_lang_code().VRAM Impact
Before:
After (SAME!):
Testing Notes
When testing:
Future Upgrades Available
With VRAM headroom still available, future options include:
🤖 Generated with Claude Code