docs(gateway): describe tensor-parallel device groups in the architecture guide - #299
Merged
Merged
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 10 included reviews per hour; 7 remain after this review. 📝 WalkthroughWalkthroughThe architecture guide now states that a model remains assigned to one child but may use multiple GPUs within that child when device grouping and tensor parallelism are configured. ChangesGeneration placement
Suggested reviewers: Priority: ⬇️ Low Merge Risk: ⚪ Minimal · up to This documentation clarification introduces no established production or deployment risk. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
The gateway architecture guide's "Multi-GPU caveat" still said a generation model is never sharded across GPUs. Since #282, a pool with
gpu.deviceGroup: truerenders one worker child that owns every GPU in the pod, and a profile that declarestensor_parallel_sizeruns one engine across those GPUs.The paragraph now says what is true:
Docs only. The wording matches the "Tensor-Parallel Device Groups" section of
deploy/helm/sie-cluster/README.md.Summary by CodeRabbit