Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
147 commits
Select commit Hold shift + click to select a range
b17248a
Move config src files into a dedicated dir (#3570)
maanug-nv Feb 25, 2026
9bb2375
chore: rotate oncall schedule
github-actions[bot] Feb 25, 2026
0dba6a0
Revert "remove encoder_and_decoder from enums (#3406)" (#3579)
ko3n1g Feb 25, 2026
b359981
Fix default cuda graph persist arg. Add persist to rl common.sh. (#3584)
yobibyte Feb 25, 2026
60a25aa
Optimize away add request overheads in dummy ep cuda-graphed forward …
sidsingh-nvidia Feb 25, 2026
e7d80bd
ci: Test docs build (#3583)
ko3n1g Feb 25, 2026
80ef2ae
ci: Fix docs build for release (#3597)
ko3n1g Feb 26, 2026
2561a59
ci: Remove secrets (#3598)
ko3n1g Feb 26, 2026
19444b9
ci: Define secrets (#3599)
ko3n1g Feb 26, 2026
a08ad25
ci: gh-release-from-tag (#3600)
ko3n1g Feb 26, 2026
ac449c4
Ko3n1g/ci/remove twine username (#3601)
ko3n1g Feb 26, 2026
9088d4f
Add training code to MCore wheel (#3573)
maanug-nv Feb 26, 2026
dd4c0b7
FP8 attention knob for nvFP4 recipe (#3363)
vasunvidia Feb 26, 2026
c2782fd
Fix error with --load-main-params-from-ckpt (#3569)
guyueh1 Feb 26, 2026
c921891
ci: Create comment (#3610)
ko3n1g Feb 26, 2026
9100119
ci: Skip cleanup-taint-node jobs during deployments (#3612)
ko3n1g Feb 26, 2026
6161f7a
ci: No comment for release workflow (#3615)
ko3n1g Feb 26, 2026
36d8a9d
ci: Re-add release tag prefix (#3619)
ko3n1g Feb 26, 2026
1b1f5c4
docs: Fix version picker urls (#3621)
chtruong814 Feb 26, 2026
027e0f3
ci: Increase changelog generation max PRs fetched (#3620)
chtruong814 Feb 26, 2026
18c94d2
Add debug info to an assert. (#3588)
yobibyte Feb 26, 2026
e1a9ac9
fix: async_utils: explicit GC in persistent checkpoint worker loop (#…
sbak5 Feb 26, 2026
5f668c1
Fix: Perform sigmoid calculation in fp32 for aux loss stability (#2765)
CodersAcademy006 Feb 26, 2026
d3c10df
fix: forward use_te_activation_func flag in non-MoE GPT layer spec (#…
saakshigupta2002 Feb 26, 2026
7e3f670
Multimodal: Limit transformer version in Dockerfile (#3448)
faradawn Feb 26, 2026
6c0b9c6
Track and plot per-token off-policy in RL (#3515)
tdene Feb 26, 2026
a7c207f
Revert "Add single-process checkpoint save to avoid forked multiproce…
ko3n1g Feb 26, 2026
6503bf8
Multimodal: fix VQA dataset selection (#3464)
faradawn Feb 26, 2026
2f549e5
Multimodal: Fix multimodal training example - tokenizer, Triton Cache…
faradawn Feb 26, 2026
afbce84
Support TP > GQA for inference (#3627)
santhnm2 Feb 26, 2026
310082a
μP: Maximal Update Parameterization (#3058)
plugyawn Feb 26, 2026
7418b1b
Add flexible virtual pipeline parallel (fVPP) to hybrid model (#3377)
duncanriach Feb 26, 2026
53c5973
Explicitly close and join Pool in preprocess_data.py (#3592)
weijiac0619 Feb 27, 2026
36a95ae
remove indexer (#3416)
dimapihtar Feb 27, 2026
f0519b7
Multimodal: add load weights only (#3452)
faradawn Feb 27, 2026
61a293d
Add single-process checkpoint save to avoid forked multiprocessing (#…
sbak5 Feb 27, 2026
6287e7f
Update oncall schedule (#3632)
Phlip79 Feb 27, 2026
93d2739
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Feb 28, 2026
e8fe068
M-FSDP: Cancel erroneous grad accumulation check (#3629)
shjwudp Mar 2, 2026
107f6ae
chore(beep boop 🤖): Bump (main) (2026-03-02)
github-actions[bot] Mar 2, 2026
c7be214
Fix MoE aux loss tracker hang with MTP enabled (#3401)
Victarry Mar 2, 2026
044f1e3
Fix test data preparation (#3652)
janEbert Mar 2, 2026
2f8c9bc
Add GPTOSS Example with Megatron-LM + Megatron Bridge (#3018)
faradawn Mar 2, 2026
63cd60b
Add thd unit test main (#3617)
kunlunl Mar 2, 2026
c9312e6
Inference | KV prefix caching. (#3063)
lmcafee-nvidia Mar 2, 2026
b969f76
[Megatron-FSDP] Add dtype customization to Megatron-FSDP. (#3067)
cspades Mar 2, 2026
0810e63
CachedMetadataFileSystemReader: shared cache (#3326)
sbak5 Mar 2, 2026
7d1c016
Inference Optimized MoEs (#3496)
sidsingh-nvidia Mar 3, 2026
98495af
Log torch_memory_saver offload/onload (#3567)
tdene Mar 3, 2026
6fc7690
Prefix caching | Mamba memory only. (#3657)
lmcafee-nvidia Mar 3, 2026
9b18de4
Prefix caching | Coordinator scheduling. (#3665)
lmcafee-nvidia Mar 3, 2026
7dee32a
Adding manual Claude reviewer (#3679)
Phlip79 Mar 3, 2026
2570947
Nemo-RL Refit (#3520)
wdykas Mar 3, 2026
fa93d79
Add extra permissions and make other changes (#3683)
Phlip79 Mar 4, 2026
77a00ec
Claude should always comment something (#3685)
Phlip79 Mar 4, 2026
2caa681
[Cleanup] Remove the deprecated GroupedMLP (#3410)
dimapihtar Mar 4, 2026
470c6ea
chore: rotate oncall schedule
github-actions[bot] Mar 4, 2026
2a931a3
Fix illegal memory access with mamba inference (#3631)
tdene Mar 4, 2026
bb31e93
Fix illegal memory access with mamba inference (bis) (#3696)
tdene Mar 4, 2026
4fa9b5a
remove duplicate rerun_state_machine.set_mode(rerun_mode) (#3279)
YangWang92 Mar 4, 2026
d3a8584
Correct indexing when cp_comms_type is a list (#3389)
jeromeku Mar 4, 2026
bfd160b
Fix optional chat_completions returnables (#3519)
tdene Mar 4, 2026
f84e84e
ci: Claude code review (#3704)
ko3n1g Mar 4, 2026
bd1406f
ci: Fix event payload (#3705)
ko3n1g Mar 4, 2026
33476ff
ci: Use issue number (#3706)
ko3n1g Mar 4, 2026
fd21af4
ci: Finalize Claude review (#3707)
ko3n1g Mar 4, 2026
b1b4df6
ci: Add codecov yml (#3455)
thomasdhc Mar 4, 2026
7ea354b
Robust signaling for coordinator inference (#3563)
tdene Mar 4, 2026
f91c4bb
Fix memory issue in mxfp8 model init (#3461)
WanZzzzzz Mar 4, 2026
c5d8f6b
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Mar 5, 2026
fa6063d
adding public_docs_features: True to get proper legal footer… (#3681)
megnvidia Mar 4, 2026
a2381d8
add --overlap-param-gather support for layer-wise optimizer. lots of …
mchrzanowski Mar 5, 2026
657d33b
ci: Mount and enforce HF_HOME (#3700)
ko3n1g Mar 5, 2026
da47e64
Add flags for changing Mamba inference state tensor dtype (#3660)
santhnm2 Mar 5, 2026
94a903b
chore: CLI launch internal CI (#3695)
ko3n1g Mar 5, 2026
485428f
Change Review Process (#3659)
Phlip79 Mar 5, 2026
d1cce0c
ci: Separate queues for internal/external contributors (#3718)
ko3n1g Mar 5, 2026
6bd5c12
Update to correct token (#3724)
Phlip79 Mar 5, 2026
f5b2ec0
build: Bump to NGC PyTorch 26.02 (#3474)
ko3n1g Mar 5, 2026
bb84493
Claude: use Opus 4.6 and auto-review on ready (#3727)
Phlip79 Mar 6, 2026
41daf81
Claude to add complexity label (#3709)
Phlip79 Mar 6, 2026
0d42bc6
Offload Flask frontend to separate process (#3648)
santhnm2 Mar 6, 2026
d3528a2
fix(moe): fix TE general_gemm API change (#3582)
hxbai Mar 6, 2026
43df309
Review process fixes (#3728)
Phlip79 Mar 6, 2026
bde8264
ci: Update golden values after PyT bump (#3733)
ko3n1g Mar 6, 2026
a979332
chore: Use PAT for CLI Launcher (#3734)
ko3n1g Mar 6, 2026
bb451db
Print more verbose error message about incorrect `model_parallel_size…
rj42 Mar 6, 2026
de63aa8
ci: Add missing gitlab rule (#3735)
ko3n1g Mar 6, 2026
37ca715
[main] Add TE CUDA Graph Support for Vision Encoder (#3293)
tomlifu Mar 6, 2026
17de0db
Optimize process management and delete operations for async save (#3262)
sbak5 Mar 6, 2026
6ec369d
Align gpt-oss window-size with 128-token sliding window (#2771)
returnL Mar 6, 2026
26f9444
fix: temperature validation error message 1000.0 -> 100.0 (#2688)
CreeperLKF Mar 6, 2026
e19fbe2
RL: Hybrid MoE training cudagraphs and fix training <-> inference tra…
mathemakitten Mar 6, 2026
c1e675f
Fix dynamic inference and GRPO functional tests (#3740)
santhnm2 Mar 6, 2026
0cfa420
Swap oncall (#3585)
janEbert Mar 6, 2026
932b767
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Mar 7, 2026
b09ee64
[bugfix] fix the bug that loss: 0 will not be printed (#1555)
leisuzz Mar 9, 2026
8318b80
Fused dLN + add in backwards pass (#3384)
CarlosGomes98 Mar 9, 2026
597721a
chore(beep boop 🤖): Bump (main) (2026-03-09)
github-actions[bot] Mar 9, 2026
116a7fa
Claude: run actions on target branch (#3745)
Phlip79 Mar 9, 2026
56158bb
revert of #2658 (#3736)
dimapihtar Mar 9, 2026
452fc11
Update README Quick Start (#3596)
ilml Mar 9, 2026
d904a68
Re-enable tests which were failing on #3373 (#3757)
mathemakitten Mar 9, 2026
94ff0dc
Check reviews properly (#3756)
Phlip79 Mar 9, 2026
ce66b22
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Mar 10, 2026
0e19bf1
Add CP + Sequence Packing support for Mimo (#2135)
mehraakash Mar 10, 2026
fca1679
MXFP8 refit (#3742)
wdykas Mar 10, 2026
397772e
Claude: update token usage (#3760)
Phlip79 Mar 10, 2026
3eea580
Handle Tool Call Argument Parsing (#3662)
sancha Mar 10, 2026
e970199
RL support for nanov3 sft checkpoint (#3741)
jon-barker Mar 10, 2026
ba497c9
add mix_hidden_states option in conversion (#3655)
yeyu-nvidia Mar 10, 2026
22c69fa
ci: Optimize release-configs for GB200 (#3541)
ko3n1g Mar 10, 2026
0f47a1a
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Mar 11, 2026
204e7d5
Add absorbed-mla (#3198)
kunlunl Mar 10, 2026
f544034
feat(checkpoint): zero-copy storage sharing in CheckpointWithoutOutpu…
Victarry Mar 11, 2026
8fd390d
Fuse MLA DOWN projection GEMMs (#3039)
cjld Mar 11, 2026
7d52694
fix: skip FSDP DTensor boundary validation under fake process group (…
Victarry Mar 11, 2026
16a8cdb
[main] fix(moe): fix the bug where gate was not sliced when kv_head <…
yuzhongw-nvidia Mar 11, 2026
b8e23d5
fix(offload): reset activation offload manager after eval as well as …
rapatel Mar 11, 2026
07e512a
chore: rotate oncall schedule
github-actions[bot] Mar 11, 2026
e20b89a
Improve error logging when invalid number of tokens is requested. (#3…
yobibyte Mar 11, 2026
5bc89f3
Add NVIDIA-Nemotron-3-Super-120B-A12B-BF16 to ModelOpt examples (#3805)
jenchen13 Mar 11, 2026
da46946
build: Bump TE2.13 (#3800)
ko3n1g Mar 11, 2026
8e64e69
Ensure dummy_forward does not attempt to run cudagraphs (#3789)
jalbericiola Mar 11, 2026
8f539df
Add speculative decoding support with MTP layers (#3594)
santhnm2 Mar 11, 2026
39472d8
Shanmugamr1992/megatron inference ultra (#3784)
shanmugamr1992 Mar 11, 2026
d997820
Fix backward compatibility issue with MFSDP `--grad-reduce-in-bf16` (…
shjwudp Mar 12, 2026
251a754
feat: add NCCL flight recorder configuration support (#3806)
sbak5 Mar 12, 2026
5a3aa17
Revert "Ensure dummy_forward does not attempt to run cudagraphs (#378…
ko3n1g Mar 12, 2026
5b25326
Fix if statement in main (#3833)
tdene Mar 12, 2026
6657173
Update golden values of weekly tests (#3829)
ko3n1g Mar 12, 2026
4736aed
build: Loosen TE restriction (#3827)
ko3n1g Mar 12, 2026
1d5e68b
Upgrade GitHub Actions for Node 24 compatibility (#3830)
ko3n1g Mar 12, 2026
46227e0
Do not let chunked prefill generate decode logprobs (#3777)
tdene Mar 12, 2026
e08dc9d
Prevent double serialization inside Flask server (#3653)
tdene Mar 12, 2026
29e798a
Allow RL to run inference-only via skip-train (#3744)
tdene Mar 12, 2026
4c1d0e4
Announce Python 3.12 migration (#3825)
ko3n1g Mar 12, 2026
fabbcdf
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Mar 13, 2026
b250472
ci: Skip test_wrong_cuda_graph_impl_returns_false in LTS (#3847)
chtruong814 Mar 13, 2026
7ca9dc5
ci: Mark TestCoordinator.test_throughput as flaky (#3849)
chtruong814 Mar 13, 2026
f261906
find optimal number of workers (#3699)
dimapihtar Mar 13, 2026
b7437fe
remove encoder_and_decoder (#3836)
dimapihtar Mar 13, 2026
8a806e5
ci: Skip more tests in test_vision_cuda_graphs for LTS (#3860)
chtruong814 Mar 13, 2026
e70916d
Merge main into dev.
ilml Mar 13, 2026
0c8dd94
Merge branch 'dev' into merge-main-into-dev
ilml Mar 19, 2026
8096db6
style: black-format test_absorbed_mla for CI lint
ilml Mar 19, 2026
c2cbe54
fix: stabilize CI dependency setup after merge
ilml Mar 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
6 changes: 4 additions & 2 deletions .github/actions/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ runs:
run: echo "node_name=$NODE_NAME" | tee -a "$GITHUB_OUTPUT"

- name: Checkout repository
uses: actions/checkout@v2
uses: actions/checkout@v6

- name: Change ownership of /home/runner/
shell: bash
Expand Down Expand Up @@ -98,7 +98,8 @@ runs:
--environment dev \
--platform dgx_h100 \
--tag ${{ inputs.tag }} \
--container-image ${{ inputs.container-image }}
--container-image ${{ inputs.container-image }} \
--hf-home /mnt/datadrive/TestData/nemo-fw/TestData/HF_HOME

RUN_TEST_EOF
)
Expand Down Expand Up @@ -186,6 +187,7 @@ runs:
--platform dgx_h100 \
--container-image ${{ inputs.container-image }} \
--data-dir /mnt/datadrive/TestData/megatron-lm/artifacts \
--hf-home /mnt/datadrive/TestData/nemo-fw/TestData/HF_HOME

RUN_TEST_EOF
)
Expand Down
2 changes: 1 addition & 1 deletion .github/copy-pr-bot.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
enabled: true
auto_sync_draft: false
auto_sync_ready: true
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "Victarry", "Wohox", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "asolergi-nv", "buptzyb", "chtruong814", "cspades", "cuichenx", "deepakn94", "dimapihtar", "dingqingy-nv", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "faradawn", "frsun-nvda", "gautham-kollu", "gdengk", "guyueh1", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kanz-nv", "kevalmorabia97", "ko3n1g", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "mkhona-nvidia", "parthmannan", "prajwal1210", "pthombre", "rogerwaleffe", "sajadn", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "trintamaki", "tylerpoon", "wdykas", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yueshen2016", "yuzhongw-nvidia", "zhongbozhu"]
trustees_override: ["AAnoosheh", "ArEsKay3", "Autumn1998", "BestJuly", "BoxiangW", "CarlosGomes98", "ChenhanYu", "FDecaYed", "HaochenYuan", "ISEEKYAN", "JRD971000", "Phlip79", "QiZhangNV", "RPrenger", "ShriyaRishab", "Victarry", "Wohox", "ZhiyuLi-Nvidia", "ahmadki", "aklife97", "ananthsub", "asolergi-nv", "buptzyb", "chtruong814", "cjld", "cspades", "cuichenx", "deepakn94", "dimapihtar", "dingqingy-nv", "duncanriach", "erhoo82", "ericharper", "fanshiqing", "faradawn", "frsun-nvda", "gautham-kollu", "gdengk", "guyueh1", "huvunvidia", "hxbai", "ilml", "jalbericiola", "janEbert", "jaredcasper", "jenchen13", "jiemingz", "jingqiny-99", "jkamalu", "jon-barker", "jstjohn", "kanz-nv", "kevalmorabia97", "ko3n1g", "ksivaman", "kunlunl", "kvareddy", "kwyss-nvidia", "layalir", "lhb8125", "lmcafee-nvidia", "maanug-nv", "mathemakitten", "matthieule", "mchrzanowski", "mehraakash", "mkhona-nvidia", "nanz-nv", "parthmannan", "prajwal1210", "pthombre", "rhewett-nv", "rogerwaleffe", "sajadn", "sanandaraj5597", "sancha", "santhnm2", "sbak5", "shanmugamr1992", "sharathts", "shengf-nv", "shifangx", "shjwudp", "sidsingh-nvidia", "skyw", "sudhakarsingh27", "tdene", "theothermike", "thomasdhc", "tomlifu", "trintamaki", "tylerpoon", "wdykas", "wplf", "xiaoyao0115", "xuwchen", "yanring", "yaox12", "yaoyu-33", "yashaswikarnati", "yeyu-nvidia", "yobibyte", "youngeunkwon0405", "yueshen2016", "yuzhongw-nvidia", "zhongbozhu"]
38 changes: 19 additions & 19 deletions .github/oncall_schedule.json
Original file line number Diff line number Diff line change
@@ -1,16 +1,4 @@
[
{
"user": "janEbert",
"date": "2026-02-18"
},
{
"user": "asolergi-nv",
"date": "2026-02-25"
},
{
"user": "BoxiangW",
"date": "2026-03-04"
},
{
"user": "maanug-nv",
"date": "2026-03-11"
Expand All @@ -20,31 +8,43 @@
"date": "2026-03-18"
},
{
"user": "gautham-kollu",
"user": "janEbert",
"date": "2026-03-25"
},
{
"user": "janEbert",
"user": "gautham-kollu",
"date": "2026-04-01"
},
{
"user": "maanug-nv",
"user": "ilml",
"date": "2026-04-08"
},
{
"user": "BoxiangW",
"user": "Phlip79",
"date": "2026-04-15"
},
{
"user": "Phlip79",
"user": "asolergi-nv",
"date": "2026-04-22"
},
{
"user": "asolergi-nv",
"user": "BoxiangW",
"date": "2026-04-29"
},
{
"user": "dimapihtar",
"user": "maanug-nv",
"date": "2026-05-06"
},
{
"user": "dimapihtar",
"date": "2026-05-13"
},
{
"user": "gautham-kollu",
"date": "2026-05-20"
},
{
"user": "ilml",
"date": "2026-05-27"
}
]
46 changes: 15 additions & 31 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,19 +5,8 @@

## Contribution process

```mermaid
flowchart LR
A[Pre-checks] --> B[PR Tests]
subgraph Code Review/Approval
C1[Expert Review] --> C2[Final Review]
end
B --> C1
C2 --> D[Merge]
```

### Pre-checks

- [ ] I want this PR in a versioned release and have added the appropriate Milestone (e.g., `Core 0.8`)
- [ ] I have added relevant unit tests
- [ ] I have added relevant functional tests
- [ ] I have added proper typing to my code [Typing guidelines](https://docs.python.org/3/library/typing.html)
Expand All @@ -26,41 +15,36 @@ flowchart LR

### Code review

The following process is enforced via the CODEOWNERS file for changes into `megatron/core`. For changes outside of `megatron/core`, it is up to the PR author whether or not to tag the Final Reviewer team.
Feel free to message or comment the [@mcore-oncall](https://github.com/orgs/NVIDIA/teams/mcore-oncall) to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

<details>
<summary>For MRs into `main` branch</summary>
All PRs start as **draft**. If you open a non-draft PR, it will be automatically converted to draft.

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!
#### Step 1: Mark PR as "Ready for Review"

#### (Step 1): Add PR label `Expert Review`
1. When your PR is ready, click **Ready for Review**.
2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
- Some PRs may jump straight to step 2. This is determined by `.github/CODEOWNERS`.

#### (Step 2): Collect the expert reviewers reviews
:warning: Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

1. Attach the `Expert Review` label when your PR is ready for review.
2. GitHub auto-assigns expert reviewers based on your changes. They will get notified and pick up your PR soon.
#### Step 2: Final Review

:warning: Only proceed to the next step once all reviewers have approved, merge-conflict are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.
For PRs that change `megatron/core`, once all expert reviewers have approved, the `Final Review` label is applied **automatically** and final reviewers are assigned.

#### (Step 3): Final Review
For PRs outside `megatron/core`, this step is skipped.

1. Add `Final Review` label
2. GitHub auto-assigns final reviewers based on your changes. They will get notified and pick up your PR soon.
#### Step 3: Approved

#### (Optional Step 4): Cherry-pick into release branch
Once all required reviewers have approved, the `Approved` label is applied **automatically**.

If this PR also needs to be merged into `core_r*` release branches, after this PR has been merged, select `Cherry-pick` to open a new PR into the release branch.
### Merge

</details>
Any member of [mcore-engineers](https://github.com/orgs/NVIDIA/teams/mcore-engineers) will be able to merge your PR.

<details>
<summary>For MRs into `dev` branch</summary>
The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either `eharper@nvidia.com` or `zijiey@nvidia.com`.
</details>

### Merging your PR

Any member of [core-adlr](https://github.com/orgs/teams/NVIDIA/core-adlr) and [`core-nemo`](https://github.com/orgs/teams/NVIDIA/core-nemo) will be able to merge your PR.
6 changes: 3 additions & 3 deletions .github/workflows/_build_test_publish_wheel.yml
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ jobs:
PUBLISH_DRYRUN: ${{ inputs.dry-run }}
steps:
- name: Checkout repository
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
ref: ${{ inputs.ref }}

Expand Down Expand Up @@ -136,7 +136,7 @@ jobs:
test "${{ steps.build-wheel.outputs.expected-release-number }}" == "$RELEASE_NUMBER"

- name: Upload wheels
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@v6
with:
name: wheels-${{ matrix.PACKAGE }}-${{ matrix.PLATFORM }}-${{ inputs.dry-run && 'dry-run' || 'release' }}
path: dist/
Expand All @@ -159,7 +159,7 @@ jobs:
PACKAGE: ${{ matrix.PACKAGE }}
steps:
- name: Download wheels
uses: actions/download-artifact@v4
uses: actions/download-artifact@v7
with:
name: wheels-${{ matrix.PACKAGE }}-${{ matrix.PLATFORM }}-${{ inputs.dry-run && 'dry-run' || 'release' }}
path: dist/
Expand Down
41 changes: 32 additions & 9 deletions .github/workflows/_release_library.yml
Original file line number Diff line number Diff line change
Expand Up @@ -53,13 +53,34 @@ on:
description: Starting tag for changelog builder (leave empty for auto-detect)
type: string
default: ""
publish-docs:
required: false
description: Publish documentation to S3 after release
type: boolean
default: true
secrets:
TWINE_PASSWORD:
required: true
SLACK_WEBHOOK:
required: true
PAT:
required: true
AWS_ASSUME_ROLE_ARN:
required: true
AWS_ACCESS_KEY_ID:
required: true
AWS_SECRET_ACCESS_KEY:
required: true
AKAMAI_HOST:
required: true
AKAMAI_CLIENT_TOKEN:
required: true
AKAMAI_CLIENT_SECRET:
required: true
AKAMAI_ACCESS_TOKEN:
required: true
S3_BUCKET_NAME:
required: true

permissions:
contents: write # To read repository content
Expand Down Expand Up @@ -89,7 +110,7 @@ jobs:
IS_DRY_RUN: ${{ inputs.dry-run }}
steps:
- name: Checkout repository
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
path: ${{ github.run_id }}
token: ${{ secrets.PAT }}
Expand Down Expand Up @@ -199,11 +220,8 @@ jobs:
# Extract PR number from URL
PR_NUMBER=$(echo $PR_URL | grep -o '[0-9]*$')

# Add comment to the newly created PR
echo gh pr comment $PR_NUMBER --body "/ok to test $(git rev-parse HEAD)"

- name: Wait for status checks on tmp branch
uses: actions/github-script@v7
uses: actions/github-script@v8
id: wait-status
with:
github-token: ${{ secrets.PAT }}
Expand Down Expand Up @@ -326,7 +344,6 @@ jobs:
ref: ${{ inputs.release-ref }}
no-publish: false
secrets:
TWINE_USERNAME: ${{ secrets.TWINE_USERNAME }}
TWINE_PASSWORD: ${{ secrets.TWINE_PASSWORD }}

create-gh-release:
Expand All @@ -344,10 +361,10 @@ jobs:
REPOSITORY: ${{ github.repository }}
PROJECT_NAME: Megatron Core
VERSION: ${{ needs.bump-next-version.outputs.release-version }}
TAG_PREFIX: ${{ inputs.gh-release-tag-prefix || '' }}
TAG_PREFIX: core_
steps:
- name: Checkout repository
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
path: ${{ github.run_id }}
ref: ${{ inputs.release-ref }}
Expand Down Expand Up @@ -455,6 +472,12 @@ jobs:
publish-docs:
needs: [bump-next-version, create-gh-release]
uses: ./.github/workflows/release-docs.yml
if: |
(
success() || !failure()
)
&& inputs.publish-docs == true
&& !cancelled()
with:
dry-run: ${{ inputs.dry-run }}
publish-as-latest: true
Expand All @@ -472,7 +495,7 @@ jobs:
VERSION: ${{ needs.build-test-publish-wheels.outputs.version }}
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
repository: NVIDIA-NeMo/FW-CI-templates
ref: v0.17.0
Expand Down
10 changes: 5 additions & 5 deletions .github/workflows/_update_dependencies.yml
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ jobs:
TARGET_BRANCH: ${{ inputs.target-branch }}
steps:
- name: Checkout repo
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
ref: ${{ env.TARGET_BRANCH }}

Expand All @@ -60,7 +60,7 @@ jobs:
fi

- name: Checkout repo
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
ref: ${{ env.SOURCE_BRANCH }}

Expand All @@ -77,7 +77,7 @@ jobs:
bash -c 'uv lock --upgrade'

- name: Upload lock file
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@v6
with:
name: lock-file-${{ env.SOURCE_BRANCH }}
path: uv.lock
Expand All @@ -90,7 +90,7 @@ jobs:
TARGET_BRANCH: ${{ inputs.target-branch }}
steps:
- name: Checkout code
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
token: ${{ secrets.PAT }}
ref: ${{ env.TARGET_BRANCH }}
Expand All @@ -103,7 +103,7 @@ jobs:
fi

- name: Download lock file
uses: actions/download-artifact@v4
uses: actions/download-artifact@v7
with:
name: lock-file-${{ env.SOURCE_BRANCH }}

Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/auto-reminder-bot.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,10 +14,10 @@ jobs:
if: github.repository == 'NVIDIA/Megatron-LM'
steps:
- name: Check out repository code
uses: actions/checkout@v4
uses: actions/checkout@v6

- name: Set up Python
uses: actions/setup-python@v5
uses: actions/setup-python@v6
with:
python-version: "3.10"

Expand Down
Loading
Loading