Skip to content

Docs/kepler vs modeling计算图ir对比分析 - #16

Closed
laksjdf wants to merge 25 commits into
jiashaokun-1:mainfrom
laksjdf:docs/kepler-vs-modeling计算图IR对比分析

Hidden character warning

The head ref may contain hidden characters: "docs/kepler-vs-modeling\u8ba1\u7b97\u56feIR\u5bf9\u6bd4\u5206\u6790"
Closed

Docs/kepler vs modeling计算图ir对比分析#16
laksjdf wants to merge 25 commits into
jiashaokun-1:mainfrom
laksjdf:docs/kepler-vs-modeling计算图IR对比分析

Conversation

@laksjdf

@laksjdf laksjdf commented May 26, 2026

Copy link
Copy Markdown
Contributor

Docs/kepler vs modeling计算图ir对比分析

laksjdf and others added 25 commits May 21, 2026 17:45
pp_p2p统计口径问题(掩盖、bwd)
寻优脚本优化,在format_result的时候增加缓存
支持pp开低精度、开overlap
支持pp开低精度、开overlap
Fix DualPipeV VPP scheduling accounting
1.fix dsv4 PARAM_SCOPE miss 2.add xlsx test
寻优允许pp=1情况,增加hidden限制
…ry 的问题;expert_agg 现在保持 MoE activation dtype,只在真正的 residual2 边界插入 FP8->BF16 cast。

修复 YAML/model config 中 attn_weight_dtype, shared_expert_weight_dtype, attn_grad_dtype, shared_expert_grad_dtype 未传入 ModelSpec 的问题。
修复 memory_breakdown() 中 shared expert weight/grad dtype 被并入全局 dtype 统计的问题;routed/shared expert 现在分别按各自 dtype 计入权重和梯度显存。
…ry 的问题;expert_agg 现在保持 MoE activation dtype,只在真正的 residual2 边界插入 FP8->BF16 cast。

修复 YAML/model config 中 attn_weight_dtype, shared_expert_weight_dtype, attn_grad_dtype, shared_expert_grad_dtype 未传入 ModelSpec 的问题。
修复 memory_breakdown() 中 shared expert weight/grad dtype 被并入全局 dtype 统计的问题;routed/shared expert 现在分别按各自 dtype 计入权重和梯度显存。
@laksjdf laksjdf closed this May 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants