Skip to content

[NVIDIA] DeepSeek V4 Perf Tracking #33636

Description

@b8zhong

Motivation

This follows DeepSeek V4 Roadmap, which covers functional enablement; this issue is perf-only.

Scope: NVIDIA SM90 / SM10X.


High priority

Attention & compression kernels

Indexer & top-k

mHC

  • FlashInfer mHC fusion feat: Add flashinfer mHC fusion for DSV4 #33616 — TileLang fusion is still faster today; this is the path when TileLang is unavailable (for future architectures), and FlashInfer should gain full pre / post+pre fusion soon

MoE & quantization

Speculative decoding (MTP / DSpark / EAGLE3)

Communication

CUDA graph & scheduling

Memory & KV capacity

Context Parallel

Docs & recipes

CI & bug tracking


If an open PR belongs here and I missed it, comment and I'll add it.

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions