100 Commits
Author SHA1 Message Date
YAMY 9b59fc5db5 [ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load (#40628) 2026-09-21 21:45:58 -07:00
YAMY 0229025127 [Spec][PP] Launch extend microbatches before the spec output exchange (#40499) 2026-09-21 15:47:03 -07:00
YAMY 9cc7da2ab0 [MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts) (#38080) 2026-09-19 16:00:38 -07:00
YAMY 24874f90a3 [Test] Fix DeepGEMM batch invariance test output dtype (#39636) 2026-09-15 12:40:28 -07:00
YAMY 55e5e21c88 [Qwen3.8-Next] Add PD state transfer for Flash Next (#36651) 2026-09-11 22:31:45 -07:00
YAMY f963d7a27c fix(qsa): make the paged sparse-decode gather memory-safe (zero-fill scratch, int64 offsets, dequant FP8 on gather) (#38851) 2026-09-11 15:47:25 -07:00
YAMY 55b45cb45a fix(qsa): dequantize FP8 cached prefixes in the sparse prefill kernels (#38855) 2026-09-10 14:07:39 -07:00
7b791c9534 [Bugfix] Keep a shared MAX_LEN prefill CUDA graph bucket when the graph captures a DP gather (MegaMoE sparse-DP hang) (#37933)
Co-authored-by: shyeh25 <206795756+shyeh25@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-09 10:52:16 -07:00
YAMY 1ad3eb09a9 [Attention] Size FlashInfer MLA indptr buffers to the padded max batch (#38590) 2026-09-09 01:25:29 -07:00
YAMY 65f7957142 fix: gather CP-sharded tokens before TP-sharded dense MLP under prefill CP (#38078) 2026-09-04 17:14:31 -07:00
YAMY db89f639ef [GDN] Amortize ReplaySSM checkpoint materialization (#35544) 2026-09-04 15:13:20 -07:00
YAMY 07199fa220 [Performance] Optimize Qwen3.5 GDN prefill projection layouts (#36267) 2026-09-04 10:32:44 -07:00
YAMY 982aa8acfc [Bugfix] Load Qwen3.5 MTP embedding under PP (#37471) 2026-09-02 15:19:40 -07:00
YAMY 3c9cea8f10 [EAGLE] Prune draft-extend logits to selected rows (#35546) 2026-09-02 15:10:08 -07:00
YAMY fe45af1e6f perf(gdn): select ReplaySSM verify loop unrolling by shape (#36970) 2026-09-02 15:07:49 -07:00
YAMY a6a19f9290 [Bugfix] Skip absent radix lock during cache cleanup (#37494) 2026-09-01 22:50:19 -07:00
YAMY 221a6273ce [Kernel] Avoid zero-bias allocation in fused softmax routing (#36811) 2026-09-01 16:29:57 -07:00
YAMY 33ed29a0ee test: update hybrid attention runner fixtures (#37345) 2026-08-31 22:00:50 -07:00
YAMY a53718e88c fix(ci): update attention backend test fixtures (#37219) 2026-08-31 02:15:17 -07:00
YAMYandgithub-actions[bot] b77cac06a9 [PP] Support prefill CUDA graph proxy tensors (#36248)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-30 23:11:45 -07:00
YAMY 2ea6d17eab fix(staging): make empty staging rings reusable (#37166) 2026-08-30 20:59:57 -07:00
YAMY 9a03bc2dc3 [CI] Fix stale GPU capability test patches (#37148) 2026-08-30 10:12:32 -07:00
YAMY 84e56982b6 [Fix] Fix transformer loader fallback test fixture (#37142) 2026-08-30 08:28:11 -07:00
YAMY 137fab2f48 fix(deps): pin compressed-tensors to 0.18.0 (#36794) 2026-08-27 22:52:47 -07:00
YAMY 6480cce9bc perf: skip redundant scheduler metadata gather for DP1 (#36568) 2026-08-27 20:04:54 -07:00
YAMY e9c9df6a52 [Performance] Tune FlashInfer EXTEND for DP prefill (#36219) 2026-08-25 08:29:57 -07:00
YAMY 834400705f perf: overlap Qwen shared expert with DeepEP routed experts (#34938) 2026-08-21 15:39:44 -07:00
YAMY ae23423b46 Split TRTLLM MHA decode batches by KV sequence length (#34888) 2026-08-20 00:44:26 -07:00
YAMY 03cf2de2e3 [Qwen3.5][MTP] Preserve online NVFP4 draft quantization for mixed checkpoints (#35545) 2026-08-19 11:12:32 -07:00
YAMY 5f12839591 [Fix] Support Kimi-K3 ModelOpt mixed NVFP4/FP8 checkpoint (#35077) 2026-08-19 08:13:45 -07:00
YAMY aa215e5523 [PD] Overlap prefill DP-rank bootstrap queries (#35071) 2026-08-19 17:25:52 +08:00
YAMY 7c423cfd41 [PD] Avoid unused PREBUILT prompt tensor transfer (#35070) 2026-08-17 16:48:12 +08:00
YAMY e6250c7c70 docs: update Qwen3.8 disaggregated serving configs (#34601) 2026-08-12 11:02:35 -07:00
YAMYandShangming Cai 667e18d99d [PD] Support pipeline-parallel prefill with Mooncake staging buffer (#33807)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-08-11 13:53:17 +08:00
YAMYandShangming Cai c971d7ac9c Refactor staging registration metadata fields (#33910)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-08-10 17:53:32 +08:00
YAMY 5a8e360e70 [PD] Skip speculative verify scratch on prefill servers (saves num_draft_tokens x mamba pool per rank) (#34191) 2026-08-09 20:56:00 -07:00
a59bb931c6 Fix DCP KV head mapping for GQA models (#32858)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 14:25:09 -07:00
YAMY 9e3f6b746b fix(mamba): widen causal_conv1d token offsets to int64 (#33665) 2026-08-07 14:31:48 -07:00
YAMY 2fc557254b fix(PP): size the mamba pool per pipeline stage, not per whole model (#33666) 2026-08-06 13:10:43 -07:00
YAMYandShangming Cai 05c7ebf64c [Disagg][StagingBuffer][2/2] Support radix cache (#30545)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-08-06 23:59:35 +08:00
YAMYandLee Nau 8b29c90218 [NVIDIA] Enable CuTe DSL BF16 GEMM on SM107 (#33617)
Co-authored-by: Lee Nau <lnau@nvidia.com>
2026-08-06 02:06:08 -07:00
YAMY 990a446773 Fix padded positions in breakable CUDA Graph attention (#33253)
TestBreakableCUDAGraphBasic and all NVIDIA CI tests pass.
2026-08-05 15:53:28 -07:00
YAMYandChunan Zeng 7bc90ab394 Select DeepGEMM standard layouts by memory budget (#33474)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
2026-08-05 15:50:43 -07:00
YAMYandSam Li 5fe97637df Support DeepGEMM for standard MoE dispatch (#33128)
Co-authored-by: Sam Li <lsam@nvidia.com>
2026-08-02 21:48:13 -07:00
YAMY 48dbc24cbf [Qwen3.5][MTP] Support FlashInfer CuTe DSL for online NVFP4 draft MoE (#31382) 2026-07-30 17:19:48 -07:00
YAMY fddfc1fb5e [GDN] Support FlashInfer GDN prefill with extra-buffer radix cache (#29735) 2026-07-30 00:47:35 +08:00
YAMYandLee Nau 86ee545388 docs(cookbook): update Kimi-K3 GB200 recipes from measured 4x4 runs (#32592)
Co-authored-by: Lee Nau <lnau@nvidia.com>
2026-07-28 16:26:40 -07:00
YAMY dd67452b4f [Cleanup] Move mamba-max-states-per-path validation into _handle_mamba_backend (#32502) 2026-07-28 14:21:21 -07:00
YAMY 91f386a5b2 fix(disagg): support pipeline-parallel hybrid-linear transfer (#32270) 2026-07-25 13:34:38 -07:00
YAMY 2428f56145 [Bugfix] Fix Kimi-Linear state transfer across heterogeneous TP (#32262) 2026-07-24 10:31:17 -07:00
YAMY de816e1eb5 [Disagg][StagingBuffer][1/2] Robustness and failure handling (#31217) 2026-07-24 17:22:55 +08:00
YAMY c18919f8f3 [Mamba] Add a per-path cap for cached states (#31230) 2026-07-23 17:58:36 +08:00
YAMY 7c9257529f [Mamba] Support speculative decoding with extra_buffer_lazy (#30437) 2026-07-21 19:38:03 +08:00
YAMYandXuwei Li 2d00e20a52 [Disagg][Qwen3.5] Fix heterogeneous attn-TP scatter transfer: GDN conv sub-block slice + GQA replicated-KV head map (#30997)
Co-authored-by: Xuwei Li <lixuwei.xy@gmail.com>
2026-07-16 02:31:37 +08:00
YAMY 41ad0d9c26 Allow prefill breakable CUDA graph for Qwen3.5 via multimodal opt-in allowlist (#30620) 2026-07-13 18:39:44 -07:00
YAMY 2e4d6368c3 [GDN] Auto-select FlashInfer GDN prefill on validated SM100 configs (#29734) 2026-07-10 00:26:00 +08:00
YAMY 8a868f8c00 [NVIDIA] Allow modelopt_mixed quantization with flashinfer_cutedsl MoE runner (#30443) 2026-07-07 23:50:52 -07:00
YAMY 48ad6a83cf [DeepSeek-V4] Enable non-paged indexer by default for large prefill chunks (#30140) 2026-07-07 15:51:24 -07:00
YAMY e81f05cf4f [dsv4] Trigger MHC prenorm prewarm at weight-load time with rank sync (#29988) 2026-07-02 23:05:03 -07:00
YAMY a6ee64d237 [DeepSeek-V4] Add an opt-in non-paged indexer for long-context prefill (#29619) 2026-07-02 22:54:40 -07:00
YAMYandYuwei An 307094dc7d [DeepSeek V4] Cover both dense and sparse prefill paths in the compress attention unittest (#29885)
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
2026-07-01 21:20:12 -07:00
YAMYandBaizhou Zhang c865347b98 [DeepSeek V4] Enable FlashMLA sparse prefill by default (#29775)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-07-01 13:50:05 -07:00
YAMY 93015a9e6b fix(runner): autotune flashinfer MoE on a decode-shaped buffer (#29069) 2026-06-23 13:31:47 -07:00
YAMY b3be2e7402 [dsv4] Pad MLA decode q-heads to 64 (not full n_heads) for FlashMLA head64 kernel (#27954) 2026-06-15 17:18:10 -07:00
YAMY f870bf1ed0 [dsv4] Prewarm MHC prenorm kernel at startup (#27986) 2026-06-15 13:26:26 -07:00
YAMY d5899b95c4 fix(qwen3.5): keep CUDA dual-stream overlap (regressed by #25885) (#27868) 2026-06-15 09:44:21 -04:00
YAMY f23f48df98 fix(moe): make FlashInfer A2A robust to collapsed global_num_tokens (moe_dense_tp_size NaN) (#27945) 2026-06-12 16:11:44 -07:00
YAMY f8b0a120b8 fix: DSV4 BCG compress-prefill plan OOB on underfilled (tiny) prefill replay (#27747) 2026-06-10 20:02:24 -07:00
YAMYandYuwei An ca66e6fb5e [BCG] Support breakable CUDA graph for DeepSeek V4 DP attention (#25195)
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
2026-06-08 13:54:58 -07:00
5af02c18ae [spec_v2] Enable trtllm_mha draft-extend CUDA graph with v2 semantics (#25002)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 17:50:12 -07:00
YAMY eae03ce3b2 refactor(dsv4): route MHC prenorm through DeepGEMM wrapper (#26238) 2026-05-27 17:45:45 -07:00
YAMY 3a6de13cd8 perf(dsv4): add MHC token-count prewarm (#25810) 2026-05-20 22:22:41 -07:00
YAMY 801d7e3eed [DSA] Make MQA logits free memory ratio configurable (#25859) 2026-05-20 12:27:16 -07:00
YAMY beaff00331 [NSA] Avoid repeated NSA MQA logits memory queries (#25299) 2026-05-19 16:04:13 -07:00
YAMY 1f9eda4ea1 Use DeepGEMM BF16 for unquantized DeepEP LL MoE (#25540) 2026-05-17 23:33:23 -07:00
YAMY 8d5b347edd Support Qwen3.5 NVFP4 MTP DeepEP (#24906) 2026-05-15 10:49:14 +08:00
YAMY 4be25f2428 fix(disagg): broadcast bootstrap port across multi-node prefill ranks (#24378) 2026-05-14 16:39:01 +08:00
2a4d382b07 [Disagg][NIXL] Add staging buffer support for heterogeneous TP KV transfer (#22536)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-05-13 19:54:02 +08:00
YAMY a4109e87ac Fix TRTLLM MHA routing for draft extend (#24856) 2026-05-12 15:48:32 -07:00
YAMY b202778aa5 Fix reduce_scatterv producer contract for SUM_LEN (#24785) 2026-05-10 16:51:13 -07:00
YAMY 560829a171 feat(scheduler): add adaptive queue-based prefill delayer trigger (#23189) 2026-05-08 16:54:30 -07:00
YAMY 6971a03fe6 fix(fa3): skip scheduler_metadata precompute under DP attention (#24632) 2026-05-08 16:19:20 -07:00
YAMY 657945c338 Replace all-reduce + dp_scatter with reduce_scatterv for DP attention (#22642) 2026-04-13 21:51:10 -07:00
YAMY c26b8b4a4b [GDN] Remove FlashInfer GDN decode + no_buffer guard and default to FlashInfer on SM100+ (#21861) 2026-04-08 11:59:15 -07:00
YAMY 5ae00ecd48 [Disagg][NIXL] Support Mamba state slice transfer for heterogeneous TP (Step 2/2 for Qwen3.5) (#22240) 2026-04-07 23:47:31 +08:00
YAMY 3148742ddb [Disagg][NIXL] Fix heterogeneous TP KV transfer for non-MLA models (same logic with mooncake, Step 1/2 for Qwen3.5 support) (#22145) 2026-04-07 14:52:02 +08:00
YAMYandShangming Cai dc125afffb Add staging buffer CI test and documentation for heterogeneous TP (#21921)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-04-06 14:00:20 +08:00
YAMY 821a8a99fb [Disagg] GPU staging buffer with dynamic ring allocator for heterogeneous TP KV transfer (#19890) 2026-04-01 14:09:18 -07:00
YAMYandShangming Cai cfead25bbf [Qwen3.5] mamba slice fix (Prefill TP != Decode TP & decode TP size>1) (#20655)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-03-17 19:30:58 +08:00
YAMY 08957c88ea [Logging] Fix prefill side logging in pd disagg (#19350) 2026-02-25 12:42:18 -08:00
YAMYandShangming Cai f75abb4521 [Fix][Qwen3.5] Fix KV cache slice transfer for GQA models with replicated KV heads (#19086)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-02-25 16:26:44 +08:00
YAMY 73fe389dd1 [Qwen3.5] Raise Exception when radix_cache and extra_buffer are enabled at the same time (#19169) 2026-02-25 15:04:37 +08:00
YAMY 5995bfec63 [Qwen3-Next] Enable fused_qkvzba_split_reshape_cat also for prefill (#18917) 2026-02-22 13:57:17 +08:00
YAMY cef353f338 [Fix] Quick fix for int32 overflow in Mooncakes' send_kvcache_slice (#19076) 2026-02-22 12:00:33 +08:00
YAMY 8541b1118d [Fix][Qwen3.5] Pass max_mamba_cache_size to mamba pool in disaggregation decode path (#19002) 2026-02-20 14:31:26 +08:00
YAMY 454676811e [Flashinfer Autotune] Fix FlashInfer FP4 MoE autotuning crash by removing incorrect flatten on hidden_states_scale (#18500) 2026-02-12 13:31:27 +08:00
YAMY a45e0e5df4 [SPEC_V2] Enable cudagraph draft_extend for trtllm_mla_backend and Acclen Fix for DP under cudagraph mode (#16974) 2026-01-18 15:56:21 +08:00
YAMY daa4841e86 [ConfigArgumentMerger] Improve ConfigArgumentMerger compatibility with external callers (#17051) 2026-01-16 15:32:42 +08:00
YAMY 2740ed1ae7 [eval] GSM8k support for run_eval (#17041) 2026-01-16 11:10:17 +08:00
YAMY 3fd88ea9b5 [MTP][spec_v2] Fix TRTLLM MLA backend crash in EAGLE draft_extend mode (#15790) 2026-01-10 19:58:23 +08:00