Commit Graph
16957 Commits
Author SHA1 Message Date
a688682f4b [AMD][CI] Fix ROCm 7.0's dead apt index fail the MORI dependency install (#35764)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
2026-08-20 22:13:09 -07:00
Ke Bao 73a2c117c6 Support mxfp8 KV cache in PD transfer (#35718) 2026-08-21 13:05:06 +08:00
Michaelandquitenode f64080fbaf [AMD] CI: cut two setup cycles from the AMD multimodal-gen lanes (#34483)
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
2026-08-20 22:01:51 -07:00
Jimmy Shong 3efa057449 [docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c (#35786) 2026-08-20 21:54:51 -07:00
Bingxu Chen 78c964d9d7 [AMD] Retry transient network failures in ROCm Dockerfile curl fetches (#35654) 2026-08-20 21:54:24 -07:00
Thomas Wang 34180a0d35 [AMD] Improve K3 dspark draft attn kernel perf (#35499) 2026-08-20 21:48:59 -07:00
bda9952377 [AMD] DeepSeek-V4 MI355X: eliminate bpreshuffle fp8-scale copies at producer sites (MoE down, MLA o_proj bmm) (#33166)
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-08-20 21:46:23 -07:00
Mick 6127d1daee [diffusion] feat: allow offloaded weights stay on the checkpoint mapping (#35701) 2026-08-21 10:54:28 +08:00
Zhangheng 44806dc507 Using unified radix tree by default for all case (#35081) 2026-08-21 10:45:46 +08:00
MickandClaude Opus 5 e0cf75d9bd [doc] standardize diffusion cookbook model pages (#34247)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 10:25:40 +08:00
Xiaoyu Zhang 7e80e889a2 [diffusion] Fuse LTX-2.5 decoder 3D RoPE (#35698) 2026-08-21 10:13:09 +08:00
ishandhananiandShangming Cai 978244d671 [P/D disagg] Decode-side radix cache for SWA hybrid models (unified radix tree) (#27770)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-08-21 09:25:06 +08:00
Mohammad Miadh Angkad aa3f766799 [CI] Temporarily disable B300 jobs (#35627) 2026-08-20 17:54:53 -07:00
Mick 5bb981dce1 [diffusion] chore: read the cgroup this process is actually in (#35707) 2026-08-21 08:48:50 +08:00
Nan Jiang f825d72936 [Sampling] Restore finite top-k requirement for sampling masks (#35205) 2026-08-20 16:55:25 -07:00
Jimmy Shong 14795dcb1a [docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev image tag (#35767) 2026-08-20 23:43:55 +00:00
ethcheandEthan Che 67f6ad61d9 fix(kernel) Fix Helion small-token prefill bug (#35197)
Co-authored-by: Ethan Che <eche@meta.com>
2026-08-20 16:41:59 -07:00
Yihao Wang ba8e601358 [docs] fix note formatting in sglang-d documentation (#35761) 2026-08-20 16:39:39 -07:00
Jimmy Shong 1a138e13b9 [docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753) 2026-08-20 23:34:44 +00:00
Shenxiu LiuandQiaolin Yu 779e593bd1 Fix _GenerationStreamAccumulator logprob_end off-by-one under retract (#26510)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-08-20 16:03:13 -07:00
Wenkai DuandHubert Lu 5a7b26c636 [AMD] [sgl-kernel] Bypass caches for peer traffic in ROCm custom all-reduce (#32832)
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
2026-08-20 15:24:18 -07:00
Jiajun Li a4ef828207 fix(openai): avoid duplicate routed expert in response when return_meta_info = True (#35323) 2026-08-20 15:21:58 -07:00
Richard Gong 94907f05c4 Add CI permissions for four contributors (#35600) 2026-08-20 15:18:18 -07:00
Liangsheng Yin 0149f56e84 [CI] Gate /rerun-test on commenter trust and remove /rerun-stage (#35750) 2026-08-20 15:05:52 -07:00
Baizhou Zhang 92eeed41d7 [Docker] Defer CUDA 13 NCCL override until after dependency resolution (#35756) 2026-08-20 14:48:07 -07:00
ishandhananiandAlex Nails 0f744b6848 feat: make mm_inputs msgpack-native (#29656)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-20 14:30:07 -07:00
Liangsheng Yin 5a100d9086 [misc] Trim restating comments and docstrings in srt/managers (#35622) 2026-08-20 14:18:40 -07:00
Lee Nau ad367d72b0 [Kimi K3] Select FlashInfer MXFP4 for SM107 auto MoE (#35554) 2026-08-20 14:09:31 -07:00
Jimmy Shong d9f6861359 [docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663) 2026-08-20 13:26:55 -07:00
eac91ac362 [Fix] Land the decode mamba checkpoint depth on the tree page under DCP (#35412)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-08-20 12:15:42 -07:00
Junlin Wuandronnie_zheng 308bc1228b 📝 [NPU] Clean up quantization comments (#34829)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-20 22:03:58 +03:00
Connor Carpenter 61fa64ae7e feat(grpc): expose KV event discovery metadata (#35714)
Signed-off-by: Connor Carpenter <connorc@nvidia.com>
2026-08-20 13:39:24 -05:00
Chao Shi 2ef0fe4669 TP/PP Consensus checker (#34406) 2026-08-21 01:36:03 +08:00
Shangming Cai 23cb04093c fix(multimodal): keep LLaVA image fetch off the CPU-preprocess timeout budget (flaky test_mixed_batch) (#35700) 2026-08-21 01:26:30 +08:00
Ke Bao ba97cc6397 Skip empty linear-attention state buffers in PD transfer (#35689) 2026-08-21 01:00:50 +08:00
R0CKSTAR 81df6f2c57 [MUSA] Harden CI dependencies and diffusion warmup (#35610) 2026-08-20 09:40:24 -07:00
Mick be373395b4 [diffusion] feat: support out-of-tree models and pipelines (#35713) 2026-08-21 00:33:34 +08:00
Mick 7f8f030000 [diffusion] feat: let every layerwise component be configurable (#35688) 2026-08-20 22:38:05 +08:00
Xiaoyu Zhang 04444ee352 [diffusion] Refresh eager optimization skills and benchmark safeguards (#35679) 2026-08-20 22:03:57 +08:00
Shuwen Wang 9b249a25a1 test: switch the Inkling-Small NVFP4 deterministic suite to DSPARK (#35293) 2026-08-20 21:54:52 +08:00
silencejade b03ac355e7 [NPU] [FIX] Fix non-contiguous parameter issue in FIA operator (#34936) 2026-08-20 20:43:40 +08:00
Estrella-xx c98f1ccedb [NPU]Ensure tensors allocated by empty_like are contiguous (#34935) 2026-08-20 20:40:57 +08:00
Mohammad Miadh Angkad a4ffb996db [Fix] Keep deterministic GDN prefill on Triton (#35632) 2026-08-20 19:38:53 +08:00
Mick 82c6fc2db9 [diffusion] quant: support pruned safetensors checkpoints for minimax-h3 (#35418) 2026-08-20 19:34:14 +08:00
Mick 97efc0507c [diffusion] feat: plan pinned host memory against the cgroup cap not the machine (#35641) 2026-08-20 19:32:29 +08:00
Jimmy Shong 710267dc4c [Quant] Load compressed-tensors kv_cache_scheme scales (#35455) 2026-08-20 19:17:59 +08:00
Mick cf3813f4ce [diffusion] feat: add weight source reader (#35668) 2026-08-20 18:37:29 +08:00
Mick 17313cf4b2 [diffusion] CI: add minimax-h3 ref2va audio consistency coverage and guard peak vram (#35511) 2026-08-20 17:22:05 +08:00
Mick f1b9a1f42a [diffusion] feat: support unverified short edge instead of rejecting it for minimax-h3 (#35664) 2026-08-20 16:52:44 +08:00
Bingxu ChenandCursor 06ad7b2b0d [AMD][CI] Run Both ROCm 7.2.4 and ROCm 7.2.0 Images on Nightly Test AMD (#35603)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-20 01:37:23 -07:00