Commit Graph
16884 Commits
Author SHA1 Message Date
Liangsheng Yin a49560ce50 [misc] Add a comment style rule to .claude/rules (#35597) 2026-08-19 18:52:48 -07:00
datdo-msft f65961844a docker: fix CUDA-13 build — rename NCCL_VERSION ARG to avoid base image ENV collision (#35587) 2026-08-19 18:41:53 -07:00
dc175b3ad2 [CI][AMD] Run the profiling suite without CUDA graphs on ROCm (#34452)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
2026-08-20 09:41:28 +08:00
chuyehandChen c7478228dd [AMD] [Docker] Upgrade Python 3.12 + torch 2.11 + triton 3.7 in ROCm 7.2.4 (#30984)
Co-authored-by: Chen <bingxche@amd.com>
2026-08-19 18:18:31 -07:00
Mick ab203663c4 [diffusion] fix: reject unsupported modelopt checkpoint algorithms (#35182) 2026-08-20 09:13:24 +08:00
Jimmy ShongandLING ZHI 1cf2b8c54d [Spec] Support quantized target lm_head in the DFlash2 selector (#35496)
Co-authored-by: LING ZHI <1747985437lz@gmail.com>
2026-08-19 18:06:41 -07:00
Mick 1f87d8f512 [diffusion] fix: stop reserving nccl device buffers for single-rank groups (#35538) 2026-08-20 09:04:35 +08:00
MichaelandCursor Agent e805a8f98e [AMD] Keep the PTX-inline-asm diffusion norm fusions off on ROCm (fix FLUX warmup crash) (#34481)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-08-20 08:55:28 +08:00
Lianmin Zheng f736895ce9 Make PR babysitter launcher fork-safe (#35575) 2026-08-19 17:48:34 -07:00
Lianmin ZhengandYichao Fu 99c12218c3 Support custom draft worker classes in DSpark (#35397)
Co-authored-by: Yichao Fu <yichaofu@meta.com>
2026-08-19 17:48:17 -07:00
Lianmin ZhengandXiaozhu Meng 9234e40aed [sampling] Fix int32 offset overflow in top-k renorm Triton kernels (#35571)
Co-authored-by: Xiaozhu Meng <mxz297@gmail.com>
2026-08-19 17:48:03 -07:00
Shiyan Deng d216737e47 [Kernel] Support wider rows in mega_moe_pre_dispatch (#35372) 2026-08-19 17:47:04 -07:00
elvischenvandClaude Opus 4.8 1df78c2cf1 chore: bump tilelang to 0.1.12 (#30874)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-19 17:01:19 -07:00
Jimmy Shong 5375babbac [Quant] Load compressed-tensors quantized lm_head instead of value-casting it (#35228) 2026-08-19 15:37:45 -07:00
c7e2c08d14 fix(constrained): reject NUL bytes in grammar specs to stop an xgrammar segfault (#34679)
Signed-off-by: Junhao Shen <junshen@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-08-19 15:31:14 -07:00
milesial 082aac8fce [Bugfix] Fix min-new-token EOS handling (#31378)
Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com>
2026-08-19 15:29:22 -07:00
Zhiqiang Xie 01814e110d [HiCache] Simple style change for buffer mode (#35574) 2026-08-19 15:23:21 -07:00
Siyuan Chen 38b74d294b Add docs for TP LMHead optimizaiton (#35283) 2026-08-19 14:59:35 -07:00
Liangsheng Yin 1270204d2c Revert "[Feature] Add DeepEPv2 (ElasticBuffer) MoE A2A backend" (#35568) 2026-08-19 14:14:38 -07:00
elvischenvandBrayden Zhong a6bc0532c9 [Fix] Fix Nemotron-H Mamba illegal memory access under DP attention with CUDA graph (#34561)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-08-19 14:13:40 -07:00
milesial ed12d6827d fix(disagg): allow fake transfer with decode DCP (#35409)
Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com>
2026-08-19 13:54:37 -07:00
Jason Wiemels defb2a3100 feat(openai): Accept the input_audio content part in chat completions (#33606) 2026-08-19 13:37:50 -07:00
DarkSharpnessandClaude Opus 5 746418a1ec [DSA] Trim top-k v2 output modes and tighten its PDL waits (#35041)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 13:09:21 -07:00
cctry 1c82955861 [HiCache] Split the host-memory budget across co-located ranks (#35540) 2026-08-19 13:01:58 -07:00
Ayushman Singh b1707996e8 fix(gemma4): quantize MTP bridge projections (#32440) 2026-08-19 12:41:52 -07:00
Po-Han Huang (NVIDIA) 6f69f927da [Scheduler] Add configurable decode interval after prefill (#35017) 2026-08-19 12:01:36 -07:00
MengYuandmenyu 4f8ecf6ae9 [Feature] Add DeepEPv2 (ElasticBuffer) MoE A2A backend (#29525)
Co-authored-by: menyu <menyu@nvidia.com>
2026-08-19 11:52:45 -07:00
YAMY 03cf2de2e3 [Qwen3.5][MTP] Preserve online NVFP4 draft quantization for mixed checkpoints (#35545) 2026-08-19 11:12:32 -07:00
Xinyuan Tong 157d8ad27a Support Intern-S2-Mobius FP8 (#34908) 2026-08-19 10:58:01 -07:00
YAMY 5f12839591 [Fix] Support Kimi-K3 ModelOpt mixed NVFP4/FP8 checkpoint (#35077) 2026-08-19 08:13:45 -07:00
Shuwen Wangandhzh0425 41c018a9ec [UnifiedTree] feat: support runtime attach/detach (#35269)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-08-19 22:49:08 +08:00
Yoray Zack 1ef7882a5b [NIXL] Query EP top-k index dtype (#35294) 2026-08-19 22:31:05 +08:00
MickandClaude Opus 5 23f2320c95 [Docs] PaddleOCR-VL: update which stage of the pipeline this serves and show real output (#35458)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:27:58 +08:00
Mick e3445ed2bd [diffusion] fix: route quantized vae component repos safely (#35184) 2026-08-19 21:02:46 +08:00
Mick 29f5d1c7c3 [diffusion] fix: fix multi-group layerwise offload startup memory (#35509) 2026-08-19 20:32:31 +08:00
Arseniy Mironovandronnie_zheng c57ada81e1 [Diffusion] Use current_platform instead of hardcoded "cuda" in cosmos3 guardrails (#34612)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-19 15:26:28 +03:00
jacky.cheng 574274660f [AMD] cookbook: serve Qwen3.5 MXFP4 on MI355X with an fp8_e4m3 KV cache (#35445) 2026-08-19 18:49:32 +08:00
Rohit Kumar SinghandSingh 3e5ce26c2d fix: fix transcription & audio-understanding for ASR/audio/speech models (#32611)
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
2026-08-19 18:44:23 +08:00
f446e853e7 [AMD] DeepSeek-V4: route decode wo_a bf16 batched matmul to aiter batched_gemm_bf16 (#33313)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-08-19 03:04:48 -07:00
karverma-amd ce1830c59b [AMD] DeepSeek-V4 MI355X: eliminate bpreshuffle fp8-scale relayout copy in dense w8a8 linear (#33165) 2026-08-19 03:02:40 -07:00
HeYaoandHeYao f22442d3a4 Add three new test cases (#35502)
Co-authored-by: HeYao <heyao@example.com>
2026-08-19 17:59:45 +08:00
Shangming Cai adca19c497 [PD] Deferred decode-side KV release for the NIXL backend (#35360) 2026-08-19 17:29:45 +08:00
YAMY aa215e5523 [PD] Overlap prefill DP-rank bootstrap queries (#35071) 2026-08-19 17:25:52 +08:00
Khoa PhamandCursor 0e4a09480c [HiCache] Support DCP with DSpark (#35221)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-19 01:42:19 -07:00
Xiaoyu ZhangandClaude Opus 5 9113fc6d93 [docs] Add a fused-kernels page for SGLang Diffusion (#35436)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 16:32:42 +08:00
8a1e6e4e46 Qwen3.8-27B Model Support (#34859)
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-08-19 16:31:43 +08:00
kangwangamdandbingxche ebec85f606 [AMD][DI][CI] Run MI355X disagg nightly at 7AM UTC (#35467)
Co-authored-by: bingxche <bingxche@users.noreply.github.com>
2026-08-19 16:27:20 +08:00
MickandYiqi Yang 73e5fa4724 [diffusion] refactor: gate native encoder quantized checkpoints (#35183)
Co-authored-by: Yiqi Yang <yangyiqi8787@gmail.com>
2026-08-19 16:18:36 +08:00
Liangsheng Yin ccbe380028 [CI] Trim the base-c 4-gpu-h100 stage from 5 shards to 4 (#35407) 2026-08-19 00:48:07 -07:00
Liangsheng Yin e614121866 [Fix] Scale the req_to_token row headroom by attn_dcp_size (#35424) 2026-08-19 00:37:26 -07:00