Commit Graph
14338 Commits
Author SHA1 Message Date
Thomas Wang cee1caaf47 [AMD] Fix nightly-8-gpu-mi35x-deepseek-v4-flash-rocm720 OOM issue (#28941) 2026-06-22 07:35:45 -07:00
Thomas Wang 04d952ea10 [AMD] deepseek-v4 clean env vars (#28920) 2026-06-22 07:32:43 -07:00
YC Yen-Ching Tseng dd39ef6cec [AMD] Pin httpx>=0.25.0 to fix anthropic SDK socket_options error (#28869) 2026-06-22 07:25:22 -07:00
Mick 4923bb93ae [diffusion] CI: fix turbo_wan/flux invisible CI cases (#28913) 2026-06-22 21:21:16 +08:00
Lianmin Zheng ad9723af03 Clean up CUDA graph capture logs (#28937) 2026-06-22 06:15:26 -07:00
Mick ead39d38fc [diffusion] refactor: refactor causal KV local head cache updates (#28888) 2026-06-22 21:00:38 +08:00
shihaozhou 1adb53f147 Fix CP page filtering by request-local position (#28718) 2026-06-22 20:47:29 +08:00
Shangming Cai 06e001347a [CI] Update nixl installation to include nixl-cu13 for h20 runner (#28930) 2026-06-22 20:35:14 +08:00
Shangming Cai b8e64fb56d [CI] Modify nixl installation to force reinstall (#28927) 2026-06-22 19:26:26 +08:00
Xiaoyu Zhang 0c9e775f2c [Diffusion] Fix FastWan2.1 default 480p resolution (#28733) 2026-06-22 18:43:24 +08:00
jianzhao-xu 4e1d25117b [NPU] update best practice docs from testcase (#28621) 2026-06-22 16:58:48 +08:00
Yuwei An 2ce32366a0 [Fix][BCG][Spec] Restore EAGLE prefill plumbing dropped by #23906 (#28870) 2026-06-22 01:54:54 -07:00
amote-i 93553a67a3 [NPU] [DOC] Create deployment tutorials for mainstream models on Ascend NPU (#27893) 2026-06-22 16:21:13 +08:00
Xinyuan Tong db12bfcdc8 [JIT] Add kpool_topk_transform JIT kernel (#28670) 2026-06-22 01:04:21 -07:00
Liangsheng Yinandthanhhao98 106d2930a6 [core] Gate the overlap WAR barrier on forward reads to recover decode throughput (#28363)
Co-authored-by: thanhhao98 <31717833+thanhhao98@users.noreply.github.com>
2026-06-22 00:40:56 -07:00
Xinyuan TongandXinyuan Tong 441ae9a5ae [Lint] Fix black formatting of DeepSeek-R1-MXFP4 MI35x tests (#28885)
Co-authored-by: Xinyuan Tong <justintong0323@gmail.com>
2026-06-22 14:20:34 +08:00
Polisetty V R K Jyothendra VarmaandMa Mingfei 62b3c8e177 [Intel GPU] Guard tvm_ffi import in dsv4 online mtp module under TYPE_CHECKING to fix import error on XPU (#28531)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-22 13:56:21 +08:00
Liangsheng Yin 64e455d4bf Fix lint break on main (#28886) 2026-06-21 22:46:08 -07:00
Bingxu Chen e2540188ce [AMD] Clean up DeepSeek-R1-MXFP4 TP2/TP4 MLA GSM8K tests (#27243) 2026-06-21 21:41:19 -07:00
Khoa PhamandCursor 0642cd5020 (chore): bump tokenspeed_mla to 0.1.7 (#28759)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-21 21:25:37 -07:00
018d0c21dc [Docs] Add Anthropic-compatible API documentation (#28522)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 04:01:09 +00:00
be774d0acd [docs][cookbook] Laguna-M.1 playground: add HiCache; refresh EP / DP-Attention notes (#28774)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-22 03:58:42 +00:00
Lianmin Zheng 886b96621d Migrate more server args to annotated style (#28830) 2026-06-21 20:50:17 -07:00
cctryandcctry 0c065671c9 [Spec] Redo: split init_backends; account draft weights in --mem-fraction-static (#28855)
Co-authored-by: cctry <cctry@fb.com>
2026-06-21 20:45:16 -07:00
Bingxu Chen fd7874d11b [AMD] Register DP attention test (#28495) 2026-06-21 20:21:41 -07:00
YC Yen-Ching Tseng 73448b0d70 [AMD] Temporarily disable deepseek V4 in AMD PR test (#28871) 2026-06-22 10:48:38 +08:00
Mick 2b2cd21783 [diffusion] fix: reject cache-dit with fsdp (#28834) 2026-06-22 10:47:40 +08:00
Trevor Morris c0bb04b67f [NVIDIA] Support NVFP4 MoE for DeepSeek-V4 (#25820) 2026-06-21 19:35:14 -07:00
amote-i 5deca2d39f [DOC] [NPU] Update features on Ascend NPU (#28643) 2026-06-22 09:50:51 +08:00
6779ca8d7f Fix Qwen MoE precision issue with PP and all-reduce fusion (#28619)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-06-22 08:20:16 +08:00
Liangsheng Yin 4f5ff39bc9 [Spec] Enable FR-Spec in EAGLE draft-extend CUDA graph by sizing logits buffer from the draft head (#28856) 2026-06-21 15:35:57 -07:00
Liangsheng Yin e6722c751b [Feature] Add graceful scheduler shutdown; free hisparse host buffer on exit (#28779) 2026-06-21 15:08:10 -07:00
Liangsheng Yin 8e890391f5 [Spec] Support FlashInfer CUDA graph for EAGLE draft-extend (#28782) 2026-06-21 14:46:25 -07:00
Lianmin Zhengandhnyls2002 a4d0ff3def [misc] Make NaN-logit sanitization opt-in (default off) (#28829)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-06-21 14:31:35 -07:00
7f67965b4d [BugFix] NCCL deadlock in HiCache writing_check by making all_reduce unconditional (#26923)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-06-22 01:47:44 +08:00
iridiumine 5351800700 [Bugfix] Fix MTP acceptance regression on plan stream by moving int64 cast before plan stream context (#28410) 2026-06-22 01:26:58 +08:00
Mohammad Miadh Angkad 643ee748c6 [PP] Pass DSA topk through PP warmup proxy buffers (#28785) 2026-06-21 23:55:39 +08:00
Lianmin Zheng 7942d546d1 Revert "[Spec] Split init_backends; account draft weights in --mem-fraction-static" (#28841) 2026-06-21 07:52:35 -07:00
Mick 320b231ea6 [diffusion] chore: bound DiffGenerator local cleanup (#28833) 2026-06-21 22:23:49 +08:00
Mick a7f31a6e1b [diffusion] fix: fix SANA-WM CFG-parallel tensor devices (#28835) 2026-06-21 22:01:00 +08:00
Mick a51d56d948 CI: Pin flash-attn-4 for diffusion CI consistency (#28838) 2026-06-21 21:13:41 +08:00
Lianmin Zheng 3975ea5ac7 Fix H20 torch import reinstall fallback (#28818) 2026-06-21 04:50:46 -07:00
cctry 9691a29fe0 [Spec] Split init_backends; account draft weights in --mem-fraction-static (#28683) 2026-06-21 01:22:26 -07:00
cctry 6d4ca9bc54 Cap SWA pool sizing with chunk cache (#28755) 2026-06-21 01:06:59 -07:00
Lianmin Zheng c9488241e9 [Refactor] Auto-derive CLI args from dataclass fields to eliminate duplication (#28814) 2026-06-21 00:51:08 -07:00
Jairo David Campaña RoseroandXinyuan Tong b4dda8b3ce fix(anthropic): handle mid-conversation system messages (#26773)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-21 04:34:47 +00:00
Mick c65f4ea692 [diffusion] fix: validate openai sampling dimensions (#28791) 2026-06-21 09:44:38 +08:00
Mick 6a16573a7f [diffusion] fix: fix Qwen-Image-Layered string image paths (#28790) 2026-06-21 09:43:56 +08:00
Lianmin Zheng d331fdd2ba Add project rule: prefer msgspec.Struct over dataclasses (#28816) 2026-06-20 18:43:16 -07:00
Lianmin Zheng 54b9b9d0c9 Remove threading atexit monkey patch (#28812) 2026-06-20 18:30:50 -07:00