Commit Graph
8639 Commits
Author SHA1 Message Date
Liwansi 55ba03db6a [NPU]use triton split_qkvgate_gemma_rmsnorm_rope for Qwen3.5 and Qwen3_next (#23925) 2026-05-20 20:22:10 +08:00
Liangsheng Yin 34d3e23232 spec_v2: consolidate seq_lens_cpu/sum maintenance into helper (#25818) 2026-05-20 04:42:26 -07:00
Liangsheng Yin 9b005d3608 disagg prebuilt: drop dead prepare_for_extend shift (#25819) 2026-05-20 04:39:47 -07:00
Liangsheng Yin 1bd4f94598 [Test] Add fwd_occupancy sanity kit (#25886) 2026-05-20 03:34:37 -07:00
Chi McIsaac 47979fb252 [diffusion] fix: fix GLM-Image /v1/images/edits support (#25697) 2026-05-20 17:11:51 +08:00
Liangsheng Yin 614672fea5 [Test] Stage-a sanity kits; consolidate core/ + models_e2e/ tests (#25831) 2026-05-20 01:58:48 -07:00
Yuhong Guo 24d27c2035 [BugFix] Fix rid_to_state leak for aborted queued requests (#24070) 2026-05-20 01:32:44 -07:00
Matt Van HornandMatt Van Horn e99f87c974 fix: add missing distro dependency to runtime docker image (#25817)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
2026-05-20 00:46:32 -07:00
044649c23a feat: Support flashinfer_cutedsl MoE runner with flashinfer alltoall backend (#22669)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 00:36:26 -07:00
Cheng WanandClaude Sonnet 4.6 8131641bc6 [Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-20 00:18:04 -07:00
Xiaobaoandgithub-actions[bot] 0d3a94f643 [Bug] Correct Weight Offloader's Attribute Name for torch.nn.Parameter (#25786)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-19 22:28:59 -07:00
Cheng WanandClaude Sonnet 4.6 052abcc0dd [Refactor] Pass PP start_layer via model constructor instead of forward_batch.token_to_kv_pool (#25825)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-19 22:16:07 -07:00
Cheng Wan a4b51d35ef Revert "[codex] Update Wan2.2 ModelOpt CI checkpoints" (#25845) 2026-05-19 21:45:20 -07:00
Xinyuan Tong 0aedc5678b loader: yield filtered MTP weights lazily to avoid OOM hang on multi-layer EAGLE (#25748) 2026-05-20 12:33:54 +08:00
Xiaoyu Zhang af22390af7 [codex] Align diffusion skills with nightly Nvidia benchmarks (#25842) 2026-05-20 12:18:05 +08:00
Shangming Cai 1fbee74fb6 [PD] Clean early abort logic in PD module (#25677)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-05-20 11:17:42 +08:00
Kevin Flansburg 3b2178c412 [PD] Un-blacklist mooncake sessions when probe succeeds (#25287) 2026-05-20 11:16:28 +08:00
mingyue300 5fe655bf26 vlm: fix shared memory bug of deep copy mm_items (#24231) 2026-05-20 11:06:51 +08:00
ybyang a8c82c652e fix(dsv4): make pool configurator PP-aware (#25750) 2026-05-20 10:58:53 +08:00
ybyang ca29c2b0e7 fix(dsv4): drop stale pp_size=1 guard for V4 PD disaggregation (#25771) 2026-05-20 10:57:23 +08:00
jy-song-hub d69e9fbfdc [diffusion] fix: honor config precisions for delight/paint (#22289) 2026-05-20 10:38:48 +08:00
jy-song-hubandjxp 549ae16c6f [diffusion] fix: respect configured precision in Qwen layered path (#21980)
Co-authored-by: jxp <jingxin.pan123@gmail.com>
2026-05-20 10:38:03 +08:00
jy-song-hub 3a9d9d5832 [diffusion] fix: fix Hunyuan3D-2 DiT checkpoint param mapping (#22729) 2026-05-20 10:37:13 +08:00
gaopengff 65fe32379e [Intel GPU]Support fused_topk for XPU (#24641) 2026-05-20 10:34:29 +08:00
Xiaoyu Zhang 80fc524809 [diffusion] quant: update Wan2.2 modelOpt CI checkpoints (#25483) 2026-05-20 09:05:39 +08:00
Liangsheng Yin 7f154ba449 drop output ids (#25774) 2026-05-19 17:50:47 -07:00
Paiiiiandzengpai 425dffbde3 DeepSeek V4 MTP Support CP (#24934)
Co-authored-by: zengpai <zengpai@baidu.com>
2026-05-19 16:51:31 -07:00
YAMY beaff00331 [NSA] Avoid repeated NSA MQA logits memory queries (#25299) 2026-05-19 16:04:13 -07:00
b9d470f4a2 Support spec v2 for FlashMLA speculative decoding (#24640)
Co-authored-by: Jackey Hua <zhendonghua@users.noreply.github.com>
Co-authored-by: Depend <yu-depend@users.noreply.github.com>
2026-05-19 15:23:17 -07:00
b9c2bf717b [BugFix] Resolve adaptive speculative decoding conflicts for Qwen3.5 (hybrid GDN) (#23331)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
2026-05-19 15:09:49 -07:00
Liangsheng Yin 16bcc4583e verify_done: wait not synchronize (#25465) 2026-05-19 14:57:45 -07:00
Ratish P fab097d66d [Gemma4]: Fix FP8 Triton scale layout (#25286) 2026-05-19 14:00:23 -07:00
ybyang 8322fe09a7 fix(dsv4): upgrade forward metadata on main stream for large PP size (#25729) 2026-05-19 20:52:00 +00:00
Yuan Luoandluoyuan.luo 4c0ce0345d Support Gemma4 Pipeline Parallelism (#25284)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-05-19 22:40:11 +08:00
Makcum888e 3b62604cec [Diffusion] Support parallelism for GLM-Image (#25645) 2026-05-19 17:27:21 +03:00
Xinyuan Tong aad00b0ed8 Upgrade transformers to 5.8.1 (#25451) 2026-05-19 22:20:30 +08:00
Yuxuan Zhang 2bcb6d2f82 [Bug Fix] Align glm4_moe_nextn NPU MTP loading with qwen3 MTP (#25524) 2026-05-19 14:47:01 +01:00
Xiaoyu Zhang 0e4d1b49d3 [Codex] Remove stale DeepSeek V4 JIT kernels (#25764) 2026-05-19 20:04:32 +08:00
Arseniy MironovandNapkin-AI 45a85efc3a [Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
2026-05-19 12:46:55 +03:00
Thomas 58b5fe3e29 [Diffusion] [NPU] Fix HunyuanVideo crash on NPU (#25592) 2026-05-19 12:40:43 +03:00
jianzhao-xu 5073c82a37 transformers v5 adapt HFRunner (#23922) 2026-05-19 17:07:38 +08:00
shiyu7 7e0818038a fix: fix deepseek v4 CP error (#25396) 2026-05-19 02:04:21 -07:00
67fd005b97 [HiSparse & PD] Support hisparse memory pool host page > 1 (#23606)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-05-19 01:29:35 -07:00
Ziang Li 78cb38ed5e [FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918) 2026-05-19 01:04:48 -07:00
Kevin Li fbfddfd5c7 fix (jit kernel): elementwise activation C++ error (#25695) 2026-05-19 15:23:52 +08:00
Yuhao Yang 79ea30d1f1 [Bug] Fix V4-Pro NaN on Blackwell by converting fp8_einsum input scale to ue8m0 (#25733) 2026-05-18 23:48:34 -07:00
Junlin Wuandronnie_zheng 4c9f31b85e [diffusion][npu][quant] Add MXFP4 quantization support for Wan2.2 Diffusion on Ascend NPU (#22338)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-19 07:46:52 +03:00
Hanming Lu 862d39e06c [Mamba] Fix extra_buffer overlap schedule races (#24954) 2026-05-19 12:13:20 +08:00
Zhonghua Deng f0763859ed perf(mimo-v2-epd): enable GPU image preprocess and parallel video decode (#25588) 2026-05-19 11:47:21 +08:00
Yuhao Yang d8e66e54e5 fix: use triton_attn as default vision attention on B300 (SM103) (#25570) 2026-05-19 11:00:07 +08:00