Commit Graph
12919 Commits
Author SHA1 Message Date
Xinyuan TongandXinyuan Tong 52eebc82ae [Docs] MiMo-V2.5 cookbook: B200 benchmarks + multi-layer EAGLE acceptance profile + long-context reference (#25359)
Co-authored-by: Xinyuan Tong <xinyuan.tong@radixark.ai>
2026-05-19 23:15:21 -07:00
Michael 7fda3caea4 [AMD] test(sgl-kernel): seed RNG on ROCm in test_moe_topk_sigmoid to fix tie-break flake (#25356) 2026-05-19 23:01:13 -07:00
Xiaobaoandgithub-actions[bot] 0d3a94f643 [Bug] Correct Weight Offloader's Attribute Name for torch.nn.Parameter (#25786)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-19 22:28:59 -07:00
Yuan Luoandluoyuan.luo b7085d3860 [fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end) (#25532)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-05-20 13:20:37 +08:00
Cheng WanandClaude Sonnet 4.6 052abcc0dd [Refactor] Pass PP start_layer via model constructor instead of forward_batch.token_to_kv_pool (#25825)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-19 22:16:07 -07:00
Cheng Wan a4b51d35ef Revert "[codex] Update Wan2.2 ModelOpt CI checkpoints" (#25845) 2026-05-19 21:45:20 -07:00
Xinyuan Tong 0aedc5678b loader: yield filtered MTP weights lazily to avoid OOM hang on multi-layer EAGLE (#25748) 2026-05-20 12:33:54 +08:00
Xiaoyu Zhang af22390af7 [codex] Align diffusion skills with nightly Nvidia benchmarks (#25842) 2026-05-20 12:18:05 +08:00
liuxianglong17andAdarsh Shirawalmath 579fed2090 Reduce excessively long logs caused by transformer version updates. (#25737)
Co-authored-by: Adarsh Shirawalmath <114558126+adarshxs@users.noreply.github.com>
2026-05-20 11:39:03 +08:00
Shangming Cai 1fbee74fb6 [PD] Clean early abort logic in PD module (#25677)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-05-20 11:17:42 +08:00
Kevin Flansburg 3b2178c412 [PD] Un-blacklist mooncake sessions when probe succeeds (#25287) 2026-05-20 11:16:28 +08:00
mingyue300 5fe655bf26 vlm: fix shared memory bug of deep copy mm_items (#24231) 2026-05-20 11:06:51 +08:00
ybyang a8c82c652e fix(dsv4): make pool configurator PP-aware (#25750) 2026-05-20 10:58:53 +08:00
ybyang ca29c2b0e7 fix(dsv4): drop stale pp_size=1 guard for V4 PD disaggregation (#25771) 2026-05-20 10:57:23 +08:00
jy-song-hub d69e9fbfdc [diffusion] fix: honor config precisions for delight/paint (#22289) 2026-05-20 10:38:48 +08:00
jy-song-hubandjxp 549ae16c6f [diffusion] fix: respect configured precision in Qwen layered path (#21980)
Co-authored-by: jxp <jingxin.pan123@gmail.com>
2026-05-20 10:38:03 +08:00
jy-song-hub 3a9d9d5832 [diffusion] fix: fix Hunyuan3D-2 DiT checkpoint param mapping (#22729) 2026-05-20 10:37:13 +08:00
gaopengff 65fe32379e [Intel GPU]Support fused_topk for XPU (#24641) 2026-05-20 10:34:29 +08:00
Baizhou Zhang 0c8049d9ba Update CI permissions and CODEOWNERS (#25826) 2026-05-19 18:19:21 -07:00
Xiaoyu Zhang 80fc524809 [diffusion] quant: update Wan2.2 modelOpt CI checkpoints (#25483) 2026-05-20 09:05:39 +08:00
Liangsheng Yin 7f154ba449 drop output ids (#25774) 2026-05-19 17:50:47 -07:00
Paiiiiandzengpai 425dffbde3 DeepSeek V4 MTP Support CP (#24934)
Co-authored-by: zengpai <zengpai@baidu.com>
2026-05-19 16:51:31 -07:00
YAMY beaff00331 [NSA] Avoid repeated NSA MQA logits memory queries (#25299) 2026-05-19 16:04:13 -07:00
Liangsheng Yin 3ef832f885 pr-states: dispatch from pr-test* notify job (fix rerun status) (#25812) 2026-05-19 15:57:16 -07:00
b9d470f4a2 Support spec v2 for FlashMLA speculative decoding (#24640)
Co-authored-by: Jackey Hua <zhendonghua@users.noreply.github.com>
Co-authored-by: Depend <yu-depend@users.noreply.github.com>
2026-05-19 15:23:17 -07:00
Liangsheng Yin 2f70902329 deflake priority below-threshold test (#25809) 2026-05-19 15:19:01 -07:00
b9c2bf717b [BugFix] Resolve adaptive speculative decoding conflicts for Qwen3.5 (hybrid GDN) (#23331)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
2026-05-19 15:09:49 -07:00
Liangsheng Yin 16bcc4583e verify_done: wait not synchronize (#25465) 2026-05-19 14:57:45 -07:00
Ratish P fab097d66d [Gemma4]: Fix FP8 Triton scale layout (#25286) 2026-05-19 14:00:23 -07:00
ybyang 8322fe09a7 fix(dsv4): upgrade forward metadata on main stream for large PP size (#25729) 2026-05-19 20:52:00 +00:00
Baizhou Zhang cd012ada58 [Fix] Fix extra uninstall of cutlass packages (#25756) 2026-05-19 10:01:38 -07:00
Yuan Luoandluoyuan.luo 4c0ce0345d Support Gemma4 Pipeline Parallelism (#25284)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-05-19 22:40:11 +08:00
Makcum888e 3b62604cec [Diffusion] Support parallelism for GLM-Image (#25645) 2026-05-19 17:27:21 +03:00
Xinyuan Tong aad00b0ed8 Upgrade transformers to 5.8.1 (#25451) 2026-05-19 22:20:30 +08:00
Yuxuan Zhang 2bcb6d2f82 [Bug Fix] Align glm4_moe_nextn NPU MTP loading with qwen3 MTP (#25524) 2026-05-19 14:47:01 +01:00
amote-i de3fc46e3d [NPU] [DOC] remove Qwen3-235B-A22B 2K+2K 100ms mixed mode benchmark (#25778) 2026-05-19 20:48:43 +08:00
Liangsheng Yin e0273dcd31 pr-test-extra: re-trigger on labeled event (#25732) 2026-05-19 05:15:55 -07:00
Xiaoyu Zhang 0e4d1b49d3 [Codex] Remove stale DeepSeek V4 JIT kernels (#25764) 2026-05-19 20:04:32 +08:00
Arseniy MironovandNapkin-AI 45a85efc3a [Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
2026-05-19 12:46:55 +03:00
Thomas 58b5fe3e29 [Diffusion] [NPU] Fix HunyuanVideo crash on NPU (#25592) 2026-05-19 12:40:43 +03:00
jianzhao-xu 5073c82a37 transformers v5 adapt HFRunner (#23922) 2026-05-19 17:07:38 +08:00
shiyu7 7e0818038a fix: fix deepseek v4 CP error (#25396) 2026-05-19 02:04:21 -07:00
67fd005b97 [HiSparse & PD] Support hisparse memory pool host page > 1 (#23606)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-05-19 01:29:35 -07:00
amote-i 1f7bf155c3 [NPU] [DOCS] Improved the usability of Ascend NPU documents (#25735) 2026-05-19 16:22:22 +08:00
Ziang Li 78cb38ed5e [FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918) 2026-05-19 01:04:48 -07:00
Kevin Li fbfddfd5c7 fix (jit kernel): elementwise activation C++ error (#25695) 2026-05-19 15:23:52 +08:00
Yuhao Yang 79ea30d1f1 [Bug] Fix V4-Pro NaN on Blackwell by converting fp8_einsum input scale to ue8m0 (#25733) 2026-05-18 23:48:34 -07:00
YC Yen-Ching Tseng 7c3f614e23 [AMD] Bump amd/Kimi-K2.5-MXFP4 revision to align with shared-experts fusion (#25740) 2026-05-18 22:47:05 -07:00
Alex Nails 1d19721394 [gRPC] Native server: Rust crate (1/N) (#23506) 2026-05-18 22:31:38 -07:00
Junlin Wuandronnie_zheng 4c9f31b85e [diffusion][npu][quant] Add MXFP4 quantization support for Wan2.2 Diffusion on Ascend NPU (#22338)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-19 07:46:52 +03:00