 Matt Van HornandMatt Van Horn
|
e99f87c974
|
fix: add missing distro dependency to runtime docker image (#25817)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
|
2026-05-20 00:46:32 -07:00 |
|
 
|
044649c23a
|
feat: Support flashinfer_cutedsl MoE runner with flashinfer alltoall backend (#22669)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-20 00:36:26 -07:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
8131641bc6
|
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-20 00:18:04 -07:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Xiaobaoandgithub-actions[bot]
|
0d3a94f643
|
[Bug] Correct Weight Offloader's Attribute Name for torch.nn.Parameter (#25786)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-05-19 22:28:59 -07:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
052abcc0dd
|
[Refactor] Pass PP start_layer via model constructor instead of forward_batch.token_to_kv_pool (#25825)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-19 22:16:07 -07:00 |
|
Cheng Wan
|
a4b51d35ef
|
Revert "[codex] Update Wan2.2 ModelOpt CI checkpoints" (#25845)
|
2026-05-19 21:45:20 -07:00 |
|
Xinyuan Tong
|
0aedc5678b
|
loader: yield filtered MTP weights lazily to avoid OOM hang on multi-layer EAGLE (#25748)
|
2026-05-20 12:33:54 +08:00 |
|
Xiaoyu Zhang
|
af22390af7
|
[codex] Align diffusion skills with nightly Nvidia benchmarks (#25842)
|
2026-05-20 12:18:05 +08:00 |
|
Shangming Cai
|
1fbee74fb6
|
[PD] Clean early abort logic in PD module (#25677)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-20 11:17:42 +08:00 |
|
Kevin Flansburg
|
3b2178c412
|
[PD] Un-blacklist mooncake sessions when probe succeeds (#25287)
|
2026-05-20 11:16:28 +08:00 |
|
mingyue300
|
5fe655bf26
|
vlm: fix shared memory bug of deep copy mm_items (#24231)
|
2026-05-20 11:06:51 +08:00 |
|
ybyang
|
a8c82c652e
|
fix(dsv4): make pool configurator PP-aware (#25750)
|
2026-05-20 10:58:53 +08:00 |
|
ybyang
|
ca29c2b0e7
|
fix(dsv4): drop stale pp_size=1 guard for V4 PD disaggregation (#25771)
|
2026-05-20 10:57:23 +08:00 |
|
jy-song-hub
|
d69e9fbfdc
|
[diffusion] fix: honor config precisions for delight/paint (#22289)
|
2026-05-20 10:38:48 +08:00 |
|
 jy-song-hubandjxp
|
549ae16c6f
|
[diffusion] fix: respect configured precision in Qwen layered path (#21980)
Co-authored-by: jxp <jingxin.pan123@gmail.com>
|
2026-05-20 10:38:03 +08:00 |
|
jy-song-hub
|
3a9d9d5832
|
[diffusion] fix: fix Hunyuan3D-2 DiT checkpoint param mapping (#22729)
|
2026-05-20 10:37:13 +08:00 |
|
gaopengff
|
65fe32379e
|
[Intel GPU]Support fused_topk for XPU (#24641)
|
2026-05-20 10:34:29 +08:00 |
|
Xiaoyu Zhang
|
80fc524809
|
[diffusion] quant: update Wan2.2 modelOpt CI checkpoints (#25483)
|
2026-05-20 09:05:39 +08:00 |
|
Liangsheng Yin
|
7f154ba449
|
drop output ids (#25774)
|
2026-05-19 17:50:47 -07:00 |
|
 Paiiiiandzengpai
|
425dffbde3
|
DeepSeek V4 MTP Support CP (#24934)
Co-authored-by: zengpai <zengpai@baidu.com>
|
2026-05-19 16:51:31 -07:00 |
|
YAMY
|
beaff00331
|
[NSA] Avoid repeated NSA MQA logits memory queries (#25299)
|
2026-05-19 16:04:13 -07:00 |
|
 
|
b9d470f4a2
|
Support spec v2 for FlashMLA speculative decoding (#24640)
Co-authored-by: Jackey Hua <zhendonghua@users.noreply.github.com>
Co-authored-by: Depend <yu-depend@users.noreply.github.com>
|
2026-05-19 15:23:17 -07:00 |
|
 
|
b9c2bf717b
|
[BugFix] Resolve adaptive speculative decoding conflicts for Qwen3.5 (hybrid GDN) (#23331)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
|
2026-05-19 15:09:49 -07:00 |
|
Liangsheng Yin
|
16bcc4583e
|
verify_done: wait not synchronize (#25465)
|
2026-05-19 14:57:45 -07:00 |
|
Ratish P
|
fab097d66d
|
[Gemma4]: Fix FP8 Triton scale layout (#25286)
|
2026-05-19 14:00:23 -07:00 |
|
ybyang
|
8322fe09a7
|
fix(dsv4): upgrade forward metadata on main stream for large PP size (#25729)
|
2026-05-19 20:52:00 +00:00 |
|
 Yuan Luoandluoyuan.luo
|
4c0ce0345d
|
Support Gemma4 Pipeline Parallelism (#25284)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-19 22:40:11 +08:00 |
|
Makcum888e
|
3b62604cec
|
[Diffusion] Support parallelism for GLM-Image (#25645)
|
2026-05-19 17:27:21 +03:00 |
|
Xinyuan Tong
|
aad00b0ed8
|
Upgrade transformers to 5.8.1 (#25451)
|
2026-05-19 22:20:30 +08:00 |
|
Yuxuan Zhang
|
2bcb6d2f82
|
[Bug Fix] Align glm4_moe_nextn NPU MTP loading with qwen3 MTP (#25524)
|
2026-05-19 14:47:01 +01:00 |
|
Xiaoyu Zhang
|
0e4d1b49d3
|
[Codex] Remove stale DeepSeek V4 JIT kernels (#25764)
|
2026-05-19 20:04:32 +08:00 |
|
 Arseniy MironovandNapkin-AI
|
45a85efc3a
|
[Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
|
2026-05-19 12:46:55 +03:00 |
|
Thomas
|
58b5fe3e29
|
[Diffusion] [NPU] Fix HunyuanVideo crash on NPU (#25592)
|
2026-05-19 12:40:43 +03:00 |
|
jianzhao-xu
|
5073c82a37
|
transformers v5 adapt HFRunner (#23922)
|
2026-05-19 17:07:38 +08:00 |
|
shiyu7
|
7e0818038a
|
fix: fix deepseek v4 CP error (#25396)
|
2026-05-19 02:04:21 -07:00 |
|
 
|
67fd005b97
|
[HiSparse & PD] Support hisparse memory pool host page > 1 (#23606)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-05-19 01:29:35 -07:00 |
|
Ziang Li
|
78cb38ed5e
|
[FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918)
|
2026-05-19 01:04:48 -07:00 |
|
Kevin Li
|
fbfddfd5c7
|
fix (jit kernel): elementwise activation C++ error (#25695)
|
2026-05-19 15:23:52 +08:00 |
|
Yuhao Yang
|
79ea30d1f1
|
[Bug] Fix V4-Pro NaN on Blackwell by converting fp8_einsum input scale to ue8m0 (#25733)
|
2026-05-18 23:48:34 -07:00 |
|
 Junlin Wuandronnie_zheng
|
4c9f31b85e
|
✨ [diffusion][npu][quant] Add MXFP4 quantization support for Wan2.2 Diffusion on Ascend NPU (#22338)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-19 07:46:52 +03:00 |
|
Hanming Lu
|
862d39e06c
|
[Mamba] Fix extra_buffer overlap schedule races (#24954)
|
2026-05-19 12:13:20 +08:00 |
|
Zhonghua Deng
|
f0763859ed
|
perf(mimo-v2-epd): enable GPU image preprocess and parallel video decode (#25588)
|
2026-05-19 11:47:21 +08:00 |
|
Yuhao Yang
|
d8e66e54e5
|
fix: use triton_attn as default vision attention on B300 (SM103) (#25570)
|
2026-05-19 11:00:07 +08:00 |
|
Xiaoyu Zhang
|
31e324391b
|
[Codex] Opt Mistral Large performace (#24611)
|
2026-05-19 10:59:51 +08:00 |
|
Mick
|
a7b3ced334
|
[diffusion] fix: fix LTX2 resident defaults and stage profiling (#25596)
|
2026-05-19 10:41:28 +08:00 |
|
 ishandhananiandShangming Cai
|
87c3c96bc8
|
[Bug][PD][NIXL] always send aux on is_last; only expects_state when truthy (#25699)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-19 10:04:04 +08:00 |
|
 huangtingweiandhzh0425
|
c2a212bfe2
|
[UnifiedTree] Support DeepSeek V4 host pool with multiple layouts. (#25282)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-05-19 09:36:00 +08:00 |
|
Lianmin Zheng
|
b45b52ee8f
|
Add spec_verify_calls_total metric for speculative decoding (#25689)
|
2026-05-18 18:35:11 -07:00 |
|
fzyzcjy
|
e4d81e48c9
|
Pull the max-prefix-len computation into its own helper and rename the matched-token argument (#25728)
|
2026-05-19 09:27:06 +08:00 |
|
 Xiaoyu ZhangandCodex
|
2424303dfb
|
[codex] Optimize hidden-size 512 RMSNorm dispatch (#24710)
Co-authored-by: Codex <codex@example.com>
|
2026-05-19 09:26:10 +08:00 |
|