Qiaolin Yu
|
1a17d753f1
|
[perf] prepare_prefill_qkv hook + fp8 quantize jit kernel (#25460)
|
2026-05-20 14:20:49 -07:00 |
|
Yuzhen Zhou
|
dac78768f0
|
[RL][TITO] Preserve whitespace in reasoning parser outputs (#24251)
|
2026-05-20 19:45:09 +00:00 |
|
YAMY
|
801d7e3eed
|
[DSA] Make MQA logits free memory ratio configurable (#25859)
|
2026-05-20 12:27:16 -07:00 |
|
Ratish P
|
5e7bf73757
|
Fix bench_serving non-stream reasoning content (#25298)
|
2026-05-20 18:41:46 +00:00 |
|
jasonjk-park
|
1f209b4433
|
Add support for generic num_tokens_per_bs in TARGET_VERIFY (#25681)
|
2026-05-20 11:34:45 -07:00 |
|
 Mark Smithandzhangjiadong1@corp.netease.com
|
33c57b8716
|
[Bug][RadixTree] Fix LRU list reference cycle leak in radix_cache (#25770)
Co-authored-by: zhangjiadong1@corp.netease.com <zhangjiadong1@corp.netease.com>
|
2026-05-21 01:06:05 +08:00 |
|
Jialin Ouyang
|
6e0b7f35ad
|
[radix cache] pluggable RadixCache factory (--radix-cache-backend) (#25101)
|
2026-05-20 10:05:04 -07:00 |
|
Xiaoyu Zhang
|
ccbbae00ea
|
[codex] Reland Wan2.2 ModelOpt CI checkpoints (#25857)
|
2026-05-20 22:15:25 +08:00 |
|
Liwansi
|
55ba03db6a
|
[NPU]use triton split_qkvgate_gemma_rmsnorm_rope for Qwen3.5 and Qwen3_next (#23925)
|
2026-05-20 20:22:10 +08:00 |
|
Liangsheng Yin
|
34d3e23232
|
spec_v2: consolidate seq_lens_cpu/sum maintenance into helper (#25818)
|
2026-05-20 04:42:26 -07:00 |
|
Liangsheng Yin
|
9b005d3608
|
disagg prebuilt: drop dead prepare_for_extend shift (#25819)
|
2026-05-20 04:39:47 -07:00 |
|
Liangsheng Yin
|
1bd4f94598
|
[Test] Add fwd_occupancy sanity kit (#25886)
|
2026-05-20 03:34:37 -07:00 |
|
Chi McIsaac
|
47979fb252
|
[diffusion] fix: fix GLM-Image /v1/images/edits support (#25697)
|
2026-05-20 17:11:51 +08:00 |
|
Liangsheng Yin
|
614672fea5
|
[Test] Stage-a sanity kits; consolidate core/ + models_e2e/ tests (#25831)
|
2026-05-20 01:58:48 -07:00 |
|
Yuhong Guo
|
24d27c2035
|
[BugFix] Fix rid_to_state leak for aborted queued requests (#24070)
|
2026-05-20 01:32:44 -07:00 |
|
 Matt Van HornandMatt Van Horn
|
e99f87c974
|
fix: add missing distro dependency to runtime docker image (#25817)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
|
2026-05-20 00:46:32 -07:00 |
|
 
|
044649c23a
|
feat: Support flashinfer_cutedsl MoE runner with flashinfer alltoall backend (#22669)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-20 00:36:26 -07:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
8131641bc6
|
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-20 00:18:04 -07:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Xiaobaoandgithub-actions[bot]
|
0d3a94f643
|
[Bug] Correct Weight Offloader's Attribute Name for torch.nn.Parameter (#25786)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-05-19 22:28:59 -07:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
052abcc0dd
|
[Refactor] Pass PP start_layer via model constructor instead of forward_batch.token_to_kv_pool (#25825)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-19 22:16:07 -07:00 |
|
Cheng Wan
|
a4b51d35ef
|
Revert "[codex] Update Wan2.2 ModelOpt CI checkpoints" (#25845)
|
2026-05-19 21:45:20 -07:00 |
|
Xinyuan Tong
|
0aedc5678b
|
loader: yield filtered MTP weights lazily to avoid OOM hang on multi-layer EAGLE (#25748)
|
2026-05-20 12:33:54 +08:00 |
|
Xiaoyu Zhang
|
af22390af7
|
[codex] Align diffusion skills with nightly Nvidia benchmarks (#25842)
|
2026-05-20 12:18:05 +08:00 |
|
Shangming Cai
|
1fbee74fb6
|
[PD] Clean early abort logic in PD module (#25677)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-20 11:17:42 +08:00 |
|
Kevin Flansburg
|
3b2178c412
|
[PD] Un-blacklist mooncake sessions when probe succeeds (#25287)
|
2026-05-20 11:16:28 +08:00 |
|
mingyue300
|
5fe655bf26
|
vlm: fix shared memory bug of deep copy mm_items (#24231)
|
2026-05-20 11:06:51 +08:00 |
|
ybyang
|
a8c82c652e
|
fix(dsv4): make pool configurator PP-aware (#25750)
|
2026-05-20 10:58:53 +08:00 |
|
ybyang
|
ca29c2b0e7
|
fix(dsv4): drop stale pp_size=1 guard for V4 PD disaggregation (#25771)
|
2026-05-20 10:57:23 +08:00 |
|
jy-song-hub
|
d69e9fbfdc
|
[diffusion] fix: honor config precisions for delight/paint (#22289)
|
2026-05-20 10:38:48 +08:00 |
|
 jy-song-hubandjxp
|
549ae16c6f
|
[diffusion] fix: respect configured precision in Qwen layered path (#21980)
Co-authored-by: jxp <jingxin.pan123@gmail.com>
|
2026-05-20 10:38:03 +08:00 |
|
jy-song-hub
|
3a9d9d5832
|
[diffusion] fix: fix Hunyuan3D-2 DiT checkpoint param mapping (#22729)
|
2026-05-20 10:37:13 +08:00 |
|
gaopengff
|
65fe32379e
|
[Intel GPU]Support fused_topk for XPU (#24641)
|
2026-05-20 10:34:29 +08:00 |
|
Xiaoyu Zhang
|
80fc524809
|
[diffusion] quant: update Wan2.2 modelOpt CI checkpoints (#25483)
|
2026-05-20 09:05:39 +08:00 |
|
Liangsheng Yin
|
7f154ba449
|
drop output ids (#25774)
|
2026-05-19 17:50:47 -07:00 |
|
 Paiiiiandzengpai
|
425dffbde3
|
DeepSeek V4 MTP Support CP (#24934)
Co-authored-by: zengpai <zengpai@baidu.com>
|
2026-05-19 16:51:31 -07:00 |
|
YAMY
|
beaff00331
|
[NSA] Avoid repeated NSA MQA logits memory queries (#25299)
|
2026-05-19 16:04:13 -07:00 |
|
 
|
b9d470f4a2
|
Support spec v2 for FlashMLA speculative decoding (#24640)
Co-authored-by: Jackey Hua <zhendonghua@users.noreply.github.com>
Co-authored-by: Depend <yu-depend@users.noreply.github.com>
|
2026-05-19 15:23:17 -07:00 |
|
 
|
b9c2bf717b
|
[BugFix] Resolve adaptive speculative decoding conflicts for Qwen3.5 (hybrid GDN) (#23331)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
|
2026-05-19 15:09:49 -07:00 |
|
Liangsheng Yin
|
16bcc4583e
|
verify_done: wait not synchronize (#25465)
|
2026-05-19 14:57:45 -07:00 |
|
Ratish P
|
fab097d66d
|
[Gemma4]: Fix FP8 Triton scale layout (#25286)
|
2026-05-19 14:00:23 -07:00 |
|
ybyang
|
8322fe09a7
|
fix(dsv4): upgrade forward metadata on main stream for large PP size (#25729)
|
2026-05-19 20:52:00 +00:00 |
|
 Yuan Luoandluoyuan.luo
|
4c0ce0345d
|
Support Gemma4 Pipeline Parallelism (#25284)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-19 22:40:11 +08:00 |
|
Makcum888e
|
3b62604cec
|
[Diffusion] Support parallelism for GLM-Image (#25645)
|
2026-05-19 17:27:21 +03:00 |
|
Xinyuan Tong
|
aad00b0ed8
|
Upgrade transformers to 5.8.1 (#25451)
|
2026-05-19 22:20:30 +08:00 |
|
Yuxuan Zhang
|
2bcb6d2f82
|
[Bug Fix] Align glm4_moe_nextn NPU MTP loading with qwen3 MTP (#25524)
|
2026-05-19 14:47:01 +01:00 |
|
Xiaoyu Zhang
|
0e4d1b49d3
|
[Codex] Remove stale DeepSeek V4 JIT kernels (#25764)
|
2026-05-19 20:04:32 +08:00 |
|
 Arseniy MironovandNapkin-AI
|
45a85efc3a
|
[Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
|
2026-05-19 12:46:55 +03:00 |
|
Thomas
|
58b5fe3e29
|
[Diffusion] [NPU] Fix HunyuanVideo crash on NPU (#25592)
|
2026-05-19 12:40:43 +03:00 |
|
jianzhao-xu
|
5073c82a37
|
transformers v5 adapt HFRunner (#23922)
|
2026-05-19 17:07:38 +08:00 |
|
shiyu7
|
7e0818038a
|
fix: fix deepseek v4 CP error (#25396)
|
2026-05-19 02:04:21 -07:00 |
|