Hanming Lu
|
b603f08c0c
|
[DP] Fix FlashInfer dispatcher workspace sizing and set_dp_buffer_len (#26643)
|
2026-06-02 13:25:20 -07:00 |
|
Bi Xue
|
9e717cae46
|
[sglang] Fix Mamba COW over-releasing SWA locks (cascade-evict assert crash) (#27038)
|
2026-06-03 01:42:04 +08:00 |
|
Muqi Li
|
6ba31e33e6
|
feat(api): add require_reasoning field for engine's generate api (#27019)
|
2026-06-02 17:36:29 +00:00 |
|
 Cheng WanandClaude Opus 4.7
|
99da43b900
|
[refactor] init_forward_metadata 3-method ABC + side-channel removal + ForwardMetadata type rename (#26735)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-06-02 10:33:33 -07:00 |
|
Ilia Yastrebov
|
6c69756fa8
|
NIXL: use prep+make API to improve performance (#26406)
|
2026-06-02 18:36:35 +02:00 |
|
Ke Bao
|
28f9c1ff24
|
Relax mamba unified cache kl threshold (#27070)
|
2026-06-02 23:28:22 +08:00 |
|
Mick
|
22043b917b
|
[diffusion] CI: ad lingbot case (#27055)
|
2026-06-02 22:26:51 +08:00 |
|
Ke Bao
|
b5e154dc73
|
Fix stale import after kl_nightly rename (#27064)
|
2026-06-02 21:46:21 +08:00 |
|
Ke Bao
|
38d4c9ba88
|
Improve type annotations in unified radix cache (#26948)
|
2026-06-02 21:17:50 +08:00 |
|
littleyellowbicycle
|
7271318dc3
|
【docs】The remote weight download function has been adjusted to be unsupported until the PTA interface is fixed. (#27050)
|
2026-06-02 20:55:23 +08:00 |
|
Mick
|
64a1dec8b6
|
[diffusion] feat: add realtime webui super resolution controls (#27026)
|
2026-06-02 20:29:21 +08:00 |
|
Mick
|
ce7da7397a
|
[diffusion] optimize: optimize cosmos3 (#27041)
|
2026-06-02 19:47:42 +08:00 |
|
Librua
|
c2eea4d7b3
|
[Bugfix] Fix orphaned aborted prefill bootstrap requests in PP disaggregation (#27028)
|
2026-06-02 19:03:46 +08:00 |
|
Zhangheng
|
ee4bf0a9d3
|
[UnifiedTree]: Add HiCache Nightly CI For GLM5 (#26927)
|
2026-06-02 19:02:50 +08:00 |
|
Mick
|
3394931044
|
[diffusion] optimize: optimize lingbot performance (#27023)
|
2026-06-02 18:33:06 +08:00 |
|
Mick
|
a777672939
|
[diffusion] feat: enable parallel decode for cosmos3(#27037)
|
2026-06-02 18:18:18 +08:00 |
|
 
|
84e1108312
|
Optimize ngram decode id computation (#24757)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-02 17:37:34 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) pengduriceandgithub-actions[bot]
|
f651b48764
|
Apply apply_group_norm_silu to LTX-2 latent upsampler (#26045)
Signed-off-by: pengdurice <pengduhit@gmail.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-06-02 17:28:57 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
3ea1ba5b15
|
[GDN] Optimize prefill QKV split dispatch (#26206)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-02 16:48:31 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
559581b383
|
[codex] Centralize Triton utility kernels (#26000)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-02 16:47:45 +08:00 |
|
 Bruce Changlong XuandKe Bao
|
172bd8e6b9
|
[scheduler] Zero gen_throughput and flush KV events on pause (#24003)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-06-02 16:43:04 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) cctryandgemini-code-assist[bot]
|
b55570d38e
|
[PD] Optimistic prefill (#26780)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-06-02 01:16:14 -07:00 |
|
Charles Chen
|
5ae8d286d2
|
perf(gemma4): single-launch fused router (topk + softmax + scale) (#26502)
|
2026-06-02 16:00:17 +08:00 |
|
fzyzcjy
|
8cea0473ea
|
Fix dp-attention token alignment in the dumper comparator e2e test (#26996)
|
2026-06-02 00:50:45 -07:00 |
|
  
|
3e993f6140
|
[PD]: Support HiCache prefetching and pd-incremental transfer on decode side (#26227)
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-06-02 15:40:10 +08:00 |
|
 Hsiu-Chun, HungandHung
|
2582134a59
|
[AMD] Add amd ci mamba state scatter test (#26677)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
|
2026-06-02 00:24:59 -07:00 |
|
 Jinyan ChenandJinyan Chen
|
301bcf0872
|
Add FP4 Indexer for DeepSeek V4 (#26209)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
|
2026-06-02 00:14:38 -07:00 |
|
Alison Shao
|
547b886b3c
|
ci: drop redundant multimodal-server jobs from nightly (Nvidia) (#26985)
|
2026-06-01 23:40:48 -07:00 |
|
Mick
|
a6985e1be0
|
[diffusion] doc: add cookbook for lingbot-world (#26958)
|
2026-06-02 14:35:22 +08:00 |
|
Mick
|
1033d835ff
|
[diffusion] optimize: reduce cosmos3 denoise overhead (#26973)
|
2026-06-02 14:23:02 +08:00 |
|
Mick
|
3b26644bc4
|
[diffusion] misc: add realtime-webui (#26959)
|
2026-06-02 14:13:02 +08:00 |
|
Mick
|
2fc548f250
|
[diffusion] model: support lingot-world (#26954)
|
2026-06-02 13:52:49 +08:00 |
|
Liangsheng Yin
|
f531bd7ff3
|
[Bug] Fix circular import in forward_batch_info from runtime cp_utils import (#27014)
|
2026-06-01 22:51:32 -07:00 |
|
Thomas Wang
|
d15a2dc72c
|
[AMD] dpsk-v4 swa loc cache support (#26931)
|
2026-06-01 22:37:07 -07:00 |
|
 
|
4226a6f13a
|
[AMD] Fix GPT-OSS MXFP4 accuracy on ROCm AITER path (#26884)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-06-01 22:30:43 -07:00 |
|
 Khoa PhamandCursor
|
08526c7fca
|
[Spec] FrozenKVMTP fold assistant seed into captured draft graph (#25539)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-01 22:27:25 -07:00 |
|
Kurkur
|
f27fa0da93
|
[NPU][Docs] Kimi-K2.5 best practice (#26774)
|
2026-06-02 13:14:14 +08:00 |
|
shuwenn
|
f40b2ca9d3
|
chore(CODEOWNERS): add allocator/ owners and @alphabetc1 to mem_cache (#27003)
|
2026-06-01 22:07:15 -07:00 |
|
 Ethan ZHUandZhangheng
|
594ec6335d
|
[Bug Fix][HiCache] Drop @lru_cache on UnifiedTreeNode.get_prefix_hash_values (#26939)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-06-02 12:38:28 +08:00 |
|
  
|
1c0019da75
|
[Docs] GLM-4.7 cookbook: add NVIDIA Blackwell (B200, GB200) + NVFP4 sections (#26384)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-01 20:47:51 -07:00 |
|
 popsiclexuandpopsiclexu
|
951fa05a09
|
[MoE] Support BF16 standard A2A with DeepGEMM runner (#26473)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
|
2026-06-01 20:40:38 -07:00 |
|
 Teng MaandZijie Xia
|
b562da0d9f
|
[PD] docs: clarify disaggregation IB device formats (#25521)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-06-02 11:38:33 +08:00 |
|
ybyang
|
9fe8b72912
|
Speed up DeepGEMM JIT warmup with per-PP-rank parallel compile (#26567)
|
2026-06-01 19:51:27 -07:00 |
|
 
|
0574d2b8a5
|
[NVIDIA] [GDN] Enable FlashInfer MTP verify on SM100+ (Blackwell) (#23273)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-01 18:56:42 -07:00 |
|
Liangsheng Yin
|
54143264bf
|
ci: disable cross-job fast-fail for run_all_tests dispatch (#26990)
|
2026-06-01 18:32:41 -07:00 |
|
  
|
98a1b58c47
|
docs(cookbook): port popular model usage guides into cookbook pages (#25813)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-06-01 17:41:49 -07:00 |
|
Liangsheng Yin
|
5e63200064
|
ci: full parallelism for run_all_tests dispatch (#26986)
|
2026-06-01 17:40:43 -07:00 |
|
Liangsheng Yin
|
f6d0beaca8
|
Revert "Support spec v2 tree drafting (eagle topk>1) with page_size==1" (#26981)
|
2026-06-01 17:16:44 -07:00 |
|
Glen Liu
|
167272e785
|
[LoRA] add lora chunked req test and fix (#23179)
|
2026-06-01 16:25:27 -07:00 |
|
 chenkaiyueandZhiqiang Xie
|
dff45411da
|
[HiCache] Prevent KV cache data loss when radix tree node is split b… (#16946)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-06-01 15:58:06 -07:00 |
|