18694 Commits
Author SHA1 Message Date
Mick 3394931044 [diffusion] optimize: optimize lingbot performance (#27023) 2026-06-02 18:33:06 +08:00
Mick a777672939 [diffusion] feat: enable parallel decode for cosmos3(#27037) 2026-06-02 18:18:18 +08:00
84e1108312 Optimize ngram decode id computation (#24757)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-02 17:37:34 +08:00
pengduriceandgithub-actions[bot] f651b48764 Apply apply_group_norm_silu to LTX-2 latent upsampler (#26045)
Signed-off-by: pengdurice <pengduhit@gmail.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-02 17:28:57 +08:00
Xiaoyu ZhangandBBuf 3ea1ba5b15 [GDN] Optimize prefill QKV split dispatch (#26206)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-02 16:48:31 +08:00
Xiaoyu ZhangandBBuf 559581b383 [codex] Centralize Triton utility kernels (#26000)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-02 16:47:45 +08:00
Bruce Changlong XuandKe Bao 172bd8e6b9 [scheduler] Zero gen_throughput and flush KV events on pause (#24003)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-06-02 16:43:04 +08:00
cctryandgemini-code-assist[bot] b55570d38e [PD] Optimistic prefill (#26780)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-02 01:16:14 -07:00
Charles Chen 5ae8d286d2 perf(gemma4): single-launch fused router (topk + softmax + scale) (#26502) 2026-06-02 16:00:17 +08:00
fzyzcjy 8cea0473ea Fix dp-attention token alignment in the dumper comparator e2e test (#26996) 2026-06-02 00:50:45 -07:00
3e993f6140 [PD]: Support HiCache prefetching and pd-incremental transfer on decode side (#26227)
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-06-02 15:40:10 +08:00
Hsiu-Chun, HungandHung 2582134a59 [AMD] Add amd ci mamba state scatter test (#26677)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
2026-06-02 00:24:59 -07:00
Jinyan ChenandJinyan Chen 301bcf0872 Add FP4 Indexer for DeepSeek V4 (#26209)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
2026-06-02 00:14:38 -07:00
Alison Shao 547b886b3c ci: drop redundant multimodal-server jobs from nightly (Nvidia) (#26985) 2026-06-01 23:40:48 -07:00
Mick a6985e1be0 [diffusion] doc: add cookbook for lingbot-world (#26958) 2026-06-02 14:35:22 +08:00
Mick 1033d835ff [diffusion] optimize: reduce cosmos3 denoise overhead (#26973) 2026-06-02 14:23:02 +08:00
Mick 3b26644bc4 [diffusion] misc: add realtime-webui (#26959) 2026-06-02 14:13:02 +08:00
Mick 2fc548f250 [diffusion] model: support lingot-world (#26954) 2026-06-02 13:52:49 +08:00
Liangsheng Yin f531bd7ff3 [Bug] Fix circular import in forward_batch_info from runtime cp_utils import (#27014) 2026-06-01 22:51:32 -07:00
Thomas Wang d15a2dc72c [AMD] dpsk-v4 swa loc cache support (#26931) 2026-06-01 22:37:07 -07:00
4226a6f13a [AMD] Fix GPT-OSS MXFP4 accuracy on ROCm AITER path (#26884)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-06-01 22:30:43 -07:00
Khoa PhamandCursor 08526c7fca [Spec] FrozenKVMTP fold assistant seed into captured draft graph (#25539)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-01 22:27:25 -07:00
Kurkur f27fa0da93 [NPU][Docs] Kimi-K2.5 best practice (#26774) 2026-06-02 13:14:14 +08:00
shuwenn f40b2ca9d3 chore(CODEOWNERS): add allocator/ owners and @alphabetc1 to mem_cache (#27003) 2026-06-01 22:07:15 -07:00
Ethan ZHUandZhangheng 594ec6335d [Bug Fix][HiCache] Drop @lru_cache on UnifiedTreeNode.get_prefix_hash_values (#26939)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-06-02 12:38:28 +08:00
1c0019da75 [Docs] GLM-4.7 cookbook: add NVIDIA Blackwell (B200, GB200) + NVFP4 sections (#26384)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 20:47:51 -07:00
popsiclexuandpopsiclexu 951fa05a09 [MoE] Support BF16 standard A2A with DeepGEMM runner (#26473)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
2026-06-01 20:40:38 -07:00
Teng MaandZijie Xia b562da0d9f [PD] docs: clarify disaggregation IB device formats (#25521)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-02 11:38:33 +08:00
ybyang 9fe8b72912 Speed up DeepGEMM JIT warmup with per-PP-rank parallel compile (#26567) 2026-06-01 19:51:27 -07:00
0574d2b8a5 [NVIDIA] [GDN] Enable FlashInfer MTP verify on SM100+ (Blackwell) (#23273)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 18:56:42 -07:00
Liangsheng Yin 54143264bf ci: disable cross-job fast-fail for run_all_tests dispatch (#26990) 2026-06-01 18:32:41 -07:00
98a1b58c47 docs(cookbook): port popular model usage guides into cookbook pages (#25813)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-01 17:41:49 -07:00
Liangsheng Yin 5e63200064 ci: full parallelism for run_all_tests dispatch (#26986) 2026-06-01 17:40:43 -07:00
Liangsheng Yin f6d0beaca8 Revert "Support spec v2 tree drafting (eagle topk>1) with page_size==1" (#26981) 2026-06-01 17:16:44 -07:00
Glen Liu 167272e785 [LoRA] add lora chunked req test and fix (#23179) 2026-06-01 16:25:27 -07:00
chenkaiyueandZhiqiang Xie dff45411da [HiCache] Prevent KV cache data loss when radix tree node is split b… (#16946)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-06-01 15:58:06 -07:00
Qiaolin Yu 4151a04d1a [Perf][Spec Decoding] Skip cat/topk/sort/gather in draft_forward for topk=1 (#26424) 2026-06-01 15:37:47 -07:00
Liangsheng Yin 1d4ee060c2 Support spec v2 tree drafting (eagle topk>1) with page_size==1 (#26866) 2026-06-01 15:37:20 -07:00
Baizhou Zhang 2fdae94e46 docs: update RTX PRO 6000 deployment snippet (#26968) 2026-06-01 14:34:27 -07:00
Yongfei Xu 5700790c05 DeepSeek V4: Support context parallelism with fused MoE (non-DeepEP) (#24947) 2026-06-01 14:25:43 -07:00
Qiaolin Yu 3bce192bd2 [misc] update adaptive spec decoding code owners (#26965) 2026-06-01 14:08:09 -07:00
eeechoandClaude Opus 4.6 524ba10eda feat: SM120 (Blackwell Desktop) support for DeepSeek-V4 inference (#24692)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-01 14:05:20 -07:00
Liangsheng Yin dfa1af99f5 Fix kill_process_tree reap wait crashing on pidfd EINVAL (#26964) 2026-06-01 13:56:59 -07:00
a0670b5ba3 [SPEC] feat: add adaptive speculative decoding metrics (#25940)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com>
2026-06-01 13:53:30 -07:00
Shu Wangandzijiexia 106092123f Update Qwen3-Coder docs_new NVIDIA guidance (#24435)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-01 13:38:34 -07:00
Khoa Pham da01f2974e [Log] include max_token_num and hidden_dim in FlashInfer workspace init log (#26605) 2026-06-01 13:26:52 -07:00
Mick f2beb7bc76 [diffusion] improve: avoid cosmos3 cpu float video postprocess (#26956) 2026-06-02 04:12:01 +08:00
Khoa PhamandClaude Opus 4.8 cb8a103b81 chore: add @pyc96 as codeowner for FrozenKVMTP module (#26953)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 13:06:58 -07:00
Mick 9a8ab2d22b [diffusion] fix: align cosmos3 text packing with official pipeline (#26950) 2026-06-02 02:07:17 +08:00
86afa21ca7 feat: optional caller-supplied mm_hashes on GenerateReqInput (#25300)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-06-01 20:04:37 +02:00