Commit Graph
9190 Commits
Author SHA1 Message Date
Zhangheng fe548f36b0 [UnifiedTree]: Fix SWA admission budget under-counts HiCache load-back consumption (#27391) 2026-06-07 10:47:08 +08:00
shuwenn e57323cae9 [mem_cache][4/N] refactor: extract MambaTokenToKVPoolAllocator into allocator/ (#27256) 2026-06-07 10:46:29 +08:00
Trevor Morris 5da265de30 [NVIDIA] Fix FP8 gemm performance with fp16 models (MInimax-M2.5) (#22300) 2026-06-07 02:45:00 +00:00
Xiaoyu Zhang 2c3e84affe [Diffusion] Enable Cosmos3 denoising profiling (#27439) 2026-06-07 10:32:28 +08:00
Yueming Yuan 4c8a022f38 Fix DeepSeek V4 DP reduce scatter when use attention DP + MoE TP (#27191) 2026-06-06 18:24:33 -07:00
Liangsheng Yin 5160f7914e Fix MLA EAGLE draft CUDA-graph kv_indices under-allocation for topk > 1 (#27460) 2026-06-06 16:28:34 -07:00
Liangsheng Yin 032c9efb46 Enable async-assert invariant probes by default in CI (#27461) 2026-06-06 16:23:21 -07:00
Qiaolin Yu 4b0f629082 [perf] reduce radix cache match overhead by changing the match algorithm (#27364) 2026-06-06 15:40:28 -07:00
Liangsheng Yin 1c7acba579 [spec] Consolidate the per-decode KV alloc reserve into one helper (#27458) 2026-06-06 14:58:50 -07:00
Cheng WanandClaude Opus 4.8 9097647090 Route the eager forward path through the CUDA graph input-buffer registry (#27407)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 14:35:53 -07:00
Liangsheng Yin 84ca0ffb8c Spec v2 tree drafting (topk>1) with page_size>1 (#26972) 2026-06-06 12:00:27 -07:00
88a7b0fd30 Classify malformed-multimodal rejects as invalid_request (#27451)
Co-authored-by: cctry <cctry@meta.com>
Co-authored-by: cctry <cctry@fb.com>
2026-06-06 10:19:22 -07:00
bd7fea0740 [diffusion] Fix LingBot-World crash on camera control with ulysses>1 (#27437)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 22:02:29 +08:00
Xiaoyu Zhang 7c6f9542c7 [Diffusion] Avoid GPU syncs in UniPC scheduler (#27440) 2026-06-06 22:01:41 +08:00
4e14b50c48 [NPU]Support torch_npu profiler patch API drift (#26356)
Co-authored-by: leland17 <lileliao@foxmail.com>
Co-authored-by: OmX <omx@oh-my-codex.dev>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-06 16:27:51 +03:00
shuwenn 280280ace9 [SPEC] fix: import copy module for eagle sampling info clone (#27390) 2026-06-06 12:12:46 +00:00
fzyzcjy 99cf0a4399 Fix flaky test_self_e2e_pd_perturb (#27426) 2026-06-06 19:31:08 +08:00
42fe025280 [HiCache] Fix the compatibility between PP and HiCache (L2). (#27285)
Co-authored-by: ybyang <ybyang7@iflytek.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: shangmingc <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-06-06 16:57:23 +08:00
Liangsheng Yin aa5213abb1 [debug] Register #27338 EAGLE draft kv_indices revert in pr_fix_toggle (#27428) 2026-06-06 00:30:13 -07:00
Liangsheng Yin 26b9053dcc [Spec] Fix fa3 EAGLE draft-decode expand page_table scatter OOB for topk>1 + page_size>1 (#27360) 2026-06-06 00:24:34 -07:00
EduardDurech e9dbbd19e9 [model] Apertus Tool/Function and Reasoning parser (#25100) 2026-06-06 00:04:31 -07:00
Qiaolin Yu 8c47b7678a [attn backend] clean legacy init_mha_chunk_metadata in trtllm_mla backend (#27403) 2026-06-05 23:30:21 -07:00
Xiaoyu ZhangandBBuf f57f8a8afd Optimize Gemma4 H200 MoE and extend attention (#26588)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-06 14:14:25 +08:00
e513c13e2e Optimize ngram decode token table update (#24756)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-06 14:13:45 +08:00
zijiexiaandClaude Opus 4.8 9da88e32e0 [Cohere2Moe] Enable flashinfer_trtllm NVFP4 fused-MoE via SigmoidRenorm routing (#27401)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 22:49:20 -07:00
Brayden ZhongandBrayden Zhong 38ae22e08c Nemotron perf changes (#26733)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-05 22:31:46 -07:00
Yuxuan Zhang 393d0e169e [Bugfix] Restore overridden HF config fields and support index_skip_topk_offset for DSA topk sharing (#27114) 2026-06-06 13:26:04 +08:00
kkandwunhuang aa55657e9e [AMD][WA] force to use gate_mode interleaved to fix tp2/tp4/tp8 acc issue (#27201)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-06-05 20:18:11 -07:00
Mick bf66b7b6da [diffusion] model: support Ideogram4 NVFP4 (#27379) 2026-06-06 11:14:28 +08:00
Mick e8668508d1 [diffusion] optimize: optimize LingBot realtime sp cache path (#27383) 2026-06-06 09:37:29 +08:00
Liangsheng Yin 58a05d3dd2 [CI] Isolate CUDA coredump dir per run to fix tracker mis-attribution (#27344) 2026-06-05 21:31:22 -04:00
Xinyi SongandHaiShaw 3030119ef7 [bugfix][AMD] AttributeError and warp mask bugs in DeepSeek V4 FP4 indexer (#27152)
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-06-05 18:26:38 -07:00
Chi McIsaac 25d8f431d1 [diffusion] optimize: cosmos3 fused qknorm rope (#27096) 2026-06-06 09:15:42 +08:00
fzyzcjy bf4f2ccc78 Add scripted-runtime unit, core integration, and chunked-prefill tests (#27413) 2026-06-06 09:08:35 +08:00
fzyzcjy 5a82db85f0 Add scripted-runtime KV-pool and lock-ref exhauster primitives (#27412) 2026-06-06 09:07:40 +08:00
fzyzcjy dd176387c6 Add scripted-runtime harness core and wire scheduler/IPC hooks (#27411) 2026-06-06 09:07:01 +08:00
fzyzcjy 21201ef718 Add kv_canary PP self-test fixture and SWA divergence coverage (#27410) 2026-06-06 09:06:19 +08:00
cctry b3e4c204fd Don't write crash dump on graceful exit (#27405) 2026-06-05 17:06:39 -07:00
DevashishLal-CBandDevashish Lal 6a3316dd1e [plugin] default device detection fixes for OOT platform plugins (#25337)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
2026-06-06 07:55:50 +08:00
ThanhhaoandHao Phan 2e7523ddbe fix(spec-dec): treat num_nextn_predict_layers=0 the same as absent for EAGLE3 drafts (#26726)
Co-authored-by: Hao Phan <htphan@nvidia.com>
2026-06-05 16:08:07 -07:00
xutizhou 29591594f5 Support Waterfill with dynamic EPLB (#27150) 2026-06-05 16:01:16 -07:00
6b180959a8 [SPEC][5/N] feat: batchsize-aware support for adaptive speculative_num_steps (#24055)
Co-authored-by: 坤钧 <maoyuhan.myh@antgroup.co>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
2026-06-05 15:43:02 -07:00
c9f582a272 [LoRA] Experimental fast LoRA path with experimental_sgl_trtllm MoE backend for FP8 and NVFP4 models (#27329)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
2026-06-05 14:45:59 -07:00
Brayden ZhongandBrayden Zhong d381ec7997 [CI] Fix Nemotron nightly mixed precision checkpoints test (#27284)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-05 13:51:26 -07:00
Brayden ZhongandBrayden Zhong 3b62286fca Reland "Support NextN = 2/4 in DSV32" (#27166)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-05 13:43:28 -07:00
Liangsheng Yin 2c2a4f243a [Bug] Fix EAGLE draft CUDA-graph kv_indices under-allocation for topk > 1 (#27338) 2026-06-05 16:22:04 -04:00
Shangming Cai 57909f731b [PD] Fix KV cache corruption on abort by notifying ongoing prefill (#27372)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-06-06 00:56:13 +08:00
xlyandMick d01cf27b7d [diffusion] model: support Ideogram 4 FP8 (#27279)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-05 23:54:48 +08:00
Michael_miaoande01448 e7f94d0d40 [Bug Fix] Fix activation.cuh JIT compilation failure on CUDA 13 due to template type/value mismatch (#26444)
Co-authored-by: e01448 <jwmao@birentech.com>
2026-06-05 22:59:24 +08:00
Aurick Qiao c06802dc16 Fix customized_info incremental streaming (#27205) 2026-06-05 21:55:01 +08:00