 DevashishLal-CBandDevashish Lal
|
6a3316dd1e
|
[plugin] default device detection fixes for OOT platform plugins (#25337)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
|
2026-06-06 07:55:50 +08:00 |
|
 ThanhhaoandHao Phan
|
2e7523ddbe
|
fix(spec-dec): treat num_nextn_predict_layers=0 the same as absent for EAGLE3 drafts (#26726)
Co-authored-by: Hao Phan <htphan@nvidia.com>
|
2026-06-05 16:08:07 -07:00 |
|
xutizhou
|
29591594f5
|
Support Waterfill with dynamic EPLB (#27150)
|
2026-06-05 16:01:16 -07:00 |
|
    
|
6b180959a8
|
[SPEC][5/N] feat: batchsize-aware support for adaptive speculative_num_steps (#24055)
Co-authored-by: 坤钧 <maoyuhan.myh@antgroup.co>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
|
2026-06-05 15:43:02 -07:00 |
|
  
|
c9f582a272
|
[LoRA] Experimental fast LoRA path with experimental_sgl_trtllm MoE backend for FP8 and NVFP4 models (#27329)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
|
2026-06-05 14:45:59 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
d381ec7997
|
[CI] Fix Nemotron nightly mixed precision checkpoints test (#27284)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
|
2026-06-05 13:51:26 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
3b62286fca
|
Reland "Support NextN = 2/4 in DSV32" (#27166)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-05 13:43:28 -07:00 |
|
Liangsheng Yin
|
2c2a4f243a
|
[Bug] Fix EAGLE draft CUDA-graph kv_indices under-allocation for topk > 1 (#27338)
|
2026-06-05 16:22:04 -04:00 |
|
Shangming Cai
|
57909f731b
|
[PD] Fix KV cache corruption on abort by notifying ongoing prefill (#27372)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-06-06 00:56:13 +08:00 |
|
 xlyandMick
|
d01cf27b7d
|
[diffusion] model: support Ideogram 4 FP8 (#27279)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-05 23:54:48 +08:00 |
|
 Michael_miaoande01448
|
e7f94d0d40
|
[Bug Fix] Fix activation.cuh JIT compilation failure on CUDA 13 due to template type/value mismatch (#26444)
Co-authored-by: e01448 <jwmao@birentech.com>
|
2026-06-05 22:59:24 +08:00 |
|
Aurick Qiao
|
c06802dc16
|
Fix customized_info incremental streaming (#27205)
|
2026-06-05 21:55:01 +08:00 |
|
Zhangheng
|
faa6286946
|
[BugFix]: Fix HiMamba HiCache prefetch hang after L3 sidecar transfer (#27366)
|
2026-06-05 20:37:40 +08:00 |
|
 XiaoTianandgongxiaotian
|
e1955bf57a
|
fix(pd): clear stale bootstrap_room when freeing metadata buffer slot (#27374)
Co-authored-by: gongxiaotian <gongxiaotian@didiglobal.com>
|
2026-06-05 19:46:35 +08:00 |
|
Wang, FangYuan
|
7f919edf00
|
[AMD] Support alt stream for Qwen3.5 on AMD platform (#25885)
|
2026-06-05 04:38:25 -07:00 |
|
R0CKSTAR
|
5d691a44f4
|
[diffusion] fix: fix LingBot World timestep error on MUSA (#27341)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-06-05 19:15:08 +08:00 |
|
Mick
|
4ef081b903
|
[diffusion] optimize: optimize LingBot realtime transport and camera conditioning (#27297)
|
2026-06-05 16:00:48 +08:00 |
|
huangtingwei
|
00fefef16b
|
[PD & HiSparse] Add DeepSeek V4 support for HiSparse direct Prefill-to-Decode DRAM (#24880)
|
2026-06-05 15:39:48 +08:00 |
|
Zhangheng
|
4df1ccdadc
|
[UnifiedTree]: Fix CP Reduce For L3 HiCache (#27330)
|
2026-06-05 14:03:54 +08:00 |
|
Qiaolin Yu
|
bd47869ba4
|
[perf] parallelize create_flashmla_kv_indices over page-blocks (#27320)
|
2026-06-04 22:11:43 -07:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
6cbc035dc9
|
FrozenKVMTPVerifyInput: add _draft_preprocess_idle call for when all requests in the verify batch finish in the same iteration (#26859)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Harmya Bhatt <harmyacs@gmail.com>
|
2026-06-04 21:47:32 -07:00 |
|
   
|
2c8357f794
|
[XPU] Enable Gemma 4 E2B / E4B / 31B/ 26B-A4B on Intel XPU (#23280)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: jmunetong <jmunetong@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: ckvermaAI <ckverma@habana.ai>
|
2026-06-05 10:05:07 +08:00 |
|
Zhangheng
|
631db6c757
|
[UnifiedTree]: Sync sidecar component hits across TP ranks and make SWA prefetch all-or-nothing (#27264)
|
2026-06-05 09:23:54 +08:00 |
|
 
|
5af02c18ae
|
[spec_v2] Enable trtllm_mha draft-extend CUDA graph with v2 semantics (#25002)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-04 17:50:12 -07:00 |
|
 Cheng WanandClaude Opus 4.8
|
7dc7376697
|
fix(attn): delegate init_mha_chunk_metadata in HybridLinearAttnBackend (#27316)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-04 17:44:23 -07:00 |
|
 Cheng WanandClaude Opus 4.8
|
0aa72a9e76
|
Replace skip_attn_backend_init with a batch-carried attention plan marker (+ staleness re-plan) (#27193)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-04 17:13:18 -07:00 |
|
Lianmin Zheng
|
448d3afb76
|
fix(spec): complete CustomSpecAlgo duck-typing interface and guard against drift (#27300)
|
2026-06-04 15:52:46 -07:00 |
|
  
|
e76d36214b
|
Changes for SM120 perf and usability for NVFP4 (#26496)
Co-authored-by: Martin Vit <martin@voipmonitor.org>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
|
2026-06-04 15:29:25 -07:00 |
|
 Bowen WangandXinyuan Tong
|
07f326c184
|
Fix multimodal synthetic benchmark prompt generation to exclude special tokens (#26864)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-06-04 22:27:43 +00:00 |
|
Ziang Li
|
4cfebbb95f
|
[FlashInfer v0.6.12] Support FlashInfer 4over6 NVFP4 (#25239)
|
2026-06-04 14:35:07 -07:00 |
|
 
|
5bf90ad988
|
Enable DeepGEMM PDL on by default (#23979)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-04 14:13:45 -07:00 |
|
Liangsheng Yin
|
cd98d97037
|
Use level-1 (quiet) busy memory check in chunked-prefill and streaming tests (#27303)
|
2026-06-04 17:04:27 -04:00 |
|
Liangsheng Yin
|
687cfe9198
|
Enable runtime busy memory check for speculation topk>1 (#27228)
|
2026-06-04 16:46:20 -04:00 |
|
Ilia Iliev
|
88a9d513e0
|
[Quant] Support asymmetric weight quant in compressed-tensors WNA16 (#25292)
|
2026-06-04 20:15:47 +00:00 |
|
YC Yen-Ching Tseng
|
69623f4b11
|
[AMD] Guard aiter greedy_sample OOB token id (fixes VLM MMMU CI) (#27247)
|
2026-06-04 12:53:58 -07:00 |
|
shuwenn
|
efe24704b7
|
[HiCache] feat: truncate mamba prefetch length to available host KV size (#26945)
|
2026-06-05 03:01:37 +08:00 |
|
 Chengze FanandHaiShaw
|
8e836e7dc9
|
Support optional kwargs in AITER fused_moe runner (#26746)
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-06-04 08:25:03 -07:00 |
|
Sam Shleifer
|
133254086b
|
Plug mamba_extra_buffer ping-pong slot leaks (#26941)
|
2026-06-04 21:46:36 +08:00 |
|
huangtingwei
|
b97a3dbb46
|
[HiSparse PD & PP]Fix HiSparse compatibility with PP decode (#27258)
|
2026-06-04 21:37:42 +08:00 |
|
liuxianglong17
|
8933ec8772
|
fix test cases failed on 5/30 in nightly pipeline (#26775)
|
2026-06-04 20:07:20 +08:00 |
|
Mick
|
e4a01d5ee2
|
[diffusion] fix: fix realtime webui recording timeline (#27249)
|
2026-06-04 19:58:56 +08:00 |
|
huangtingwei
|
a5c7e9d236
|
[HiCache] fix PD L3 cache hit details from decode responses (#27046)
|
2026-06-04 18:01:21 +08:00 |
|
Mick
|
9e2ad6054c
|
[diffusion] feat: speed up lossless realtime rgb transport (#27236)
|
2026-06-04 16:59:23 +08:00 |
|
Liangsheng Yin
|
ef170c27b6
|
Add quiet mode for busy mem check (level 1: buffer + dump on leak) (#27238)
|
2026-06-04 04:31:57 -04:00 |
|
jacky.cheng
|
7aee2ff31b
|
[AMD] Remove BF16-to-FP32 elementwise cast from compressor GEMM on HIP (#26914)
|
2026-06-04 00:58:02 -07:00 |
|
YC Yen-Ching Tseng
|
ff93a576e5
|
[AMD] Minimax M25 : FP8 block-scale GEMM dispatch for ROCm 7.0 on gfx950 (#27111)
|
2026-06-04 00:41:42 -07:00 |
|
shuwenn
|
1af53f67b2
|
[mem_cache][2/N] refactor: move SWATokenToKVPoolAllocator to allocator/swa.py (#26676)
|
2026-06-04 15:10:37 +08:00 |
|
Leon Gao
|
a10bd785be
|
Reduce mamba prefill allocation overhead (#25000)
|
2026-06-04 15:10:16 +08:00 |
|
 Alex SunandHaiShaw
|
e4191708c9
|
[Qwen3.5][AMD] Fix shared-expert ×ep_size over-count under allreduce-EP (#26845)
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-06-03 23:04:39 -07:00 |
|
Heyang Huang
|
858e5a5109
|
[diffusion] chore: disagg server args, launch helpers, and warmup utils (#26119)
|
2026-06-04 13:40:39 +08:00 |
|