Ziang Li
|
4cfebbb95f
|
[FlashInfer v0.6.12] Support FlashInfer 4over6 NVFP4 (#25239)
|
2026-06-04 14:35:07 -07:00 |
|
 
|
5bf90ad988
|
Enable DeepGEMM PDL on by default (#23979)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-04 14:13:45 -07:00 |
|
Liangsheng Yin
|
cd98d97037
|
Use level-1 (quiet) busy memory check in chunked-prefill and streaming tests (#27303)
|
2026-06-04 17:04:27 -04:00 |
|
Liangsheng Yin
|
687cfe9198
|
Enable runtime busy memory check for speculation topk>1 (#27228)
|
2026-06-04 16:46:20 -04:00 |
|
Ilia Iliev
|
88a9d513e0
|
[Quant] Support asymmetric weight quant in compressed-tensors WNA16 (#25292)
|
2026-06-04 20:15:47 +00:00 |
|
YC Yen-Ching Tseng
|
69623f4b11
|
[AMD] Guard aiter greedy_sample OOB token id (fixes VLM MMMU CI) (#27247)
|
2026-06-04 12:53:58 -07:00 |
|
 Douglas YangandClaude Opus 4.8
|
75be922451
|
docs(cookbook): add Docker install option for Gemma 4 (#27287)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-04 12:45:32 -07:00 |
|
shuwenn
|
efe24704b7
|
[HiCache] feat: truncate mamba prefetch length to available host KV size (#26945)
|
2026-06-05 03:01:37 +08:00 |
|
 Chengze FanandHaiShaw
|
8e836e7dc9
|
Support optional kwargs in AITER fused_moe runner (#26746)
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-06-04 08:25:03 -07:00 |
|
Sam Shleifer
|
133254086b
|
Plug mamba_extra_buffer ping-pong slot leaks (#26941)
|
2026-06-04 21:46:36 +08:00 |
|
huangtingwei
|
b97a3dbb46
|
[HiSparse PD & PP]Fix HiSparse compatibility with PP decode (#27258)
|
2026-06-04 21:37:42 +08:00 |
|
liuxianglong17
|
8933ec8772
|
fix test cases failed on 5/30 in nightly pipeline (#26775)
|
2026-06-04 20:07:20 +08:00 |
|
liuxianglong17
|
4167f211b1
|
Solving the problem of test case failures caused by timeouts (#27027)
|
2026-06-04 20:07:04 +08:00 |
|
liuxianglong17
|
364c2bfea8
|
Test case restoration in the full test. (#26933)
|
2026-06-04 20:06:44 +08:00 |
|
Mick
|
e4a01d5ee2
|
[diffusion] fix: fix realtime webui recording timeline (#27249)
|
2026-06-04 19:58:56 +08:00 |
|
huangtingwei
|
a5c7e9d236
|
[HiCache] fix PD L3 cache hit details from decode responses (#27046)
|
2026-06-04 18:01:21 +08:00 |
|
Mick
|
9e2ad6054c
|
[diffusion] feat: speed up lossless realtime rgb transport (#27236)
|
2026-06-04 16:59:23 +08:00 |
|
Liangsheng Yin
|
ef170c27b6
|
Add quiet mode for busy mem check (level 1: buffer + dump on leak) (#27238)
|
2026-06-04 04:31:57 -04:00 |
|
Alison Shao
|
0783813fdf
|
ci: cache HF hub for base-a-test-cpu to avoid Hub 429 flakes (#27092)
|
2026-06-04 04:25:00 -04:00 |
|
jacky.cheng
|
7aee2ff31b
|
[AMD] Remove BF16-to-FP32 elementwise cast from compressor GEMM on HIP (#26914)
|
2026-06-04 00:58:02 -07:00 |
|
 Jinyan ChenandJinyan Chen
|
10b6b45cad
|
docs: add DeepSeek V4 FP4 indexer usage (#27035)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
|
2026-06-04 00:44:21 -07:00 |
|
YC Yen-Ching Tseng
|
ff93a576e5
|
[AMD] Minimax M25 : FP8 block-scale GEMM dispatch for ROCm 7.0 on gfx950 (#27111)
|
2026-06-04 00:41:42 -07:00 |
|
zijiexia
|
b89686710d
|
[Docs] re-organize nemotron cookbook (#27240)
|
2026-06-04 00:40:13 -07:00 |
|
  
|
1463e5fbdd
|
docs: add Nemotron 3 Ultra cookbook entry (#26969)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com>
|
2026-06-04 00:14:19 -07:00 |
|
shuwenn
|
1af53f67b2
|
[mem_cache][2/N] refactor: move SWATokenToKVPoolAllocator to allocator/swa.py (#26676)
|
2026-06-04 15:10:37 +08:00 |
|
Leon Gao
|
a10bd785be
|
Reduce mamba prefill allocation overhead (#25000)
|
2026-06-04 15:10:16 +08:00 |
|
 
|
04c16fc1e5
|
[AMD][CI] Remove transformers pin from GLM-5.x nightly jobs (#27232)
Co-authored-by: bingxche <bingxche@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-04 14:34:07 +08:00 |
|
Shangming Cai
|
f3acb6d4de
|
[CI] Fix multimodal-gen path filter for shared trace code (#27222)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-06-04 14:20:27 +08:00 |
|
 Alex SunandHaiShaw
|
e4191708c9
|
[Qwen3.5][AMD] Fix shared-expert ×ep_size over-count under allreduce-EP (#26845)
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-06-03 23:04:39 -07:00 |
|
Baizhou Zhang
|
cc67f922cd
|
[misc] Update Codeowner for Lora (#27224)
|
2026-06-03 22:50:15 -07:00 |
|
Heyang Huang
|
858e5a5109
|
[diffusion] chore: disagg server args, launch helpers, and warmup utils (#26119)
|
2026-06-04 13:40:39 +08:00 |
|
 Jitendra PatilandMa Mingfei
|
0e0ecc11ff
|
Add nightly Intel XPU Docker release workflow (#27182)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-04 13:39:57 +08:00 |
|
YC Yen-Ching Tseng
|
6dcd78a37f
|
[AMD] Add MiniMax-M2.5 TP=4 nightly accuracy test for MI355X (#27126)
|
2026-06-03 22:17:40 -07:00 |
|
Lianmin Zheng
|
ff5c4d7b57
|
Add zx3xyy to CI_PERMISSIONS.json (#27213)
|
2026-06-03 22:12:00 -07:00 |
|
Zhonghua Deng
|
e541bc3881
|
[EPD] feat: encoder DP mode with per-rank subprocess workers (#26576)
|
2026-06-04 12:37:41 +08:00 |
|
Yinghai Lu
|
71c759ebb7
|
[loader] Reduce transient allocations in NVFP4 MoE setup (#26861)
|
2026-06-03 21:13:25 -07:00 |
|
 
|
5c8a04ac4e
|
[XPU CI] Expand stage-a and consolidate stage-b tests into stage-a (#27156)
Co-authored-by: vshekhawat-hlab <vshekhawat@habana.ai>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-04 12:00:25 +08:00 |
|
 Cheng WanandClaude Opus 4.8
|
10ab7c919f
|
[refactor] Retire DecodeInputBuffers / PrefillInputBuffers in favor of CudaGraphBufferRegistry (#27192)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-03 20:52:56 -07:00 |
|
 Zaili WangandMa Mingfei
|
3b7a258f63
|
[CPU] upgrade dependent torch ver to PT2.12 (#21456)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-04 11:04:11 +08:00 |
|
王鹤男
|
29d23e198f
|
[diffusion] fix: preserve _explicit_fields across dataclasses.replace in DiffGenerator (#25308)
|
2026-06-04 11:02:23 +08:00 |
|
Mick
|
11605767e0
|
[diffusion] optimize: skip unused wanvae halo send copies (#27151)
|
2026-06-04 10:23:01 +08:00 |
|
 Zhanghengand晟海
|
736263f3dc
|
[UnifiedTree]: Support l3 storage for swa and deepseek v4 (#26881)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-06-04 10:17:34 +08:00 |
|
shuwenn
|
1a57145975
|
[codex] Fix adaptive metrics test flake (#27135)
|
2026-06-03 19:07:17 -07:00 |
|
 Yuzhen ZhouandJiajun Li
|
e03dfa8182
|
[3/N][Sync sglang-miles] TITO Support (#23751)
Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com>
|
2026-06-03 21:45:33 -04:00 |
|
Jonny Kong
|
084c6a7e2a
|
Refactor simulated acceptance length generation (#26768)
|
2026-06-03 18:31:32 -07:00 |
|
 
|
f6cd1a9822
|
Add num_waiting_uncached_tokens load metric (#27174)
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-03 18:29:49 -07:00 |
|
 ybyangandLianmin Zheng
|
687baf9471
|
fix(load-snapshot): avoid duplicate zmq bind in multi-tokenizer mode (#27145)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-03 18:24:56 -07:00 |
|
 
|
14ed9b448e
|
Add ZMQ IPv6 support, bench_serving sampling params, and reduce routed_dp_rank log noise (#27180)
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Grigory Sizov <grisha.sizov@gmail.com>
|
2026-06-03 17:49:34 -07:00 |
|
Mick
|
5dbc52c2b7
|
[diffusion] doc: add ernie Image diffusion (#27195)
|
2026-06-04 08:45:56 +08:00 |
|
Yinghai Lu
|
7bb5c96685
|
Trigger scheduler diagnostics on health failure (#26757)
|
2026-06-03 17:19:55 -07:00 |
|