       
|
85712fa5b0
|
Fix Responses API request handling (#25881)
Co-authored-by: Kai-Hsun Chen <kaihsun@apache.org>
Co-authored-by: Kristin Cowalcijk <kristincowalcijk@gmail.com>
Co-authored-by: aerosta <63026763+aerosta@users.noreply.github.com>
Co-authored-by: glaziermag <glaziermag@users.noreply.github.com>
Co-authored-by: Blake Ledden <blake.ledden@gmail.com>
Co-authored-by: PanJason <pyyjason@gmail.com>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: kennyu <966806+kennyu@users.noreply.github.com>
|
2026-06-12 14:47:55 -07:00 |
|
+3        
|
b3270264e4
|
Fix Anthropic Messages API compatibility (#25876)
Co-authored-by: Jairo David Campaña Rosero <jairocampana10001@gmail.com>
Co-authored-by: Karan Bansal <3264937+karanb192@users.noreply.github.com>
Co-authored-by: eason <85663565+mango766@users.noreply.github.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: qingchanghan <17794466+qingchanghan@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ajay Anubolu <124525760+AjAnubolu@users.noreply.github.com>
Co-authored-by: Ravitez Dondeti <13931987+dondetir@users.noreply.github.com>
Co-authored-by: Ratish P <114130421+Ratish1@users.noreply.github.com>
Co-authored-by: Xiaoshuai Zhang <15795935+jetd1@users.noreply.github.com>
Co-authored-by: Ricardo-M-L <69202550+Ricardo-M-L@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuan.tong@radixark.ai>
|
2026-06-12 14:46:57 -07:00 |
|
Liangsheng Yin
|
caf59759ea
|
[Spec] Centralize dummy verify-input capture; add carries_draft_hidden_states (#28032)
|
2026-06-12 14:23:29 -07:00 |
|
ybyang
|
1e71c1a859
|
fix(pd): disable overlap for spec+grammar in disagg decode loop (#28039)
|
2026-06-12 14:08:36 -07:00 |
|
Mohammad Miadh Angkad
|
cb9140ee61
|
Enable PDL for GPT-OSS tinygemm router (#27941)
|
2026-06-12 13:51:50 -07:00 |
|
shuwenn
|
9d37e710b7
|
[Bench] Add consistent p90/p95/p99 percentiles for all latency metrics (#27662)
|
2026-06-12 13:44:38 -07:00 |
|
Zach Zhu
|
627ed3476b
|
Fix invalid KVFP4QuantizeUtil references (#28013)
Signed-off-by: Zach Zhu <zzqshu@126.com>
|
2026-06-12 13:44:11 -07:00 |
|
 Vedant V JhaveriandVedant Jhaveri
|
3be5a7ec89
|
Respect explicit --max-running-requests instead of clamping to heuristic (#27399)
Co-authored-by: Vedant Jhaveri <vjhaveri@linkedin.com>
|
2026-06-12 12:30:47 -07:00 |
|
 
|
6e0fa5afe1
|
Support Nemotron DP attention and MTP (#24955)
Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com>
Co-authored-by: Zhichenzzz <northwesterniemsteaching@gmail.com>
|
2026-06-12 12:21:12 -07:00 |
|
David Wang
|
bb33594c1a
|
flashinfer swa kv pool fix (dflash gemma 4) (#27737)
|
2026-06-12 11:35:05 -07:00 |
|
cctry
|
75998d0421
|
Fix --mem-fraction-static not accounting for EAGLE draft model KV cache (#23862)
|
2026-06-12 10:35:48 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
c80d8fe78a
|
[Perf] Skip per-call mat_a/scales_a padding in cutlass FP8 blockwise GEMM (#27896)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-06-12 12:01:53 -04:00 |
|
 qiaozpandishandhanani
|
fcae6767d5
|
Add bucketed multi-dir layout for NIXL file storage (#27672)
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
|
2026-06-12 05:34:55 -07:00 |
|
luoroger37
|
18989f3d48
|
[PD] Fix resource leak on prealloc/transfer abort and idle check (#28022)
|
2026-06-12 18:20:24 +08:00 |
|
Xinyi Song
|
371b96e210
|
[AMD] Fix DeepSeek-V4-Flash-FP8 on MI300 (#27972)
|
2026-06-12 01:51:32 -07:00 |
|
Joectwm
|
694cea8656
|
Add EPD disaggregated encode tracing (#25994)
|
2026-06-12 16:16:14 +08:00 |
|
iridiumine
|
60e4f14953
|
[NPU][Bugfix] Fix accuracy issue in no-graph with MTP (#27752)
|
2026-06-12 16:05:06 +08:00 |
|
Liangsheng Yin
|
a52ccd2179
|
[Spec] Install EagleDraftExtendInput as the V2 draft-extend spec_info (#24860)
|
2026-06-12 00:46:16 -07:00 |
|
 Bi Xueandispobock
|
3c1f9eafa5
|
[sgl] proactively release out-of-window SWA slots after chunked prefill (#27402)
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2026-06-12 14:17:18 +08:00 |
|
 Xinyi SongandThomas Wang
|
36d61613a1
|
[AMD] Cache unified_kv swa_loc once per step instead of per layer (#27978)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
|
2026-06-11 23:13:15 -07:00 |
|
Jumiar
|
ca17bd8347
|
perf: eliminate CUDA syncs in VLM embed path (#26082)
|
2026-06-12 13:15:56 +08:00 |
|
Mick
|
8038806557
|
[diffusion] optimize: optimize flux1 tensor parallel sharding (#27826)
|
2026-06-12 13:13:18 +08:00 |
|
  
|
0a1fb0da86
|
add LRU eviction for mooncacke embedding cache (#25954)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: siyu <liusy58@linux.alibaba.com>
Co-authored-by: Yuang Chen <1131578721@qq.com>
|
2026-06-12 12:38:13 +08:00 |
|
Liangsheng Yin
|
e1164a6dfc
|
[Spec] Remove dead prepare_for_verify / prepare_extend_after_decode + extend-decode kernel (#27761)
|
2026-06-11 20:59:30 -07:00 |
|
syy-hw
|
3a3a759464
|
[NPU] Add Gemma4 Sliding Window Attention support on Ascend backend (#26147)
|
2026-06-12 11:44:38 +08:00 |
|
 Brayden ZhongandBrayden Zhong
|
8bfcc0c39c
|
Use the correct wrapper for fp4_quantize (#27956)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-11 20:23:09 -07:00 |
|
Yihao Wang
|
f5c9f88ee2
|
[plugin][distributed] use active platform's backend in get_default_distributed_backend (#23969)
|
2026-06-11 20:05:17 -07:00 |
|
 
|
7074704c0c
|
[Spec] Dedup post-verify mamba state commit into shared spec_utils helpers (#27966)
Co-authored-by: xbfs <xuebf1@lenovo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-11 18:33:40 -07:00 |
|
Liangsheng Yin
|
3ffe72517f
|
[Spec] Remove the dead spec V1 scheduler paths (#27977)
|
2026-06-11 18:31:13 -07:00 |
|
Cheng Wan
|
2e74ff192c
|
[DSV4] Use int64 for compressor out_loc tensors (#27973)
|
2026-06-11 17:45:34 -07:00 |
|
Cheng Wan
|
97a0031799
|
[lint] Enable Ruff UP037 to drop redundant quoted annotations (#27984)
|
2026-06-11 17:38:05 -07:00 |
|
Liangsheng Yin
|
c0480a88be
|
[Spec] Retire Spec V1 (#27964)
|
2026-06-11 16:15:15 -07:00 |
|
Oguz Ulgen
|
949326d922
|
Add SGLANG_ENABLE_WAR_BARRIER to force-enable the overlap scheduler WAR barrier on non-CUDA (e.g. AMD) (#27967)
|
2026-06-11 15:38:37 -07:00 |
|
 
|
d71e9bede6
|
[bugfix] commit Mamba states after NGRAM target verify (#26351)
Co-authored-by: xbfs <xuebf1@lenovo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-11 15:11:18 -07:00 |
|
Yueming Yuan
|
cd075d1f64
|
[RL] convert DeepSeek V4 APE layout through weight loader (#27307)
|
2026-06-11 15:05:27 -07:00 |
|
Liangsheng Yin
|
fee717f303
|
[Spec] Fold the DFLASH worker base into DFlashWorkerV2 on BaseSpecWorker (#27950)
|
2026-06-11 14:54:28 -07:00 |
|
Liangsheng Yin
|
acdb39edd2
|
[Spec] Remove the DFLASH V1 worker path (#27959)
|
2026-06-11 14:37:58 -07:00 |
|
Wang, FangYuan
|
6e885c844f
|
Revert "[AMD] Fix DeepSeek V4 Pro c128 state tensor dtype mismatch error and c4_sparse_raw_indices attribute error in cuda graph phase" (#27919)
|
2026-06-11 14:25:32 -07:00 |
|
Liangsheng Yin
|
df5055e00f
|
Bump spec logprob match delta for the bf16 eagle fixture (#27952)
|
2026-06-11 14:21:07 -07:00 |
|
  ![github-actions[bot]](/assets/img/avatar_default.png)
|
ec0eb6cce8
|
Support MiMo v2 ASR (#26278)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: yanyihan <yanyihan@xiaomi.com>
Co-authored-by: zhuqingchao <zhuqingchao@xiaomi.com>
|
2026-06-11 13:43:13 -07:00 |
|
Jackey Hua
|
24c5d76f74
|
fix: per-sequence last-token embedding in EAGLE3/MTP draft for batched multimodal spec decoding (#27846)
|
2026-06-11 13:33:16 -07:00 |
|
 cctryandcctry
|
10219bd9d6
|
[PD] Fix negative prefill kv_transfer_alloc_ms under optimistic prefill (#27885)
Co-authored-by: cctry <cctry@fb.com>
|
2026-06-11 13:12:25 -07:00 |
|
Yuwei An
|
880e6f66fc
|
[BCG] Share output buffers across capture sizes + typed ShapeKey (#27857)
|
2026-06-11 11:58:05 -07:00 |
|
Brian Chao
|
7f57b344c9
|
[diffusion] feat: progressive resolution growing for Ideogram 4 via GPU DCT upsampling with up to 1.56× speedup (#27736)
|
2026-06-11 23:16:53 +08:00 |
|
 Chi McIsaacandMick
|
b2728bda9d
|
[diffusion] feat: use fused w8a8 kernel for Ideogram4 weight-only linear as an opt-in (#27590)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-11 23:15:27 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
06e0df5899
|
Optimize Qwen3 Next FP8 MoE on H200 (#26204)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-11 22:18:43 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
1a6b5561db
|
Fix MLA scaling when YARN scaling is disabled (#26203)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-11 22:17:55 +08:00 |
|
Xiaoyu Zhang
|
d571e076fa
|
[codex] Centralize more inline Triton kernels (#27429)
|
2026-06-11 22:17:26 +08:00 |
|
Mick
|
d9110d971e
|
[diffusion] fix: fix wan ti2v sp timestep padding (#27876)
|
2026-06-11 22:13:25 +08:00 |
|
ybyang
|
8077fb1df7
|
fix(deepgemm): align PP-parallel warmup bs to CP padding (#27922)
|
2026-06-11 20:52:45 +08:00 |
|