 huangtingweiandhzh0425
|
c2a212bfe2
|
[UnifiedTree] Support DeepSeek V4 host pool with multiple layouts. (#25282)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-05-19 09:36:00 +08:00 |
|
Lianmin Zheng
|
b45b52ee8f
|
Add spec_verify_calls_total metric for speculative decoding (#25689)
|
2026-05-18 18:35:11 -07:00 |
|
fzyzcjy
|
e4d81e48c9
|
Pull the max-prefix-len computation into its own helper and rename the matched-token argument (#25728)
|
2026-05-19 09:27:06 +08:00 |
|
 Xiaoyu ZhangandCodex
|
2424303dfb
|
[codex] Optimize hidden-size 512 RMSNorm dispatch (#24710)
Co-authored-by: Codex <codex@example.com>
|
2026-05-19 09:26:10 +08:00 |
|
fzyzcjy
|
170fe57cf0
|
Encapsulate the pending-flush bookkeeping in a small wrapper (#25727)
|
2026-05-19 09:23:10 +08:00 |
|
fzyzcjy
|
5067da3eb6
|
Confine req-pool-idx assignment to the pool allocator (#25726)
|
2026-05-19 09:22:51 +08:00 |
|
fzyzcjy
|
6bcbf4de35
|
Fix the misnamed request finish-check method to reflect its mutating semantics (#25725)
|
2026-05-19 09:22:28 +08:00 |
|
fzyzcjy
|
2d2dff28da
|
Return a mamba tracking entry from the cache lookup instead of mutating caller lists (#25724)
|
2026-05-19 09:21:55 +08:00 |
|
fzyzcjy
|
3fd6a58e6c
|
Inline the single-use split-prefill setup at its caller (#25722)
|
2026-05-19 09:21:29 +08:00 |
|
fzyzcjy
|
1f3e5aa1e0
|
Publish elastic-EP active ranks from a dedicated step (#25721)
|
2026-05-19 09:21:04 +08:00 |
|
fzyzcjy
|
fa37b68653
|
Rename the request mid-chunk flag to describe what it actually tracks (#25720)
|
2026-05-19 09:20:43 +08:00 |
|
fzyzcjy
|
32f1259c91
|
Confine max-prefix-len to where it is used and drop the leftover variable (#25719)
|
2026-05-19 09:20:21 +08:00 |
|
fzyzcjy
|
1cba3ab467
|
Stop returning the unused prefix-computed flag from priority calc (#25718)
|
2026-05-19 09:20:00 +08:00 |
|
fzyzcjy
|
2d868656d0
|
Move the retract-decode ratio estimation onto the new-token-ratio tracker (#25717)
|
2026-05-19 09:19:42 +08:00 |
|
fzyzcjy
|
1a882c5c63
|
Pack scattered new-token-ratio state into a dedicated tracker (#25716)
|
2026-05-19 09:19:20 +08:00 |
|
fzyzcjy
|
954b5c5846
|
Pack scattered scheduler IPC channel state into a dedicated container (#25714)
|
2026-05-19 09:19:02 +08:00 |
|
fzyzcjy
|
07b4f262b7
|
Set up the idle sleeper outside of the IPC channel initialization (#25713)
|
2026-05-19 09:18:36 +08:00 |
|
fzyzcjy
|
2d40f45193
|
Pack scattered request logprob state into a dedicated container (#25712)
|
2026-05-19 09:18:19 +08:00 |
|
fzyzcjy
|
fb7e49d4eb
|
Expose can-run-cuda-graph as a regular property on the embedding result (#25711)
|
2026-05-19 09:18:00 +08:00 |
|
fzyzcjy
|
a740f8de33
|
Remove the dead hasattr fallback around the test-only crash counter (#25710)
|
2026-05-19 09:17:40 +08:00 |
|
fzyzcjy
|
0e198f0f4f
|
Refactor batch_result_processor into per-step prefill/decode helpers (#25709)
|
2026-05-19 09:17:17 +08:00 |
|
fzyzcjy
|
7e7cb969e9
|
Route streaming output through the accumulator's payload method instead of an inline send (#25708)
|
2026-05-19 09:16:42 +08:00 |
|
fzyzcjy
|
da50e3d943
|
Log per-request time stats in a dedicated tail step (#25707)
|
2026-05-19 09:16:22 +08:00 |
|
fzyzcjy
|
b911fd1673
|
Route streaming-accept decisions through the accumulator instead of an inline gate (#25706)
|
2026-05-19 09:15:45 +08:00 |
|
fzyzcjy
|
d8f190dfba
|
Pack scattered output-streamer state into a dedicated accumulator (#25705)
|
2026-05-19 09:14:59 +08:00 |
|
fzyzcjy
|
e8e55bb19b
|
Split the request-reception loop into smaller phases (#25703)
|
2026-05-19 09:12:47 +08:00 |
|
Kurkur
|
d90bc65e30
|
[NPU] Fix TypeError in get_state_buf_infos when index_head_dim is None on MLA (#25383)
|
2026-05-19 09:09:11 +08:00 |
|
Kurkur
|
d028697d17
|
[NPU][Docs] Add Kimi-K2.5-W4A8 instance doc on NPU (#25269)
|
2026-05-19 09:08:28 +08:00 |
|
Ratish P
|
dbac464726
|
[Spec]: Make Triton standalone spec test deterministic (#25303)
|
2026-05-18 17:36:10 -07:00 |
|
 Baizhou Zhangandhnyls2002
|
b79e4b1e68
|
[Fix] Try to fix error caused by latest cutedsl packages (#25690)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-05-18 16:51:32 -07:00 |
|
 
|
314dedf7c6
|
Use SGLANG_CACHE_DIR env for gpu_p2p_access_cache path (#25686)
Co-authored-by: Ian O'Connell <ianoc@meta.com>
Co-authored-by: ianoc <ianoc@fb.com>
|
2026-05-18 16:00:54 -07:00 |
|
 Lianmin ZhengandHanming Lu
|
745abd6cc0
|
Add no_combine support to cutlass_moe_fp4 (#25688)
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
|
2026-05-18 15:59:55 -07:00 |
|
 
|
878e6b8886
|
[SP] Fix runtime_max_tokens_per_rank for sequence parallelism (#25685)
Co-authored-by: Ming Yang <minos.future@gmail.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
|
2026-05-18 15:59:33 -07:00 |
|
Baizhou Zhang
|
6f892047ec
|
[misc] Throw error when single batch overlap is enabled on Hopper (#25509)
|
2026-05-18 14:51:32 -07:00 |
|
Liangsheng Yin
|
c904fdd20e
|
ci: pr-states match renamed "PR Test Base" workflow_run (#25687)
|
2026-05-18 14:42:29 -07:00 |
|
Liangsheng Yin
|
9e3bb9a307
|
[Spec] fold can_run_cuda_graph into EagleVerifyOutput; drop dead extend-after-decode check (#25566)
|
2026-05-18 14:24:54 -07:00 |
|
Cheng Wan
|
b7267e8fce
|
[CI] Enable weight prefetch for 8-gpu-h200 basic tests (#25684)
|
2026-05-18 14:12:45 -07:00 |
|
Qiaolin Yu
|
1f185c6ba8
|
Support draft extend cuda graph for tokenspeed_mla attention backend (#25489)
|
2026-05-18 11:26:16 -07:00 |
|
Qiaolin Yu
|
f5049709b3
|
fix(eagle3): drop +1 offset on aux layer ids when first id != 1 (#25454)
|
2026-05-18 11:25:51 -07:00 |
|
Shangming Cai
|
86c6c77f2f
|
[Bugfix] Fix missing group arg in get dp buffer (#25585)
|
2026-05-19 02:08:10 +08:00 |
|
 Khoa PhamandClaude Opus 4.7
|
b29e41e8b3
|
Respect user override for Gemma4 attention backend (#25547)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-18 10:46:25 -07:00 |
|
Lianmin Zheng
|
f21fe6ad4d
|
Fix PD disaggregation warmup: set request_name and improve error logging (#25542)
|
2026-05-18 09:49:56 -07:00 |
|
Gaoji Liu
|
3e2a109636
|
Add output_gate_type to Qwen3NextConfig and update models to utilize it (#25401)
|
2026-05-19 00:18:08 +08:00 |
|
 
|
d96e593fd0
|
[Benchmark] Add SGLANG_SIMULATE_UNIFORM_EXPERTS for balanced expert routing with dummy weights (#25571)
Co-authored-by: Byron Hsu <byronhsu@Byrons-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-18 09:16:12 -07:00 |
|
+3        
|
866793c502
|
Amd/deepseek v4 rebase main 0509 (#24933)
Co-authored-by: root <root@smci355-ccs-aus-m12-33.cs-aus.dcgpu>
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-18 09:15:07 -07:00 |
|
Mick
|
110bbdcad7
|
[diffusion] fix: use dynamic LoRA for LTX2 original stage-two (#25591)
|
2026-05-18 23:05:29 +08:00 |
|
ybyang
|
d1acd62d29
|
fix(disagg): unstuck decode aborts under prealloc pressure (#25561)
|
2026-05-18 22:57:22 +08:00 |
|
 grunerandAmit Gruner
|
0ab427d0e1
|
[SMG] Add /v1/models fallback for model name discovery (#25293)
Co-authored-by: Amit Gruner <agruner@crusoe.ai>
|
2026-05-18 22:02:35 +08:00 |
|
  
|
ba2ffcf156
|
Add DeepSeekV4 fused MoE Triton autotune support (#25569)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
|
2026-05-18 21:35:32 +08:00 |
|
Shangming Cai
|
f04c522534
|
[PD] Add conclude_state to fake KV backend (#25599)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-18 19:56:42 +08:00 |
|