 FilipandClaude Opus 4.8
|
19d3f86895
|
[LoRA] Laguna: per-layer LoRA hidden-dim resolution for packed attention (#30298)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-08-04 14:48:05 -07:00 |
|
 Xinyuan TongandAlex Nails
|
a9c3b55435
|
[Refactor] Keep chat template validation out of ServerArgs dispatcher (#33392)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-04 14:45:47 -07:00 |
|
 Lianmin Zhengandcctry
|
34af3ff386
|
Allow optimistic prefill with L2 hierarchical cache and write-back policy (#33545)
Co-authored-by: cctry <cctry@meta.com>
|
2026-08-04 13:23:08 -07:00 |
|
+26        
|
abddb1c7e9
|
[Kimi] Support kimi-k3 (#32541)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai>
Co-authored-by: Zijie Xia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: zhangxiaohao <1024393531@qq.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: RolaoDenthu <xinyisong0111@gmail.com>
Co-authored-by: pigeonsoup <32922982+pigeonsoup@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com>
Co-authored-by: Lee Nau <lee.nau@gmail.com>
Co-authored-by: HMING <126185151+Hearum@users.noreply.github.com>
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: Hanming Lu <hanminglu@meta.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
|
2026-08-04 13:22:49 -07:00 |
|
Xingyu Liu
|
aa06433709
|
Avoid TRTLLM prefill output copy (#33306)
|
2026-08-04 12:54:04 -07:00 |
|
 
|
38dc2d6cf8
|
[metrics] Split tokenizer request metrics by stream (#32734)
Co-authored-by: wpc <wpc@devvm23443.cco0.facebook.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-08-04 12:53:52 -07:00 |
|
Lianmin Zheng
|
4794b401d5
|
[Observability] Add startup, memory, and hybrid SWA diagnostics (#33375)
|
2026-08-04 12:50:09 -07:00 |
|
 Lianmin ZhengandJialin Ouyang
|
5081c063c0
|
fix(metrics): clear forward occupancy on idle (#33562)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
|
2026-08-04 12:49:17 -07:00 |
|
 Lianmin ZhengandItai Gat
|
dea2be5ae3
|
[CUDA Graph] Allow custom decode graph runners (#33553)
Co-authored-by: Itai Gat <itaigat.mail@gmail.com>
|
2026-08-04 12:48:56 -07:00 |
|
 
|
e76d0acdc9
|
migrate NPU PR/nightly test cases to a3-560T (#33346)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-08-05 01:17:58 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
95d0e57e83
|
[diffusion] Fuse DiT FFN tanh-GELU into up-proj GEMM (cublasLt epilogue) behind quality=high (Qwen-Image 1024^2 denoise 12.36 -> 12.05 s on H200) (#33536)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-04 23:49:43 +08:00 |
|
Xiaoyu Zhang
|
0d0c7d853f
|
[diffusion] FLUX.2 VAE decoder fast path behind quality=high (H200: 1024^2 97.6->29.2 ms, 2048^2 437.2->168.5 ms) (#33451)
|
2026-08-04 23:48:19 +08:00 |
|
Lianmin Zheng
|
7adf2f4a9a
|
Inline _set_gc into _set_envs_and_config (#33538)
|
2026-08-04 07:31:56 -07:00 |
|
 Lianmin ZhengandYinghai Lu
|
d257b58e67
|
[Router] Report accelerator count in /v1/loads (#33548)
Co-authored-by: Yinghai Lu <yinghai@meta.com>
|
2026-08-04 05:55:35 -07:00 |
|
Lianmin Zheng
|
8f2a3ad6d7
|
[mem_cache] Label HiCache host pools and clarify post-capture KV sizing logs (#33445)
|
2026-08-04 04:21:36 -07:00 |
|
jacky.cheng
|
723c277640
|
[AMD] [Fix] Enable aiter hd256 FP8 prefill FMHA on gfx950 (#33399)
|
2026-08-04 02:40:33 -07:00 |
|
Xingyu Liu
|
5e6c37f2b4
|
[cuda_graph] Gate breakable-CG capture_inputs retention to DP-gather paths (#32678)
Signed-off-by: xingyuliu <charlotteliu12x@gmail.com>
|
2026-08-04 02:23:34 -07:00 |
|
 
|
b57721ccf7
|
Enable post-capture KV sizing with DP attention (#33427)
Co-authored-by: cctry <cctry@meta.com>
Co-authored-by: cctry <cctry@fb.com>
|
2026-08-04 02:20:24 -07:00 |
|
Lianmin Zheng
|
16d3b118a2
|
Reduce startup log noise and fix Dynamo / CUDA-graph edge cases (#33428)
|
2026-08-04 02:19:56 -07:00 |
|
Liangsheng Yin
|
b6d548afd7
|
[Fix] Resolve VLM test image placeholders from the model's own chat template (#33509)
|
2026-08-04 02:01:39 -07:00 |
|
 ashwini rathiandMa Mingfei
|
53804d609c
|
[CI][XPU] Stabilize XPU CI: pin UMD/IGC, retry infra flakes, right-size EAGLE3 (#32438)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-08-04 16:28:52 +08:00 |
|
 Kan WuandClaude Fable 5
|
17d19081d9
|
[mm] sglang-mm: server vision pipeline core (fetch/driver/pipeline) + Qwen VL (#32364)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-04 00:47:00 -07:00 |
|
 Vladislav NosivskoyandXinyuan Tong
|
154f0ac662
|
Fix DSpark and DP/EP (#33098)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-08-04 00:35:57 -07:00 |
|
 Alex NailsandClaude Opus 5
|
bfa4e4a57b
|
[Nemotron] Hoist mamba track-mask host syncs out of the per-layer prefill path (#32589)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-03 23:58:28 -07:00 |
|
 weireweireandweireweire
|
23ea7b6481
|
Prewarm DSV4 MHC post kernel at model load (#30741)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-08-03 23:41:56 -07:00 |
|
 TobyMintandMick
|
101bb2327c
|
[diffusion] fix: fix local-path detection for MiniMax-H3 and other non-diffusers models (#33365)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-04 14:03:35 +08:00 |
|
 Liangsheng YinandMick
|
afc868517b
|
[Perf] Speed up the Kimi-K2.5 vision path and match PIL bicubic in the GPU resize (#33349)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-03 21:57:49 -07:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
c6f2a9c1d4
|
[diffusion] Restrict request-level quality to two validated tiers: lossless (default) and high (#33453)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-04 11:43:16 +08:00 |
|
 Jinchen HanandMick
|
614825fd38
|
[vla] fix: pi05 models does not apply scale factor for language embeddings (#33367)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-04 11:25:54 +08:00 |
|
Mick
|
b058dc9106
|
[diffusion] fix: reject ring parallelism where it would silently miscompute (#33353)
|
2026-08-04 11:15:23 +08:00 |
|
EchO
|
1e64fc1563
|
[Fix] Honor FlashMLA natural-log LSE in DCP reduction (#33065)
|
2026-08-03 19:55:52 -07:00 |
|
Khoa Pham
|
91fae8a72c
|
[DCP] Bound a request by the aggregate KV pool, not one rank's share (#33448)
|
2026-08-03 19:36:25 -07:00 |
|
 Oguz UlgenandCheng Wan
|
c113ead98a
|
Bump helion version to 1.4 (#32562)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-08-03 19:10:31 -07:00 |
|
 Yuwei AnandClaude Opus 5
|
92087ef4d2
|
fix(mem_cache): state the MLA KV bound in the DCP index space (#33432)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-03 19:07:49 -07:00 |
|
Baizhou Zhang
|
eb31a53338
|
Revert "Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120" (#33455)
|
2026-08-03 19:03:56 -07:00 |
|
Xingyu Liu
|
572924634b
|
[mem_cache] Build empty-prefix last_loc sentinel on-device to avoid per-call H2D sync (#32575)
Signed-off-by: xingyuliu <charlotteliu12x@gmail.com>
|
2026-08-03 18:55:03 -07:00 |
|
Baizhou Zhang
|
7f6a2e2b50
|
[Refactor] Clean up and split DSA indexer (#33443)
|
2026-08-03 18:05:01 -07:00 |
|
 
|
3960983753
|
Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120 (#32994)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-04 08:37:42 +08:00 |
|
Xiaoyu Zhang
|
f829fb3ff7
|
[diffusion] Fix component accuracy topology reuse (#33317)
|
2026-08-04 08:35:03 +08:00 |
|
 
|
1307968605
|
[JIT] Drop redundant per-kernel arch overrides (#32952)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-04 08:33:54 +08:00 |
|
Po-Han Huang (NVIDIA)
|
7eb27372b3
|
fix(server): capture legal multi-request prefill CUDA graph batches (#30206)
|
2026-08-03 16:56:59 -07:00 |
|
 zijiexiaandClaude Opus 4.8
|
b819d2fb5b
|
[Docs] Rename docs_new/ to docs/ (#32123)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-08-03 16:51:00 -07:00 |
|
Hanming Lu
|
bc7e1a07c3
|
Bound prefill delayer all-branch delay and decay the max_prefill_bs high-watermark (#32880)
|
2026-08-03 14:43:02 -07:00 |
|
cctry
|
db0fe370b7
|
[PD] Fix false health-503 during decode retraction re-admission (#33118)
|
2026-08-03 09:44:25 -07:00 |
|
cctry
|
6cf661117d
|
[PD] Add a queues.prealloc_ready counter to the load snapshot (#33133)
|
2026-08-03 09:44:11 -07:00 |
|
Mick
|
0ba46c88e5
|
[diffusion] CI: add minimax-h3 2-gpu consistency coverage (#33281)
|
2026-08-03 22:40:08 +08:00 |
|
 huangtingweiandjackyYang6
|
3953788596
|
[HiSparse]Fix DeepSeek V4 HiSparse PD Transfers with Separate Host and Device KV Indices (#31901)
Co-authored-by: jackyYang6 <82102811+jackyYang6@users.noreply.github.com>
|
2026-08-03 20:35:04 +08:00 |
|
Mohammad Miadh Angkad
|
d48ab2d386
|
Fix BCG circular import during server startup (#33371)
|
2026-08-03 01:53:16 -07:00 |
|
Xun Sun
|
a2d1003b18
|
[Mooncake] Fix ProcessGroup API imports (#32403)
|
2026-08-03 01:16:41 -07:00 |
|
amd-oshkarav
|
92999d84f4
|
[AMD]Qwen3.5 integration gfx950 fmha fp8 hd256 (#32046)
|
2026-08-03 00:25:31 -07:00 |
|