Thomas Wang
|
27596abdc0
|
[AMD] Update amd k3 cookbook for PR#34580 (#35263)
|
2026-08-17 23:19:36 -07:00 |
|
 Alexandrandronnie_zheng
|
61600c9f39
|
[Diffusion][Refactor] Refactor and extract complex RoPE implementation to layers/rotary_embedding for MOVA DiT (#31453)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-08-18 09:14:35 +03:00 |
|
Jiajun Li
|
e6df23f3c2
|
refactor: rename chat response token IDs (#35225)
|
2026-08-17 23:07:43 -07:00 |
|
Shuwen Wang
|
0077f84d37
|
[mem_cache][8/N] refactor: move MambaPoolHost to pool_host.mamba (#31180)
|
2026-08-18 05:22:45 +00:00 |
|
Colin Z
|
ea27e3ddab
|
[AMD] Fix Quark Shared Experts Fusion Gate after load-time-override Removal (#35200)
|
2026-08-17 22:10:01 -07:00 |
|
Zhiyao Jiang
|
9401db3f29
|
[AMD] Scope the EAGLE greedy-verify TP broadcast to ROCm only (#35195)
|
2026-08-17 21:38:24 -07:00 |
|
 
|
8ea5229d42
|
[AMD] Add the Kimi-K3 MI35x perf benchmarks in nightly (#34985)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
|
2026-08-17 21:23:51 -07:00 |
|
Ke Bao
|
d528192bf9
|
Skip inkling sheared bias under batch invariance (#35161)
|
2026-08-18 12:22:58 +08:00 |
|
 amd-danli103andThomas Wang
|
d01812d89e
|
[AMD] Optimize KIMI-K3 with Triton MLA decode kernel by tuning the stage-1 geometry for gfx950 (#34580)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-08-17 21:16:14 -07:00 |
|
Baizhou Zhang
|
53621818e4
|
[Docs] Enable PD disaggregation for DSV4 low-latency recipes (#35224)
|
2026-08-17 20:07:27 -07:00 |
|
 Khoa PhamandClaude Opus 5
|
f44a130c5e
|
[DCP] Drop the prefill index-selection syncs by taking each rank's rows by stride (#35084)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-17 19:55:51 -07:00 |
|
Siyuan Chen
|
fcdaaf8a5d
|
[Feature] Optimize TP LMHead with All-to-All (#32313)
|
2026-08-17 19:55:27 -07:00 |
|
 jiayisunxandMa Mingfei
|
d6c837489a
|
[XPU] Enable fused GDN QKV split Triton kernel on XPU (#30144)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-08-18 10:42:58 +08:00 |
|
Ma Mingfei
|
7f51e6bba0
|
[CPU] Explicitly import sgl_kernel in CPU kernel tests (#35119)
|
2026-08-18 10:20:28 +08:00 |
|
 MickandYiqi Yang
|
d55f1c28e2
|
[diffusion] feat: load quantized H3 text encoder checkpoints (#34986)
Co-authored-by: Yiqi Yang <yangyiqi8787@gmail.com>
|
2026-08-18 09:10:54 +08:00 |
|
Zaili Wang
|
0ea262e6e5
|
[XPU] xpu kernel release workflow (#33679)
|
2026-08-17 18:07:48 -07:00 |
|
Baizhou Zhang
|
c2c1c4d3eb
|
[CI] Move DSA PD+MTP+CP Layersplit test to basic B300 test suite (#35220)
|
2026-08-17 17:44:31 -07:00 |
|
 sglang-botandsglang-bot
|
9ffc2856fb
|
docs: sync LMSYS SGLang blog cards (#35218)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-08-17 17:30:28 -07:00 |
|
Faradawn Yang
|
91144797c5
|
Update Qwen3.5 H200 FP8 for AgentX HiCache MTP (#35194)
|
2026-08-17 17:12:07 -07:00 |
|
Liangsheng Yin
|
c0b6474b43
|
[Spec] Reduce host-side overhead in ngram draft prep (#35207)
|
2026-08-17 16:40:06 -07:00 |
|
Cheng Wan
|
b3c8f0d923
|
config: one control-plane log for the process (#35028)
|
2026-08-17 16:19:00 -07:00 |
|
Cheng Wan
|
c70c7d72a8
|
config: the readback and the resolving view say what they are (#35027)
|
2026-08-17 16:18:19 -07:00 |
|
Cheng Wan
|
cba3c5d5ac
|
config: the per-instance families read the bags (#35026)
|
2026-08-17 16:17:53 -07:00 |
|
Cheng Wan
|
a97bc8db32
|
config: the DP/EP topology reads come from the parallel bag (#35025)
|
2026-08-17 16:17:26 -07:00 |
|
Cheng Wan
|
d2bc697396
|
spec: size the speculative buffers from the bags, not the startup record (#35024)
|
2026-08-17 16:16:56 -07:00 |
|
Cheng Wan
|
3d7ec00179
|
config: publish before a process reads configuration (#35023)
|
2026-08-17 16:16:20 -07:00 |
|
Cheng Wan
|
2b278b4ac4
|
config: retire the multi-engine accommodation in the runtime context (#35022)
|
2026-08-17 16:15:33 -07:00 |
|
Baizhou Zhang
|
bc312d185d
|
Clean deprecated DeepSeek V4 Environs (#34926)
|
2026-08-17 16:07:00 -07:00 |
|
Lennox Fu
|
b42abbb1ba
|
[metrics] Fix prefill FLOPs estimate to count prefix and per-request causal pairs (#34316)
|
2026-08-17 15:18:54 -07:00 |
|
 Jimmy ShongandClaude Opus 5
|
b956e916ae
|
docs(cookbook): add Qwen3.8-27B DGX Spark configs (#35121)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-17 15:04:17 -07:00 |
|
 
|
816ea65058
|
[AMD] Add Kimi-K3 8-GPU MI35x nightly accuracy CI (#32568)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
|
2026-08-17 14:48:29 -07:00 |
|
Lianmin Zheng
|
198a7b2fc9
|
[Misc] Clean up python/sglang package structure (#35062)
|
2026-08-17 14:24:35 -07:00 |
|
 
|
770e7b47a2
|
[Spec] Support output logprobs with DSpark (#34478)
Co-authored-by: zhisbug <1654062+zhisbug@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-08-17 14:23:24 -07:00 |
|
Liangsheng Yin
|
032fe9c891
|
[Spec] Relay ngram accept tokens through the FutureMap (#35198)
|
2026-08-17 14:21:07 -07:00 |
|
Yuhao Yang
|
861eca8e25
|
docs: add NVFP4 quantization option to Kimi-K3 deploy panel (#35168)
|
2026-08-17 11:01:47 -07:00 |
|
 cctryandcctry
|
2e7c85da68
|
[PD] Preserve decode KV across retraction in HiCache (#34801)
Co-authored-by: cctry <cctry@fb.com>
|
2026-08-17 08:49:11 -07:00 |
|
Lianmin Zheng
|
af743371cc
|
Clean up environ.py: remove dead env vars, unify deprecation handling, move examples to a unit test (#35060)
|
2026-08-17 06:53:34 -07:00 |
|
Mick
|
d97b796c16
|
[diffusion] chore: reuse SRT CLIP encoder blocks (#35004)
|
2026-08-17 19:51:51 +08:00 |
|
Ke Bao
|
6e8a4abb57
|
Add bit-exact class for MTP (#35143)
|
2026-08-17 19:51:39 +08:00 |
|
Mick
|
e9ad8102a2
|
[diffusion] chore: reuse SRT SigLIP in Pi0.5 (#34992)
|
2026-08-17 19:33:36 +08:00 |
|
WenhaoZhang
|
f33b83b4cc
|
[diffusion] fix: fix h3 swap peft SwiGLU lora_B halves when loading FFN Lora (#34940)
|
2026-08-17 18:47:31 +08:00 |
|
Mohammad Miadh Angkad
|
82995a001b
|
Stabilize GB300 nightly tests (#35044)
|
2026-08-17 03:30:16 -07:00 |
|
 
|
744740dbea
|
[XPU] upgrade sglang xpu backend to PyTorch 2.13 (#31751)
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-08-17 18:29:15 +08:00 |
|
kangwangamd
|
c82e928fe5
|
[AMD] diffusion: normalize ModelOpt-FP8 weights to e4m3fnuz on gfx942 (#35111)
|
2026-08-17 02:35:10 -07:00 |
|
Alan Kao
|
056808723e
|
[AMD] Guard ROCm 7.0 build from using hipMemcpyBatchAsync (#35128)
|
2026-08-17 02:22:58 -07:00 |
|
 
|
92bce3d7bb
|
[AMD] [GLM5] fp8 MLA absorbed bmm for GLM-5.2 on gfx950 (#30519)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: sogalin_codegen <39478626+sogalin@users.noreply.github.com>
|
2026-08-17 02:15:16 -07:00 |
|
 
|
8cc112d486
|
[DSA] Skip indexer KV cache for skip-topk layers (#30531)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
|
2026-08-17 02:02:23 -07:00 |
|
YAMY
|
7c423cfd41
|
[PD] Avoid unused PREBUILT prompt tensor transfer (#35070)
|
2026-08-17 16:48:12 +08:00 |
|
Liangsheng Yin
|
711bdacb82
|
[Spec] Resolve shared-read ends from the backend declaration alone (#35059)
|
2026-08-17 01:35:29 -07:00 |
|
    
|
b83d507cd7
|
[NPU] Support DeepSeek-V4 DSpark and refactor DSV4 cache management (#33676)
Co-authored-by: JiaruiChang5268 <jc5268@columbia.edu>
Co-authored-by: Kelon <kelonlu@163.com>
Co-authored-by: unknown <z8ruev42yk@gmail.com>
Co-authored-by: Talantan1102 <545811257@qq.com>
Co-authored-by: Talantan1102 <44429302+Talantan1102@users.noreply.github.com>
|
2026-08-17 16:27:44 +08:00 |
|