YAMY
|
48dbc24cbf
|
[Qwen3.5][MTP] Support FlashInfer CuTe DSL for online NVFP4 draft MoE (#31382)
|
2026-07-30 17:19:48 -07:00 |
|
 Trang DoandCheng Wan
|
a1c30701aa
|
Integrate pplx a2a backend (#30756)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-07-30 15:33:19 -07:00 |
|
Mohammad Miadh Angkad
|
3a53c26c27
|
[CI] Fix MoE compile and DSA indexer regressions (#32937)
|
2026-07-30 15:21:51 -07:00 |
|
 sglang-botandsglang-bot
|
85f9998524
|
docs: sync LMSYS SGLang blog cards (#32838)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-30 21:56:22 +00:00 |
|
Trevor Morris
|
a6221d776f
|
feat: Support nvidia/MiniMax-M3-NVFP4 (#31989)
|
2026-07-30 14:32:03 -07:00 |
|
Henning Thieß
|
c4af6cf263
|
Qwen3.5-MoE: support modelopt_fp4 checkpoints that quantize attention (+ load baked FP8 KV scales) (#31220)
|
2026-07-30 14:30:26 -07:00 |
|
saatwiknagpal
|
5339450ed4
|
Support SGLANG_SIMULATE_ACC_LEN for DFLASH (#32595)
|
2026-07-30 14:10:06 -07:00 |
|
Rain Jiang
|
3312645a30
|
wire the rust server modules into lib, runtime, and tokenizer manager (#32877)
|
2026-07-30 12:46:19 -07:00 |
|
Rain Jiang
|
047635ee35
|
add the rust server native api handlers and runtime threads (#32876)
|
2026-07-30 12:46:19 -07:00 |
|
Rain Jiang
|
30643f88bc
|
add the rust server api frame codec and http server entry (#32875)
|
2026-07-30 12:46:19 -07:00 |
|
Rain Jiang
|
4facc0e18a
|
add the rust server ingress tests, guard, and submit modules (#32874)
|
2026-07-30 12:46:18 -07:00 |
|
Rain Jiang
|
e2c65af229
|
add the rust server ingress request validation and api server common types (#32873)
|
2026-07-30 12:46:17 -07:00 |
|
Rain Jiang
|
922d6e5542
|
add the rust server tokenizer, detokenizer, and egress modules (#32872)
|
2026-07-30 12:46:17 -07:00 |
|
Rain Jiang
|
35f2e6ab58
|
update Cargo.lock for the rust sglang-server dependencies (#32871)
|
2026-07-30 12:32:27 -07:00 |
|
 
|
04edadb34d
|
Add Inkling-Small cookbook (#32951)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
|
2026-07-31 01:59:11 +08:00 |
|
Broduker
|
b61cb5f9de
|
Fix DeepSeek V4 loading with RunAI Model Streamer. (#30240)
|
2026-07-30 23:03:34 +08:00 |
|
Xiaoyu Zhang
|
c5bd3d7dce
|
[diffusion][benchmark] Add reproducible request-manifest offline benchmark (#32917)
|
2026-07-30 22:11:18 +08:00 |
|
Xiaoyu Zhang
|
7784ac8f91
|
[diffusion][docs] Fix Cosmos3 model sizes (#32916)
|
2026-07-30 22:10:23 +08:00 |
|
Xiaoyu Zhang
|
2e9c82b359
|
[Kernel] Remove unreachable AOT headers (#32842)
|
2026-07-30 22:08:57 +08:00 |
|
 Bingxu ChenandCursor Agent
|
48c1b37a33
|
[AMD] Update ROCm AITER pin to d9e5ef7 (#32939)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
|
2026-07-30 22:05:03 +08:00 |
|
cen121212
|
b78d3999b5
|
【NPU】fix decode MTP + eagle shape error (#32791)
|
2026-07-30 21:34:16 +08:00 |
|
Mick
|
b129e8a299
|
[diffusion] docs: surface diffusion AR and PE guides (#32932)
|
2026-07-30 21:05:54 +08:00 |
|
Mick
|
db3da62333
|
[diffusion] feat: unify encoder folding and batch data-parallel encoding (#30211)
|
2026-07-30 20:15:22 +08:00 |
|
Yi (Vincent) Zhong
|
1f04eaab6a
|
Fix LFM 2 tool parser. (#27614)
|
2026-07-30 10:15:08 +00:00 |
|
YC Yen-Ching Tseng
|
fd86795107
|
[AMD] MiniMax-M3: opt-in custom/quick all-reduce on ROCm (#32230)
|
2026-07-30 02:56:07 -07:00 |
|
Liangsheng Yin
|
9f56553408
|
[Perf] Fast-path chain-style draft token organization in multi-layer EAGLE (#32887)
|
2026-07-30 02:55:21 -07:00 |
|
YC Yen-Ching Tseng
|
04d6fb4d6c
|
[AMD] Minimax-M3 : unblock mxfp8 block convert on gfx950 (#32036)
|
2026-07-30 02:43:27 -07:00 |
|
Zheng Wengang
|
4ba7d5ad93
|
[BugFix][EPD] Early-release mooncake GPU embeddings; fix gpu_id via scheduler.ps (#31591)
|
2026-07-30 17:36:55 +08:00 |
|
 Liangsheng YinandKaixi
|
6ab3231b97
|
[Perf] Skip the target-verify tree mask fill when the backend never reads it (#32886)
Co-authored-by: Kaixi <kaiximatteoc@nvidia.com>
|
2026-07-30 02:32:38 -07:00 |
|
kangwangamd
|
4b52758c76
|
[AMD] Skip test_update_weights_from_disk on ROCm pending reload fix (#31924) (#31925)
|
2026-07-30 02:32:09 -07:00 |
|
 Ding Yinandyinding
|
fc007e1f00
|
Add SM90 FP8 MegaMoE support for DeepSeek-V4 (#29016)
Co-authored-by: yinding <yinding@bytedance.com>
|
2026-07-30 01:48:10 -07:00 |
|
 
|
f46d5f25b4
|
[4/N][CP] Support interleave strategy for cp v2 (#30482)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-07-30 01:45:32 -07:00 |
|
Liangsheng Yin
|
c192145830
|
[Kernel] Fuse KV-cache writes for asymmetric K/V (head_dim != v_head_dim) (#32813)
|
2026-07-30 00:26:10 -07:00 |
|
Liangsheng Yin
|
2625fdfe6b
|
[Fix] Count multi-layer draft-extend replays in the fwd-occupancy device timer (#32867)
|
2026-07-30 00:21:34 -07:00 |
|
Yanbin Jiang
|
92b3a51ba6
|
[LoRA] Fix Marlin MoE kernel import (#32884)
|
2026-07-30 15:17:29 +08:00 |
|
 Ho-Ren (Jack) ChuangandClaude Fable 5
|
e4a40a71f8
|
[DSA] Q8KV8 FP8 Sparse Prefill on GLM-5.2 & DeepSeek-V3.2: Q8-Path & Shared-Path Optimizations (#31888)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-30 15:15:11 +08:00 |
|
Jinwoo Jeong
|
4f51dad1da
|
fix: prevent ReqTimeStats from being dropped during IPC serialization (#31339)
|
2026-07-30 14:42:38 +08:00 |
|
Shangming Cai
|
f6ff5e8bb0
|
[PD] Handle abort requests in PP mode (#32797)
|
2026-07-30 14:39:44 +08:00 |
|
 wenxuewuhdandronnie_zheng
|
36afd442c7
|
[DLLM] vectorized joint/low-confidence decoding and skip redundant attn init (#21094)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-30 09:13:14 +03:00 |
|
Ke Bao
|
07a087bf45
|
Fix Inkling tool-call parsing recovery, content handling, and streaming (#32861)
|
2026-07-30 14:11:09 +08:00 |
|
Liangsheng Yin
|
f4e0ac382e
|
[misc] Remove unused multi_layer_draft_forward_cg module (#32881)
|
2026-07-29 21:17:16 -07:00 |
|
Bingxu Chen
|
3d6e1e6f81
|
[AMD] Revert ROCm AITER pin to 9127c94 (#32879)
|
2026-07-30 11:53:49 +08:00 |
|
Jimmy Shong
|
ed361ae7f0
|
Fix attention backends for models with per-layer head counts (num_attention_heads_per_layer) (#32625)
|
2026-07-29 20:03:00 -07:00 |
|
Mick
|
22faf9fef8
|
embedding: centralize capabilities and complete OpenAI compatibility (#32481)
|
2026-07-30 10:28:52 +08:00 |
|
Liangsheng Yin
|
313a518bee
|
[Spec] Emit step trace span for multi-layer draft-extend graph replays (#32850)
|
2026-07-29 19:22:58 -07:00 |
|
Mick
|
2aa86e9130
|
[diffusion] docs: add diffusion cookbook model tags (#32836)
|
2026-07-30 10:03:55 +08:00 |
|
 Sam ShleiferandClaude Fable 5
|
62dfaaa0e0
|
[Nemotron] Fix decode track-save reading the stale tail of the CUDA-graph track buffer (#32555)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-29 19:03:44 -07:00 |
|
amote-i
|
20d5b91e5b
|
[NPU] [DOC] update feature name to follow the code changement (#32749)
|
2026-07-30 09:54:08 +08:00 |
|
Rain Jiang
|
d7a4c830e5
|
sglang rust server tokenizer manager, ring and runtime (#32358)
|
2026-07-29 18:42:12 -07:00 |
|
 Xuanyi LiandR0CKSTAR
|
8fbf960980
|
[MLX] Size request capacity by attention DP (#32115)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-07-29 18:18:22 -07:00 |
|