Commit Graph
15878 Commits
Author SHA1 Message Date
amote-i c039e1a7ee [NPU][DOC] Restructure ascend-npus docs into layered navigation (#32857) 2026-07-31 09:49:46 +08:00
DarkSharpnessandClaude Fable 5 3abbc565e4 [Docs] Add a Conventions section to the add-jit-kernel skill (#32956)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 09:11:49 +08:00
Mick a149717308 feat: log multimodal encoder DP tradeoffs (#30903) 2026-07-31 08:50:20 +08:00
weireweireandweireweire 55c1963df4 Remove unused draft-extend CUDA graph top-k (#31430)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
2026-07-30 17:45:01 -07:00
Xinyuan Tong 68d442945f Flush dropped reasoning at stream end when stream_reasoning=False (#32225) 2026-07-31 08:25:54 +08:00
YAMY 48dbc24cbf [Qwen3.5][MTP] Support FlashInfer CuTe DSL for online NVFP4 draft MoE (#31382) 2026-07-30 17:19:48 -07:00
Trang DoandCheng Wan a1c30701aa Integrate pplx a2a backend (#30756)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2026-07-30 15:33:19 -07:00
Mohammad Miadh Angkad 3a53c26c27 [CI] Fix MoE compile and DSA indexer regressions (#32937) 2026-07-30 15:21:51 -07:00
sglang-botandsglang-bot 85f9998524 docs: sync LMSYS SGLang blog cards (#32838)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-30 21:56:22 +00:00
Trevor Morris a6221d776f feat: Support nvidia/MiniMax-M3-NVFP4 (#31989) 2026-07-30 14:32:03 -07:00
Henning Thieß c4af6cf263 Qwen3.5-MoE: support modelopt_fp4 checkpoints that quantize attention (+ load baked FP8 KV scales) (#31220) 2026-07-30 14:30:26 -07:00
saatwiknagpal 5339450ed4 Support SGLANG_SIMULATE_ACC_LEN for DFLASH (#32595) 2026-07-30 14:10:06 -07:00
Rain Jiang 3312645a30 wire the rust server modules into lib, runtime, and tokenizer manager (#32877) 2026-07-30 12:46:19 -07:00
Rain Jiang 047635ee35 add the rust server native api handlers and runtime threads (#32876) 2026-07-30 12:46:19 -07:00
Rain Jiang 30643f88bc add the rust server api frame codec and http server entry (#32875) 2026-07-30 12:46:19 -07:00
Rain Jiang 4facc0e18a add the rust server ingress tests, guard, and submit modules (#32874) 2026-07-30 12:46:18 -07:00
Rain Jiang e2c65af229 add the rust server ingress request validation and api server common types (#32873) 2026-07-30 12:46:17 -07:00
Rain Jiang 922d6e5542 add the rust server tokenizer, detokenizer, and egress modules (#32872) 2026-07-30 12:46:17 -07:00
Rain Jiang 35f2e6ab58 update Cargo.lock for the rust sglang-server dependencies (#32871) 2026-07-30 12:32:27 -07:00
04edadb34d Add Inkling-Small cookbook (#32951)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
2026-07-31 01:59:11 +08:00
Broduker b61cb5f9de Fix DeepSeek V4 loading with RunAI Model Streamer. (#30240) 2026-07-30 23:03:34 +08:00
Xiaoyu Zhang c5bd3d7dce [diffusion][benchmark] Add reproducible request-manifest offline benchmark (#32917) 2026-07-30 22:11:18 +08:00
Xiaoyu Zhang 7784ac8f91 [diffusion][docs] Fix Cosmos3 model sizes (#32916) 2026-07-30 22:10:23 +08:00
Xiaoyu Zhang 2e9c82b359 [Kernel] Remove unreachable AOT headers (#32842) 2026-07-30 22:08:57 +08:00
Bingxu ChenandCursor Agent 48c1b37a33 [AMD] Update ROCm AITER pin to d9e5ef7 (#32939)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-07-30 22:05:03 +08:00
cen121212 b78d3999b5 【NPU】fix decode MTP + eagle shape error (#32791) 2026-07-30 21:34:16 +08:00
Mick b129e8a299 [diffusion] docs: surface diffusion AR and PE guides (#32932) 2026-07-30 21:05:54 +08:00
Mick db3da62333 [diffusion] feat: unify encoder folding and batch data-parallel encoding (#30211) 2026-07-30 20:15:22 +08:00
Yi (Vincent) Zhong 1f04eaab6a Fix LFM 2 tool parser. (#27614) 2026-07-30 10:15:08 +00:00
YC Yen-Ching Tseng fd86795107 [AMD] MiniMax-M3: opt-in custom/quick all-reduce on ROCm (#32230) 2026-07-30 02:56:07 -07:00
Liangsheng Yin 9f56553408 [Perf] Fast-path chain-style draft token organization in multi-layer EAGLE (#32887) 2026-07-30 02:55:21 -07:00
YC Yen-Ching Tseng 04d6fb4d6c [AMD] Minimax-M3 : unblock mxfp8 block convert on gfx950 (#32036) 2026-07-30 02:43:27 -07:00
Zheng Wengang 4ba7d5ad93 [BugFix][EPD] Early-release mooncake GPU embeddings; fix gpu_id via scheduler.ps (#31591) 2026-07-30 17:36:55 +08:00
Liangsheng YinandKaixi 6ab3231b97 [Perf] Skip the target-verify tree mask fill when the backend never reads it (#32886)
Co-authored-by: Kaixi <kaiximatteoc@nvidia.com>
2026-07-30 02:32:38 -07:00
kangwangamd 4b52758c76 [AMD] Skip test_update_weights_from_disk on ROCm pending reload fix (#31924) (#31925) 2026-07-30 02:32:09 -07:00
Ding Yinandyinding fc007e1f00 Add SM90 FP8 MegaMoE support for DeepSeek-V4 (#29016)
Co-authored-by: yinding <yinding@bytedance.com>
2026-07-30 01:48:10 -07:00
f46d5f25b4 [4/N][CP] Support interleave strategy for cp v2 (#30482)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-07-30 01:45:32 -07:00
Liangsheng Yin c192145830 [Kernel] Fuse KV-cache writes for asymmetric K/V (head_dim != v_head_dim) (#32813) 2026-07-30 00:26:10 -07:00
Liangsheng Yin 2625fdfe6b [Fix] Count multi-layer draft-extend replays in the fwd-occupancy device timer (#32867) 2026-07-30 00:21:34 -07:00
Yanbin Jiang 92b3a51ba6 [LoRA] Fix Marlin MoE kernel import (#32884) 2026-07-30 15:17:29 +08:00
Ho-Ren (Jack) ChuangandClaude Fable 5 e4a40a71f8 [DSA] Q8KV8 FP8 Sparse Prefill on GLM-5.2 & DeepSeek-V3.2: Q8-Path & Shared-Path Optimizations (#31888)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 15:15:11 +08:00
Jinwoo Jeong 4f51dad1da fix: prevent ReqTimeStats from being dropped during IPC serialization (#31339) 2026-07-30 14:42:38 +08:00
Shangming Cai f6ff5e8bb0 [PD] Handle abort requests in PP mode (#32797) 2026-07-30 14:39:44 +08:00
wenxuewuhdandronnie_zheng 36afd442c7 [DLLM] vectorized joint/low-confidence decoding and skip redundant attn init (#21094)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-30 09:13:14 +03:00
Ke Bao 07a087bf45 Fix Inkling tool-call parsing recovery, content handling, and streaming (#32861) 2026-07-30 14:11:09 +08:00
Liangsheng Yin f4e0ac382e [misc] Remove unused multi_layer_draft_forward_cg module (#32881) 2026-07-29 21:17:16 -07:00
Bingxu Chen 3d6e1e6f81 [AMD] Revert ROCm AITER pin to 9127c94 (#32879) 2026-07-30 11:53:49 +08:00
Jimmy Shong ed361ae7f0 Fix attention backends for models with per-layer head counts (num_attention_heads_per_layer) (#32625) 2026-07-29 20:03:00 -07:00
Mick 22faf9fef8 embedding: centralize capabilities and complete OpenAI compatibility (#32481) 2026-07-30 10:28:52 +08:00
Liangsheng Yin 313a518bee [Spec] Emit step trace span for multi-layer draft-extend graph replays (#32850) 2026-07-29 19:22:58 -07:00