This website requires JavaScript.
9dcaf6bfdf
rust server build release artifacts (#33096 )
Rain Jiang and Alex Nails
2026-07-31 12:36:25 -07:00
94743f934c
[Docs] Add DeepSeek-V4 Flash Official (0731) recipe (#33083 )
Xinyuan Tong and zijiexia
2026-08-01 03:14:53 +08:00
5df193b4ac
[Speculative Decoding] Fix GPT-OSS EAGLE3 hidden states (#32334 )
Po-Han Huang (NVIDIA)
2026-08-01 02:58:49 +08:00
4af8ddb576
support rust sglang server (#29799 )
Rain Jiang
2026-07-31 11:56:31 -07:00
77c77a3da8
feat(inkling): migrate short convs onto the ShortConv attention backend (#33023 )
Cheng Wan
2026-07-31 11:52:12 -07:00
d3222bcc3a
[unified-memory] Support fa3, the default MLA backend on pre-Blackwell hosts (#33046 )
Cheng Wan
2026-07-31 11:46:46 -07:00
26486a957d
Fix --hicache-size allocating ~2x host memory on hybrid Mamba (#32915 )
luchangli
2026-08-01 02:37:57 +08:00
89f4a80c1f
Support fastsafetensors no-GDS loading and page-cache release (#31859 )
Nan Jiang
2026-07-31 08:12:32 -07:00
5f9b0db18c
Fix async loading of RunAI-streamed tensors (#32896 )
Danila Shtan and Danila Shtan
2026-07-31 15:46:33 +02:00
690de097c4
[fix]reject media input for text-only models (#32914 )
luoroger37 and Xinyuan Tong
2026-07-31 20:43:25 +08:00
e3d4f48e55
[Fix] missing max_context_len on HybridAttnBackend (#32690 )
Peng Wu
2026-07-31 04:43:09 -07:00
754b692afc
[diffusion] optimization: support cuda-ipc zero-staging all-to-all for 2-rank Ulysses (#31854 )
Mick and Claude Fable 5
2026-07-31 19:35:48 +08:00
fd28242b68
[CI] Pin NCCL ports for GB300 PR tests (#33044 )
Baizhou Zhang
2026-07-31 02:28:37 -07:00
2573190b93
feat: support Kimi Linear PD disaggregation with DCP (#32837 )
Khoa Pham and Yangmin Li
2026-07-31 02:14:09 -07:00
33c27d8e7f
[unified-memory] Let Kimi-Linear use the paged MLA attention backends (#32972 )
Cheng Wan
2026-07-31 01:32:08 -07:00
937c77cf50
[Fix] Clear stale FlashInfer BF16 MoE index cache (#33016 )
Ziang Li
2026-07-31 00:35:15 -07:00
585a7d05e3
[Diffusion] Return scheduler sigmas snapshot in rollout dit_trajectory (#32683 )
Kangrui Du
2026-07-31 00:29:07 -07:00
0d6bef6b6d
[NPU] [DOC] renew triton-ascend installation guide location (#32986 )
amote-i
2026-07-31 14:46:04 +08:00
f94d2c5663
[Fix] Restore online MXFP8 quantization for linear layers (#32953 )
Brayden Zhong
2026-07-30 23:42:06 -07:00
5c6635d8f3
[Spec] Compact the target-verify mask when nothing reads it (#32920 )
Liangsheng Yin and Kaixi Matteo Chen
2026-07-30 23:13:21 -07:00
09193bf36f
[Fix]: render tool_reference schema regardless of tool_result part order (#32522 )
Shijin Zhang and Xinyuan Tong
2026-07-31 13:57:52 +08:00
e23ccb15f0
[unified-memory] Support MLA-hybrid-Mamba (Kimi-Linear) on the Triton backend (#32971 )
Cheng Wan
2026-07-30 22:10:34 -07:00
06ccaef24a
Fix silently wrong EPLB output with --moe-a2a-backend none (rank-invariant dispatch) (#32962 )
Cheng Wan
2026-07-30 22:10:04 -07:00
f3fd869494
[gdn] support replayssm with extra buffer (#32692 )
Qiaolin Yu
2026-07-30 21:34:37 -07:00
afeaeccfa2
perf(hisparse): eliminate redundant swap output fill (#32483 )
DAI0818 and Xiaoyu Zhang
2026-07-31 12:09:44 +08:00
425349b799
[Perf][DSA] Pass topk_length to flash_mla_sparse_fwd in the sparse attention path (#31128 )
zky
2026-07-31 11:25:59 +08:00
0aefba7283
fix(dsa): correct packed FlashInfer top-k and backend selection semantics (#32490 )
Ziang Li and Parth Chadha
2026-07-30 20:20:16 -07:00
c039e1a7ee
[NPU][DOC] Restructure ascend-npus docs into layered navigation (#32857 )
amote-i
2026-07-31 09:49:46 +08:00
3abbc565e4
[Docs] Add a Conventions section to the add-jit-kernel skill (#32956 )
DarkSharpness and Claude Fable 5
2026-07-31 09:11:49 +08:00
a149717308
feat: log multimodal encoder DP tradeoffs (#30903 )
Mick
2026-07-31 08:50:20 +08:00
55c1963df4
Remove unused draft-extend CUDA graph top-k (#31430 )
weireweire and weireweire
2026-07-31 08:45:01 +08:00
68d442945f
Flush dropped reasoning at stream end when stream_reasoning=False (#32225 )
Xinyuan Tong
2026-07-31 08:25:54 +08:00
48dbc24cbf
[Qwen3.5][MTP] Support FlashInfer CuTe DSL for online NVFP4 draft MoE (#31382 )
YAMY
2026-07-30 17:19:48 -07:00
a1c30701aa
Integrate pplx a2a backend (#30756 )
Trang Do and Cheng Wan
2026-07-31 05:33:19 +07:00
3a53c26c27
[CI] Fix MoE compile and DSA indexer regressions (#32937 )
Mohammad Miadh Angkad
2026-07-31 06:21:51 +08:00
85f9998524
docs: sync LMSYS SGLang blog cards (#32838 )
sglang-bot and sglang-bot
2026-07-30 14:56:22 -07:00
a6221d776f
feat: Support nvidia/MiniMax-M3-NVFP4 (#31989 )
Trevor Morris
2026-07-30 14:32:03 -07:00
c4af6cf263
Qwen3.5-MoE: support modelopt_fp4 checkpoints that quantize attention (+ load baked FP8 KV scales) (#31220 )
Henning Thieß
2026-07-30 23:30:26 +02:00
5339450ed4
Support SGLANG_SIMULATE_ACC_LEN for DFLASH (#32595 )
saatwiknagpal
2026-07-30 14:10:06 -07:00
3312645a30
wire the rust server modules into lib, runtime, and tokenizer manager (#32877 )
Rain Jiang
2026-07-30 12:46:19 -07:00
047635ee35
add the rust server native api handlers and runtime threads (#32876 )
Rain Jiang
2026-07-30 12:46:19 -07:00
30643f88bc
add the rust server api frame codec and http server entry (#32875 )
Rain Jiang
2026-07-30 12:46:19 -07:00
4facc0e18a
add the rust server ingress tests, guard, and submit modules (#32874 )
Rain Jiang
2026-07-30 12:46:18 -07:00
e2c65af229
add the rust server ingress request validation and api server common types (#32873 )
Rain Jiang
2026-07-30 12:46:17 -07:00
922d6e5542
add the rust server tokenizer, detokenizer, and egress modules (#32872 )
Rain Jiang
2026-07-30 12:46:17 -07:00
35f2e6ab58
update Cargo.lock for the rust sglang-server dependencies (#32871 )
Rain Jiang
2026-07-30 12:32:27 -07:00
04edadb34d
Add Inkling-Small cookbook (#32951 )
2026-07-31 01:59:11 +08:00
b61cb5f9de
Fix DeepSeek V4 loading with RunAI Model Streamer. (#30240 )
Broduker
2026-07-30 23:03:34 +08:00
c5bd3d7dce
[diffusion][benchmark] Add reproducible request-manifest offline benchmark (#32917 )
Xiaoyu Zhang
2026-07-30 22:11:18 +08:00
7784ac8f91
[diffusion][docs] Fix Cosmos3 model sizes (#32916 )
Xiaoyu Zhang
2026-07-30 22:10:23 +08:00
2e9c82b359
[Kernel] Remove unreachable AOT headers (#32842 )
Xiaoyu Zhang
2026-07-30 22:08:57 +08:00
48c1b37a33
[AMD] Update ROCm AITER pin to d9e5ef7 (#32939 )
Bingxu Chen and Cursor Agent
2026-07-30 22:05:03 +08:00
b78d3999b5
【NPU】fix decode MTP + eagle shape error (#32791 )
cen121212
2026-07-30 21:34:16 +08:00
b129e8a299
[diffusion] docs: surface diffusion AR and PE guides (#32932 )
Mick
2026-07-30 21:05:54 +08:00
db3da62333
[diffusion] feat: unify encoder folding and batch data-parallel encoding (#30211 )
Mick
2026-07-30 20:15:22 +08:00
1f04eaab6a
Fix LFM 2 tool parser. (#27614 )
Yi (Vincent) Zhong
2026-07-30 03:15:08 -07:00
fd86795107
[AMD] MiniMax-M3: opt-in custom/quick all-reduce on ROCm (#32230 )
YC Yen-Ching Tseng
2026-07-30 17:56:07 +08:00
9f56553408
[Perf] Fast-path chain-style draft token organization in multi-layer EAGLE (#32887 )
Liangsheng Yin
2026-07-30 02:55:21 -07:00
04d6fb4d6c
[AMD] Minimax-M3 : unblock mxfp8 block convert on gfx950 (#32036 )
YC Yen-Ching Tseng
2026-07-30 17:43:27 +08:00
4ba7d5ad93
[BugFix][EPD] Early-release mooncake GPU embeddings; fix gpu_id via scheduler.ps (#31591 )
Zheng Wengang
2026-07-30 02:36:55 -07:00
6ab3231b97
[Perf] Skip the target-verify tree mask fill when the backend never reads it (#32886 )
Liangsheng Yin and Kaixi
2026-07-30 02:32:38 -07:00
4b52758c76
[AMD] Skip test_update_weights_from_disk on ROCm pending reload fix (#31924 ) (#31925 )
kangwangamd
2026-07-30 17:32:09 +08:00
fc007e1f00
Add SM90 FP8 MegaMoE support for DeepSeek-V4 (#29016 )
Ding Yin and yinding
2026-07-30 16:48:10 +08:00
f46d5f25b4
[4/N][CP] Support interleave strategy for cp v2 (#30482 )
2026-07-30 16:45:32 +08:00
c192145830
[Kernel] Fuse KV-cache writes for asymmetric K/V (head_dim != v_head_dim) (#32813 )
Liangsheng Yin
2026-07-30 00:26:10 -07:00
2625fdfe6b
[Fix] Count multi-layer draft-extend replays in the fwd-occupancy device timer (#32867 )
Liangsheng Yin
2026-07-30 00:21:34 -07:00
92b3a51ba6
[LoRA] Fix Marlin MoE kernel import (#32884 )
Yanbin Jiang
2026-07-30 00:17:29 -07:00
e4a40a71f8
[DSA] Q8KV8 FP8 Sparse Prefill on GLM-5.2 & DeepSeek-V3.2: Q8-Path & Shared-Path Optimizations (#31888 )
Ho-Ren (Jack) Chuang and Claude Fable 5
2026-07-30 00:15:11 -07:00
4f51dad1da
fix: prevent ReqTimeStats from being dropped during IPC serialization (#31339 )
Jinwoo Jeong
2026-07-30 15:42:38 +09:00
f6ff5e8bb0
[PD] Handle abort requests in PP mode (#32797 )
Shangming Cai
2026-07-30 14:39:44 +08:00
36afd442c7
[DLLM] vectorized joint/low-confidence decoding and skip redundant attn init (#21094 )
wenxuewuhd and ronnie_zheng
2026-07-30 14:13:14 +08:00
07a087bf45
Fix Inkling tool-call parsing recovery, content handling, and streaming (#32861 )
Ke Bao
2026-07-30 14:11:09 +08:00
f4e0ac382e
[misc] Remove unused multi_layer_draft_forward_cg module (#32881 )
Liangsheng Yin
2026-07-29 21:17:16 -07:00
3d6e1e6f81
[AMD] Revert ROCm AITER pin to 9127c94 (#32879 )
Bingxu Chen
2026-07-30 11:53:49 +08:00
ed361ae7f0
Fix attention backends for models with per-layer head counts (num_attention_heads_per_layer) (#32625 )
Jimmy Shong
2026-07-29 20:03:00 -07:00
22faf9fef8
embedding: centralize capabilities and complete OpenAI compatibility (#32481 )
Mick
2026-07-30 10:28:52 +08:00
313a518bee
[Spec] Emit step trace span for multi-layer draft-extend graph replays (#32850 )
Liangsheng Yin
2026-07-29 19:22:58 -07:00
2aa86e9130
[diffusion] docs: add diffusion cookbook model tags (#32836 )
Mick
2026-07-30 10:03:55 +08:00
62dfaaa0e0
[Nemotron] Fix decode track-save reading the stale tail of the CUDA-graph track buffer (#32555 )
Sam Shleifer and Claude Fable 5
2026-07-29 22:03:44 -04:00
20d5b91e5b
[NPU] [DOC] update feature name to follow the code changement (#32749 )
amote-i
2026-07-30 09:54:08 +08:00
d7a4c830e5
sglang rust server tokenizer manager, ring and runtime (#32358 )
Rain Jiang
2026-07-29 18:42:12 -07:00
8fbf960980
[MLX] Size request capacity by attention DP (#32115 )
Xuanyi Li and R0CKSTAR
2026-07-29 18:18:22 -07:00
1d9c292547
[Kernel] Add inventory guards and clean benchmark layout (#32788 )
Xiaoyu Zhang
2026-07-30 09:03:24 +08:00
5efbb18a6f
[docs] Rotate popular models on the landing pages, lead the Cookbook nav with Kimi (#32835 )
zijiexia and Claude Opus 5
2026-07-29 17:46:50 -07:00
3c9efaf3e1
[docs] Kimi-K3: widen the H200 High-Throughput recipe to 4x8 TP32/EP32 (#32834 )
zijiexia and Claude Opus 5
2026-07-29 17:18:56 -07:00
6e48c13497
sglang rust server egress message (#32342 )
Rain Jiang
2026-07-29 17:16:06 -07:00
a55e1764a2
Enable GPT-OSS FlashInfer MXFP4 on SM120 (#32668 )
Mohammad Miadh Angkad
2026-07-30 08:04:23 +08:00
e5c46ff07d
[Fix] Route asymmetric-KV models to fa4 on SM100 and pin MiMoV2 FP8 MoE to flashinfer_trtllm (#32818 )
Liangsheng Yin
2026-07-29 16:37:43 -07:00
3c1717d9b6
Follow up on #30157 post-merge review (#32672 )
cctry
2026-07-29 15:03:59 -07:00
8fc54d46ef
Fix MoE reduce-scatterv eligibility check (#32663 )
Sam (Kesen Li)
2026-07-30 05:58:55 +08:00
d24de56995
sglang rust server sampling message (#32343 )
Rain Jiang
2026-07-29 14:55:06 -07:00
ffd4705baa
fix(reasoning): honor Poolside template thinking defaults (#32540 )
Willow Lopez and Jiminator
2026-07-30 05:02:19 +08:00
d0e69d3881
[feat] Optional base64 encoding for the flat prompt top logprob arrays (#31960 )
Sam Shleifer
2026-07-29 15:15:56 -04:00
e4f7f7b380
fix(qwen3.5): restrict MoE weights to local PP layers (#32022 )
2026-07-30 03:01:09 +08:00
f73d6f2789
Update Inkling cookbook install command (#32799 )
Ke Bao
2026-07-30 02:23:23 +08:00
2ca2ca753a
sglang rust server request message (#32242 )
Rain Jiang
2026-07-29 10:54:56 -07:00
eefb434d17
[PD+PP] Honor PP consensus for bootstrap and prealloc (#31869 )
ziang663 and Chao Shi
2026-07-30 01:19:17 +08:00
62d0f81f16
[2/N] elastic-ep: Enable EPLB after scale-up (#30553 )
Yoray Zack
2026-07-29 20:06:56 +03:00
d19999b755
[LFM2] Wire Lfm2MoeForCausalLM into the LFM2 serving override tables (#30780 )
Piotr Mazurek
2026-07-29 19:02:14 +02:00
f69af7b7ad
[Bugfix] compressed-tensors: mixed-precision checkpoints silently load unquantized (#32736 )
Hert4 and Mohammad Miadh Angkad
2026-07-29 23:55:33 +07:00