This website requires JavaScript.
ff228d11fe
[HiCache] Release buffer prefetch anchor locks during storage cleanup (#38483 )
Shuwen Wang
2026-09-14 04:47:21 +08:00
ddc1df1203
fix(router): simplify bucket context limit check (#39259 )
ishandhanani
2026-09-13 12:27:59 -07:00
6220f45d8e
[PD] Add /v1/responses support to the HTTP PD router (#36141 )
2026-09-13 23:48:13 +09:00
7f1f8c706a
[mem_cache][10/N] refactor: drop the redundant _component suffix in unified_cache/components (#35644 )
Shuwen Wang and Claude Opus 5
2026-09-13 21:38:20 +08:00
14a131ad5b
[PD][OpenAI] Gate /v1/responses persistence behind --enable-response-store, default off (#39122 )
Shangming Cai and Xinyuan Tong
2026-09-13 21:20:47 +08:00
f9fca05803
[diffusion] fix: make the sp sequence gather pass contiguous shards (#39291 )
2026-09-13 20:30:52 +08:00
9ffe548738
[diffusion] fix: don't route an unreadable checkpoint into the native fallback (#39292 )
2026-09-13 20:30:24 +08:00
14b647cf27
chore: add HiSparse coordinator and allocator code owners (#38682 )
Shuwen Wang
2026-09-13 18:42:33 +08:00
d24eacaecd
[Test] Add unit test for muse_glimmer_format (#37015 )
Hritik Raj
2026-09-13 16:10:02 +05:30
7078e5ffbc
[HiCache] Publish a host store event for storage-prefetch refills (#38486 )
Jundong Liu and Shuwen Wang
2026-09-13 18:29:14 +08:00
a7cf4a6fbc
[Unified Cache][7/N] Support MTP, EAGLE, and DSpark draft KV caches in the external linker (#37914 )
huangtingwei and hzh0425
2026-09-13 18:22:06 +08:00
cebca698e2
[Qwen3.8] Enable NVIDIA NVFP4 on DGX Spark with file-backed PLE and PDL router fix (#39126 )
2026-09-13 01:23:41 -07:00
d6fabb74b4
[AMD][Fix] Fix aiter bpreshuffle GEMM for output sizes it cannot dispatch for qwen3.5 mxfp-attn-fp8-v2 TP4 (#37564 )
jacky.cheng and HAI
2026-09-13 14:11:44 +08:00
d2f054d916
[diffusion] CI: make diffusion GT generation usable without the publish token, and cover the 5090 lane (#39255 )
2026-09-13 13:59:43 +08:00
492f346d88
[diffusion] CI: set an explicit x264 preset for video output (#38657 )
2026-09-13 13:47:41 +08:00
6d6d42d1c0
[Fix] Preserve GLM tool argument types across JSON Schema unions (#39136 )
William Hu and Xinyuan Tong
2026-09-13 01:47:13 -04:00
34b2904741
[Session] Work with PD and Fix empty continuations (#39038 )
2026-09-12 22:24:32 -07:00
24b6c1c7f5
Fix multimodal embedding cache retaining full batches through views (#39120 )
cctry
2026-09-12 21:37:14 -07:00
7e3d18bbcc
Add a provider hook for prefill-buffer ceilings (#39182 )
cctry and cctry
2026-09-12 21:37:10 -07:00
7763f666f3
Fix DeepGEMM release MegaMoE validation without RDMA (#39245 )
Baizhou Zhang
2026-09-12 20:27:29 -07:00
ec5fba5777
[AMD] Add dspark config and agentic workload section for deepseek-v4 model (#39252 )
Thomas Wang
2026-09-13 10:32:01 +08:00
206034e520
Keep graph-pool borrows on their allocation stream (#39180 )
cctry and cctry
2026-09-12 19:26:09 -07:00
18cc55dc0b
Expose a capacity check for graph-pool borrows (#39178 )
cctry and cctry
2026-09-12 19:24:57 -07:00
23bc4c6ed9
[Diffusion] Return Qwen-Image-Layered outputs and preserve CFG2 rounding (#38549 )
Xiaoyu Zhang and Mick Qian
2026-09-13 09:15:47 +08:00
3e035a3513
[Diffusion] Optimize Qwen-Image-Edit attention on Hopper (#38584 )
Xiaoyu Zhang and Mick Qian
2026-09-13 09:15:01 +08:00
fa663e7297
config: delete the redundant full stamp in initialize_model_parallel (#39202 )
Cheng Wan
2026-09-12 17:27:21 -07:00
fa260f26da
config: delete dead ensure_model_parallel_initialized (#39137 )
Cheng Wan
2026-09-12 17:27:12 -07:00
6804eeaabe
config: an out-of-tree replacement point for every resolution-pipeline step (#39134 )
Cheng Wan
2026-09-12 17:26:54 -07:00
a8b5616303
Fix DeepGEMM release dependencies and bound GPU validation (#39241 )
Baizhou Zhang
2026-09-12 17:00:58 -07:00
6657f7d844
[MoE][ROCm] Admit the unified Triton router on ROCm, including single-group routing (#38328 )
2026-09-13 07:21:28 +08:00
a66451c058
[GLM-5.3 Flash] Restore and enable KPool metadata fusion (#38845 )
2026-09-12 16:01:11 -07:00
288627e400
[AMD] Document GLM-5.2 MXFP4 recipe update on MI355X (#39230 )
Zhang, Jiejing
2026-09-12 15:57:15 -07:00
21289cfd50
[AMD] Fix Dspark accept length and reduce host bubble on DSV4 (#39116 )
Xinyi Song
2026-09-12 15:21:43 -07:00
b5a2aebc7e
[Docs] GLM-5.3-Flash cookbook: fixed MTP 5/1/6, EP1 + flashinfer_trtllm on Blackwell (#39213 )
Xinyuan Tong
2026-09-13 03:56:57 +08:00
7ae4af8187
Scope graph-pool borrowing to the runtime and reduce fragmentation (#39177 )
cctry and cctry
2026-09-12 11:45:17 -07:00
2784a86062
Reuse live CUDA graph executables during dedup registration (#39176 )
cctry and cctry
2026-09-12 11:44:17 -07:00
25a5641cf2
[Session + MM] Fix text positions in session continuations (#39144 )
2026-09-12 10:31:55 -07:00
ae1acf822d
[GraniteMoE] Load split per-expert quantized MoE weights (#37679 )
2026-09-12 11:10:20 -04:00
7b89b95168
[diffusion] feat: pick the attention backend by measuring it (#38689 )
2026-09-12 21:52:16 +08:00
dc3171c322
[diffusion] CI: remove mova ulysses two-gpu CI case (#39097 )
Mick and Mick Qian
2026-09-12 21:48:04 +08:00
925e684a88
[Feat][Responses API] Support custom tools, encrypted reasoning replay, developer tier and model validation (#38690 )
2026-09-12 21:29:12 +08:00
6dc7b3421b
Inference Support Mamba 2 and 1 (#34556 )
desmond-intel
2026-09-12 18:21:18 +05:30
fd32226706
Fix gpt-oss RunAI streamer weight ownership (#38908 )
Kunal and Claude Fable 5.1
2026-09-12 08:21:59 -04:00
bd45cd50ca
[Unified Tree] Port SWA Branching-Point Caching to the Rust TreeCore (#37584 )
Shuwen Wang
2026-09-12 18:37:57 +08:00
0b415fa573
[HiCache][LoRA] Isolate storage pages by extra key (#38577 )
Yanbin Jiang and Shuwen Wang
2026-09-12 03:18:14 -07:00
b9cb96496d
[Unified Tree] Preserve aux LRU recency when splitting nodes (#38482 )
Shuwen Wang
2026-09-12 18:17:38 +08:00
6953dae005
[Qwen 3.8 Next] Remove unused tokenwise QSA implementation and tests (#38960 )
Qiaolin Yu
2026-09-12 01:32:01 -07:00
7bc4eb3740
[AMD] Update MI355X MXFP4 HiCache defaults and quick-reduce quantization for Qwen3.5 cookbook (#39104 )
ChangLiu0709
2026-09-12 08:59:32 +01:00
7c195b9151
[AMD] Fix Quark load of MiniMax-M3 MXFP4 index_qkv_proj (#37254 )
Tuan Nguyen Gia
2026-09-12 10:50:41 +03:00
bf3305b65e
[diffusion] feat: read mapped layers directly when the host cannot cache them (#39022 )
Mick and Claude Opus 5
2026-09-12 15:48:03 +08:00
0a57403468
[Perf] Optimize Qwen3-VL unique-image serving on H100 (#36411 )
Xiaoyu Zhang and Cursor
2026-09-12 15:16:59 +08:00
6ba96d329f
[DCP] Resolve --dcp-comm-backend to fi_a2a/a2a by default for every model (#39165 )
Khoa Pham and Claude Fable 5.1
2026-09-11 23:48:42 -07:00
981b947568
[Fix] Seed raw tokenizer_path for smg-grpc-servicer in gRPC mode (#39105 )
Shangming Cai and Claude Fable 5
2026-09-12 14:41:15 +08:00
56a4f47ca6
[CI] Fix base-a wait for skipped CPU matrix (#39195 )
Baizhou Zhang
2026-09-11 23:39:53 -07:00
3bb2a5231a
[CI] Extend DeepGEMM SM90 test timeout to four hours (#39194 )
Baizhou Zhang
2026-09-11 23:36:54 -07:00
1cdc5bca5e
fix(qsa): clamp the compress gather to the rows this forward has (#38346 )
ehuaa
2026-09-12 14:24:17 +08:00
0d08668821
[Cookbook] Kimi-K3: keep DCP under HiCache L1+L2 with DSPARK (#39190 )
Khoa Pham and Claude Fable 5.1
2026-09-11 23:14:24 -07:00
55e5e21c88
[Qwen3.8-Next] Add PD state transfer for Flash Next (#36651 )
YAMY
2026-09-11 22:31:45 -07:00
dbd4302bbb
Keep NVFP4 blockscale swizzle padding on the input device (#39141 )
Ziang Li
2026-09-11 22:18:02 -07:00
c389dd0863
Fix RunAI object-storage checkpoint index filtering (#38988 )
Aurick Qiao
2026-09-11 21:25:57 -07:00
a984c78330
[NPU] Support DFlash speculative decoding for MiMo-V2.5-Pro (mxfp4) (#37565 )
iridiumine
2026-09-12 11:59:23 +08:00
9c365e97b1
[CI] Fix DeepGEMM sanitizer setup and Blackwell test timeouts (#39163 )
Baizhou Zhang
2026-09-11 20:54:52 -07:00
ff1ce11348
[diffusion] model: support VDN-H3 with a hybrid_window_attn_h3 backend (#37903 )
2026-09-11 20:36:32 -07:00
e91c948057
feat: support TP>1 Domino rollout for DFlash V2 (#37069 )
Francis and Qiaolin-Yu
2026-09-12 09:51:11 +08:00
0d1bea77da
[AMD] Allow aiter attention backend for Gemma-4 (#38758 )
Vignesh Sethuraman
2026-09-11 18:48:26 -07:00
e1d364dd1f
[AMD] aiter: route head_dim>256 prefill through Triton unified_attention (#38757 )
Vignesh Sethuraman
2026-09-11 18:40:11 -07:00
b805cc5014
[Bugfix] Track DFlash Mamba state at checkpoint boundaries (#37818 )
2026-09-12 09:38:09 +08:00
29cc9d2dc9
[CI] Install elfutils headers for DeepGEMM wheel builds (#39152 )
Baizhou Zhang
2026-09-11 17:29:40 -07:00
a207786205
[PD] Transfer the DCP-replicated DSPARK draft KV in DCP1->DCP-N relayouts (#37709 )
2026-09-11 17:15:18 -07:00
7d9c57da6e
[Session] Fix session idle timeout after rejected requests (#39035 )
2026-09-11 16:59:43 -07:00
5f3606c7b2
[Cookbook][AMD] Kimi-K3 MI350X/MI355X: pin a ROCm image with the DSPARK graph-capture fix, add measured cell numbers (#39029 )
Kevin Mi
2026-09-11 16:10:44 -07:00
6671cfc775
[AMD] Fix weight checking for AITER-shuffled block FP8 weights (#34330 )
Xinyu Kang
2026-09-11 18:49:56 -04:00
f963d7a27c
fix(qsa): make the paged sparse-decode gather memory-safe (zero-fill scratch, int64 offsets, dequant FP8 on gather) (#38851 )
YAMY
2026-09-11 15:47:25 -07:00
ec30f19e4a
[LoRA] Support MoE in full and breakable prefill CUDA graphs (#38578 )
Yanbin Jiang
2026-09-11 14:40:02 -07:00
45715e7f20
[Fix] Disable NCCL graph buffer registration for the TP LM-head all-to-all (pure-DP decode hang under request bursts) (#38936 )
Hanming Lu
2026-09-11 13:10:48 -07:00
a338a9a01c
[AMD][CI] Skip failing Wave test and relax multi-LoRA output check (#38585 )
Bingxu Chen
2026-09-12 03:25:20 +08:00
df6424967a
[docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611 )
2026-09-11 11:58:30 -07:00
d7c284b894
[AMD] Use the triton DSA backend for GLM-5.2 MXFP4 on MI355X (#39106 )
Zhang, Jiejing
2026-09-11 11:48:50 -07:00
833bce9df5
[AMD][DCP 1/N] add dcp support for aiter backend (#34432 )
billishyahao and HAI
2026-09-12 01:35:51 +08:00
f69d6fc28a
[OpenAI] Propagate PD routing metadata through /v1/responses (#35503 )
Jeremy Zhang
2026-09-12 01:33:38 +08:00
ddf02a4f58
[AMD] aiter: resolve SWA KV pool for draft workers + guard paged decode (#38756 )
Vignesh Sethuraman
2026-09-11 10:13:48 -07:00
4309c7ce19
Fix stale DSV4 indexer metadata names in the TopK v2 dispatch test (#39101 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-11 23:31:42 +08:00
d6b5dca90c
[Diffusion] Optimize SANA-WM convolution post-processing and streaming GDN (#38529 )
Xiaoyu Zhang
2026-09-11 23:30:17 +08:00
2c10f87991
[PD] Read nixl TransferInfo.is_dummy as a field in unit tests (#39100 )
Shangming Cai
2026-09-11 23:20:02 +08:00
2a46cf2ca0
[PD] Share the prefill->decode failure notification across backends (#36612 )
2026-09-11 23:16:17 +08:00
e016de462c
[diffusion] CI: expose nightly server telemetry coverage (#38782 )
Mick and Mick Qian
2026-09-11 23:07:37 +08:00
7f09fbcd25
[Diffusion] Preserve BF16 rounding in Hopper LTX QKNorm and RoPE fusion (#38533 )
Xiaoyu Zhang
2026-09-11 23:02:54 +08:00
dd67a42634
[diffusion] UX: quiet request-path cache diagnostics (#38783 )
Mick and Mick Qian
2026-09-11 22:45:21 +08:00
593c7a900d
Add granite_thinking_parser reasoning parser for Granite 4.2 (#38693 )
2026-09-11 10:18:35 -04:00
335f6aab27
[DSV4] Support raw-index output in TopK v2 (#33672 )
2026-09-11 22:17:36 +08:00
ab9750fb35
[diffusion] optimization: pin layerwise host stores in place at their exact size (#39021 )
Mick and Claude Opus 5
2026-09-11 22:04:40 +08:00
e8a36d339c
Auto-detect GLM-5.3 chat templates as glm45/glm47 parsers (#38297 )
Xinyuan Tong
2026-09-11 21:02:44 +08:00
358c163250
[diffusion] fix: recover ipc jit initialization after interrupted builds (#39034 )
Mick and Mick Qian
2026-09-11 19:34:09 +08:00
165d8dd177
[NPU][Hicache] Add Ascend Memcache Hicache L3 storage backend (#38827 )
James
2026-09-11 18:54:00 +08:00
747734dce4
[NPU]glm5.2 fp8 memory opt (#38807 )
Liwansi
2026-09-11 18:48:45 +08:00
a4ff5634b8
[CI] Add CI permissions for PP contributor stepinto (#39077 )
Shangming Cai
2026-09-11 18:04:57 +08:00
8e7deb329e
[Router] Preserve global cache affinity with bucket routing (#38814 )
Vincent Gao
2026-09-11 17:41:27 +08:00
822e73ccdd
[Unified Cache][AMD] Support DeepSeek-V4 unified KV in direct external linkers (#38269 )
2026-09-11 16:49:29 +08:00
0bae67648a
[NPU] Change npu.Dockerfile working directory to /sgl-workspace (#38968 )
Jensen
2026-09-11 16:30:07 +08:00
ad5af539cd
[CI] Trim DSV4 trtllm B200 tests (#39013 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-11 16:04:16 +08:00