Commit Graph
3724 Commits
Author SHA1 Message Date
Rain Jiang 4af8ddb576 support rust sglang server (#29799) 2026-07-31 11:56:31 -07:00
Cheng Wan 77c77a3da8 feat(inkling): migrate short convs onto the ShortConv attention backend (#33023) 2026-07-31 11:52:12 -07:00
Cheng Wan d3222bcc3a [unified-memory] Support fa3, the default MLA backend on pre-Blackwell hosts (#33046) 2026-07-31 11:46:46 -07:00
luchangli 26486a957d Fix --hicache-size allocating ~2x host memory on hybrid Mamba (#32915) 2026-08-01 02:37:57 +08:00
Nan Jiang 89f4a80c1f Support fastsafetensors no-GDS loading and page-cache release (#31859) 2026-07-31 23:12:32 +08:00
Danila ShtanandDanila Shtan 5f9b0db18c Fix async loading of RunAI-streamed tensors (#32896)
Co-authored-by: Danila Shtan <dan@nebius.com>
2026-07-31 21:46:33 +08:00
luoroger37andXinyuan Tong 690de097c4 [fix]reject media input for text-only models (#32914)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-07-31 12:43:25 +00:00
Peng Wu e3d4f48e55 [Fix] missing max_context_len on HybridAttnBackend (#32690) 2026-07-31 19:43:09 +08:00
Baizhou Zhang fd28242b68 [CI] Pin NCCL ports for GB300 PR tests (#33044) 2026-07-31 02:28:37 -07:00
Khoa PhamandYangmin Li 2573190b93 feat: support Kimi Linear PD disaggregation with DCP (#32837)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-07-31 02:14:09 -07:00
Cheng Wan 33c27d8e7f [unified-memory] Let Kimi-Linear use the paged MLA attention backends (#32972) 2026-07-31 01:32:08 -07:00
Liangsheng YinandKaixi Matteo Chen 5c6635d8f3 [Spec] Compact the target-verify mask when nothing reads it (#32920)
Co-authored-by: Kaixi Matteo Chen <kaiximatteoc@nvidia.com>
2026-07-30 23:13:21 -07:00
Shijin ZhangandXinyuan Tong 09193bf36f [Fix]: render tool_reference schema regardless of tool_result part order (#32522)
Signed-off-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-07-31 13:57:52 +08:00
Cheng Wan e23ccb15f0 [unified-memory] Support MLA-hybrid-Mamba (Kimi-Linear) on the Triton backend (#32971) 2026-07-30 22:10:34 -07:00
Cheng Wan 06ccaef24a Fix silently wrong EPLB output with --moe-a2a-backend none (rank-invariant dispatch) (#32962) 2026-07-30 22:10:04 -07:00
Qiaolin Yu f3fd869494 [gdn] support replayssm with extra buffer (#32692) 2026-07-30 21:34:37 -07:00
DAI0818andXiaoyu Zhang afeaeccfa2 perf(hisparse): eliminate redundant swap output fill (#32483)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-07-31 12:09:44 +08:00
Ziang LiandParth Chadha 0aefba7283 fix(dsa): correct packed FlashInfer top-k and backend selection semantics (#32490)
Co-authored-by: Parth Chadha <parth@humansand.ai>
2026-07-30 20:20:16 -07:00
Mick a149717308 feat: log multimodal encoder DP tradeoffs (#30903) 2026-07-31 08:50:20 +08:00
Xinyuan Tong 68d442945f Flush dropped reasoning at stream end when stream_reasoning=False (#32225) 2026-07-31 08:25:54 +08:00
Trang DoandCheng Wan a1c30701aa Integrate pplx a2a backend (#30756)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2026-07-30 15:33:19 -07:00
Mohammad Miadh Angkad 3a53c26c27 [CI] Fix MoE compile and DSA indexer regressions (#32937) 2026-07-30 15:21:51 -07:00
Trevor Morris a6221d776f feat: Support nvidia/MiniMax-M3-NVFP4 (#31989) 2026-07-30 14:32:03 -07:00
Henning Thieß c4af6cf263 Qwen3.5-MoE: support modelopt_fp4 checkpoints that quantize attention (+ load baked FP8 KV scales) (#31220) 2026-07-30 14:30:26 -07:00
Broduker b61cb5f9de Fix DeepSeek V4 loading with RunAI Model Streamer. (#30240) 2026-07-30 23:03:34 +08:00
Liangsheng YinandKaixi 6ab3231b97 [Perf] Skip the target-verify tree mask fill when the backend never reads it (#32886)
Co-authored-by: Kaixi <kaiximatteoc@nvidia.com>
2026-07-30 02:32:38 -07:00
Ding Yinandyinding fc007e1f00 Add SM90 FP8 MegaMoE support for DeepSeek-V4 (#29016)
Co-authored-by: yinding <yinding@bytedance.com>
2026-07-30 01:48:10 -07:00
f46d5f25b4 [4/N][CP] Support interleave strategy for cp v2 (#30482)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-07-30 01:45:32 -07:00
Liangsheng Yin c192145830 [Kernel] Fuse KV-cache writes for asymmetric K/V (head_dim != v_head_dim) (#32813) 2026-07-30 00:26:10 -07:00
Ho-Ren (Jack) ChuangandClaude Fable 5 e4a40a71f8 [DSA] Q8KV8 FP8 Sparse Prefill on GLM-5.2 & DeepSeek-V3.2: Q8-Path & Shared-Path Optimizations (#31888)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 15:15:11 +08:00
wenxuewuhdandronnie_zheng 36afd442c7 [DLLM] vectorized joint/low-confidence decoding and skip redundant attn init (#21094)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-30 09:13:14 +03:00
Ke Bao 07a087bf45 Fix Inkling tool-call parsing recovery, content handling, and streaming (#32861) 2026-07-30 14:11:09 +08:00
Mick 22faf9fef8 embedding: centralize capabilities and complete OpenAI compatibility (#32481) 2026-07-30 10:28:52 +08:00
Sam ShleiferandClaude Fable 5 62dfaaa0e0 [Nemotron] Fix decode track-save reading the stale tail of the CUDA-graph track buffer (#32555)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 19:03:44 -07:00
Xuanyi LiandR0CKSTAR 8fbf960980 [MLX] Size request capacity by attention DP (#32115)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-07-29 18:18:22 -07:00
Xiaoyu Zhang 1d9c292547 [Kernel] Add inventory guards and clean benchmark layout (#32788) 2026-07-30 09:03:24 +08:00
Mohammad Miadh Angkad a55e1764a2 Enable GPT-OSS FlashInfer MXFP4 on SM120 (#32668) 2026-07-30 00:04:23 +00:00
Liangsheng Yin e5c46ff07d [Fix] Route asymmetric-KV models to fa4 on SM100 and pin MiMoV2 FP8 MoE to flashinfer_trtllm (#32818) 2026-07-29 16:37:43 -07:00
Sam (Kesen Li) 8fc54d46ef Fix MoE reduce-scatterv eligibility check (#32663) 2026-07-29 14:58:55 -07:00
Willow LopezandJiminator ffd4705baa fix(reasoning): honor Poolside template thinking defaults (#32540)
Co-authored-by: Jiminator <69131491+Jiminator@users.noreply.github.com>
2026-07-29 21:02:19 +00:00
Sam Shleifer d0e69d3881 [feat] Optional base64 encoding for the flat prompt top logprob arrays (#31960) 2026-07-29 12:15:56 -07:00
ziang663andChao Shi eefb434d17 [PD+PP] Honor PP consensus for bootstrap and prealloc (#31869)
Co-authored-by: Chao Shi <chao.shi@alibaba-inc.com>
2026-07-30 01:19:17 +08:00
Yoray Zack 62d0f81f16 [2/N] elastic-ep: Enable EPLB after scale-up (#30553) 2026-07-30 01:06:56 +08:00
Hert4andMohammad Miadh Angkad f69af7b7ad [Bugfix] compressed-tensors: mixed-precision checkpoints silently load unquantized (#32736)
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-07-29 09:55:33 -07:00
YAMY fddfc1fb5e [GDN] Support FlashInfer GDN prefill with extra-buffer radix cache (#29735) 2026-07-30 00:47:35 +08:00
huangtingweiandalphabetc1 50029f05a3 [HiCache] Merge HiCache event checks to reduce decode overhead (#30511)
Co-authored-by: alphabetc1 <47200617+alphabetc1@users.noreply.github.com>
2026-07-29 23:48:48 +08:00
Bingxu Chen bfc450248e [AMD] Replace MI325 with MI300 CI Runners (#31409) 2026-07-29 08:18:07 -07:00
Ke Bao 50b029257f Skip mamba lock during decoding (#32228) 2026-07-29 21:53:28 +08:00
pllimax d004a15a3e Fix GLM4-7B-Flash accuracy test configuration, tune Qwen3.6-27B/35B performance test parameters, and harden Ascend NPU multi-node E2E test utilities against pod name format errors. (#32371) 2026-07-29 21:44:17 +08:00
LZW 4c82bb3252 Add Mooncake tenant id support (#30256) 2026-07-29 18:17:41 +08:00