Commit Graph
17892 Commits
Author SHA1 Message Date
AMD-yanfeiwangandkk 7825e5ffca [AMD] Fix DSV4 unified attention sink TP slice (#35092)
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
2026-09-03 22:59:53 -07:00
Alex NailsandClaude Opus 5 ebae8ee21e [Fix] Register triton.runtime.cache.triton_key in the MPS stub so torch.compile keeps working (#37937)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 22:56:14 -07:00
Xiaoyu Zhang 06b8749803 [Diffusion][Docs] Add single-GPU large-VRAM performance notes (B300) (#37891) 2026-09-04 13:33:28 +08:00
Mick ed67b73081 [diffusion] UX: reduce hot-path server log noise (#37804) 2026-09-04 13:31:00 +08:00
Thomas Wang 225129fe44 [AMD] Update v4 amd cookbook 0903 (#37829) 2026-09-03 22:11:13 -07:00
Mohammad Miadh AngkadandMohammad Angkad ca7b8efd87 [CI] Fix handle_platform_cp_compatibility reading legacy CP flags off the record (#37930)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-03 22:07:04 -07:00
f478b2bb2d Fix: abort handling for dispatched requests after client disconnect (#35255)
Signed-off-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
2026-09-03 21:51:57 -07:00
ashwini rathi e787de5478 [XPU][CI] Move XPU tests to nightly and add per-subclass server launch timeout (#37532) 2026-09-04 12:51:26 +08:00
Mohammad Miadh Angkad 3ad3f23ed5 [Comm] Drop the in-tree MNNVL CuTe DSL port in favor of FlashInfer 0.6.18 (#37206) 2026-09-03 21:49:48 -07:00
Zhangheng e3305b3b87 Add code owner for sglang-simulator (#37922) 2026-09-04 12:20:48 +08:00
Mick 97d081ac76 [diffusion] chore: remove unreachable cosmos3 transfer encoding (#37805) 2026-09-04 12:14:03 +08:00
Baizhou Zhang ff1285cc28 [CP V1 Deprecation 2/5] Make strategy prefill CP canonical (#36223) 2026-09-03 20:24:19 -07:00
59799a3687 [Simulator] Add high-fidelity CPU-based inference simulator (#33824)
Co-authored-by: zhouhaizhu.zhz <zhouhaizhu.zhz@alibaba-inc.com>
Co-authored-by: LinSiyuan814 <linsiyuan.lsy@alibaba-inc.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-09-04 11:12:11 +08:00
Mohammad Miadh AngkadandMohammad Angkad a5f07b1241 [CI] Build the Rust extensions for aarch64 too (#37820)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-03 19:54:08 -07:00
junduandMa Mingfei 8770c1db1f [CPU][CI]: rename Xeon CPU CI suites to stage-*-intel (#37395)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-09-04 10:08:41 +08:00
Mohammad Miadh AngkadandMohammad Angkad c1b4d535d7 [CI] Fix lint (#37895)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-03 18:37:50 -07:00
Aleksi Vesanto 5e81d462c3 [diffusion] fix: only use fused qk_norm on nvidia gpu (#33994) 2026-09-04 09:25:08 +08:00
Li Jinliang 6c63d678d5 [diffusion] webui: fix minimax h3 webui inference settings (#36320) 2026-09-04 09:24:14 +08:00
Kedi WuandKedi Wu 94eb15eb6c [diffusion] quant: support fp8 mixed precision for cosmos3 (#36380)
Co-authored-by: Kedi Wu <kediw@nvidia.com>
2026-09-04 09:23:38 +08:00
Yihao Wang 667bc043dc [CI] remove MINIMAX_H3_HF_TOKEN (#37687) 2026-09-04 09:22:32 +08:00
Mick 2a0602c7ac [diffusion] optimization: unfused w2 bias on SM12.x (cuBLAS 16x16 kernel mis-dispatch) for minimax-h3 vae decoder: (#37835) 2026-09-04 09:17:23 +08:00
Mick 795dd7abce [diffusion] feat: compose third-party component bundles safely (#37816) 2026-09-04 09:15:05 +08:00
eigenandYingyi Huang 516cfbd362 Fix UNO test adapter subdirectory resolution (#37872)
Co-authored-by: Yingyi Huang <averyh@nvidia.com>
2026-09-03 18:07:09 -07:00
heziiop 9ed2721c6d [NPU] fix extend_seq_lens_cpu shape in eager mode (#36843) 2026-09-04 09:03:30 +08:00
Liangsheng Yin 6147a54ddf [PD] Bound transfer engine init with SGLANG_DISAGGREGATION_ENGINE_INIT_TIMEOUT (#37874) 2026-09-03 17:16:55 -07:00
Yueming Yuan 66d60433c1 state_capturer: pin the exact host-cache size via mmap + cudaHostRegister (#37285) 2026-09-03 16:20:59 -07:00
Liangsheng Yin 8b0501399e [CI] Improve Lark CI cards: structured layout, PDT timestamps, slow-only queue digest (#37884) 2026-09-03 16:18:56 -07:00
Zhiqiang Xie b44496389c [HiCache] Count hit allocations and in-flight backups in the buffer pipeline idle check (#37883) 2026-09-03 16:08:50 -07:00
Zhiqiang Xie a480f388b2 [HiCache] L3 storage prefetch lifecycle metrics and cross-tier attribution fixes (#37503) 2026-09-03 16:00:04 -07:00
Liangsheng Yin 0610a6539d [CI] Add Lark notifications for CUDA CI status, runner health, and queue time (#37881) 2026-09-03 15:46:33 -07:00
Yonghao Zhuangandyhzhuang 4dc9cda5f9 [PD] Gate deferred decode KV release on backend capability (#37454)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
2026-09-03 15:42:21 -07:00
Baizhou Zhang fd70325c10 [CI] Add dspark + dsv4 e2e test (#37665) 2026-09-03 14:26:25 -07:00
Mohammad Miadh Angkad 4372b8efa7 [1/N] Quantization Refactor: remove dead code and dedup the FP4 marlin helpers (#37552) 2026-09-03 14:22:31 -07:00
Liangsheng Yin 0e5414fd2f [Test] Allow top-k cutoff ties in test_sampling_mask_matches_topk_logprobs (#37873) 2026-09-03 14:13:26 -07:00
Lianmin Zheng 05dbe64dff Fix buffer-mode idle tracking and VLM memory sizing (#37567) 2026-09-03 13:56:12 -07:00
Liangsheng Yin 2a980cbf10 [mem_cache] Require page-aligned starts in free_segment and drop the boundary trim (#37729) 2026-09-03 13:28:33 -07:00
Martin Hickey 3ffacf949b [Docs] [BugFix] Sync --tool-call-parser and --reasoning-parser lists with the code (#37788)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
2026-09-03 13:17:33 -07:00
Zhiqiang Xie d0c95f6c91 [HiCache] buffer mode: anchor-lock staged prefetches by default (#37464) 2026-09-03 12:08:22 -07:00
Zhiqiang Xie 68978f8d52 [Scheduler] Count the parked chunked-prefill request in the busy mem check (#37502) 2026-09-03 12:08:17 -07:00
Jimmy ShongandClaude Fable 5.1 2da5802bfa [Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-03 11:43:01 -07:00
Zhanghengand晟海 abed680320 [Unified Cache][5/N]: Integrate external linker mode end to end (#37381)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-09-04 02:02:58 +08:00
Lianmin Zheng 619ab2bcce Fix block-scale swizzling device placement (#37849) 2026-09-03 10:47:04 -07:00
23ab10a63e Support speculative decoding with unified SWA memory (#36403)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-03 10:44:57 -07:00
cctry 33a22b1b08 [Cache] Forward fast prefix matching capability (#37844) 2026-09-03 10:33:30 -07:00
Vincent Gaoandinkcherry 54cadad151 [Router] Add composable scoring and eligibility policies (#37731)
Co-authored-by: inkcherry <mingzhi.liu@amd.com>
2026-09-04 00:01:41 +08:00
Yash AkhauriandBBuf 392841f47c [Bugfix] Support K2 Horizon MoE without MoVA (#37825)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-03 23:53:16 +08:00
triple-muandmickqian bf71035d39 [diffusion] MiniMax-H3: tiered AdaLN plan cache (pinned-host tier + per-plan LRU) (#37266)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
2026-09-03 22:09:21 +08:00
Quanli Liand全力 4e37882a93 [Diffusion][minimax-h3] Add SM120 support for SubBlock sparse attention (#37332)
Co-authored-by: 全力 <liquanli.lql@antgroup.com>
2026-09-03 22:05:22 +08:00
pllimax 3239baef25 [CI][NPU] Fix kimi_k2_6 16p in64k perf test and dsv4-flash testcases (#37760) 2026-09-03 21:57:14 +08:00
dujifengandmickqian 9cb38a3d57 [diffusion] feat: filter duplicate precision variants across custom loaders (#37616)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
2026-09-03 21:54:11 +08:00