Zhiqiang Xie
|
d0c95f6c91
|
[HiCache] buffer mode: anchor-lock staged prefetches by default (#37464)
|
2026-09-03 12:08:22 -07:00 |
|
Zhiqiang Xie
|
68978f8d52
|
[Scheduler] Count the parked chunked-prefill request in the busy mem check (#37502)
|
2026-09-03 12:08:17 -07:00 |
|
 Jimmy ShongandClaude Fable 5.1
|
2da5802bfa
|
[Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-03 11:43:01 -07:00 |
|
 Zhanghengand晟海
|
abed680320
|
[Unified Cache][5/N]: Integrate external linker mode end to end (#37381)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-09-04 02:02:58 +08:00 |
|
Lianmin Zheng
|
619ab2bcce
|
Fix block-scale swizzling device placement (#37849)
|
2026-09-03 10:47:04 -07:00 |
|
 
|
23ab10a63e
|
Support speculative decoding with unified SWA memory (#36403)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-09-03 10:44:57 -07:00 |
|
cctry
|
33a22b1b08
|
[Cache] Forward fast prefix matching capability (#37844)
|
2026-09-03 10:33:30 -07:00 |
|
 Vincent Gaoandinkcherry
|
54cadad151
|
[Router] Add composable scoring and eligibility policies (#37731)
Co-authored-by: inkcherry <mingzhi.liu@amd.com>
|
2026-09-04 00:01:41 +08:00 |
|
 Yash AkhauriandBBuf
|
392841f47c
|
[Bugfix] Support K2 Horizon MoE without MoVA (#37825)
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-03 23:53:16 +08:00 |
|
 triple-muandmickqian
|
bf71035d39
|
[diffusion] MiniMax-H3: tiered AdaLN plan cache (pinned-host tier + per-plan LRU) (#37266)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
|
2026-09-03 22:09:21 +08:00 |
|
 Quanli Liand全力
|
4e37882a93
|
[Diffusion][minimax-h3] Add SM120 support for SubBlock sparse attention (#37332)
Co-authored-by: 全力 <liquanli.lql@antgroup.com>
|
2026-09-03 22:05:22 +08:00 |
|
pllimax
|
3239baef25
|
[CI][NPU] Fix kimi_k2_6 16p in64k perf test and dsv4-flash testcases (#37760)
|
2026-09-03 21:57:14 +08:00 |
|
 dujifengandmickqian
|
9cb38a3d57
|
[diffusion] feat: filter duplicate precision variants across custom loaders (#37616)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
|
2026-09-03 21:54:11 +08:00 |
|
 Yihao WangandClaude Fable 5
|
f5bed255c0
|
[diffusion] feat: support key masks on USPAttention's replicated-prefix path (#36735)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-03 20:13:33 +08:00 |
|
 Yihao WangandClaude Fable 5
|
14444c6a04
|
[diffusion] feat: add maybe_record_function profiler spans for request phases (#35922)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-03 20:12:54 +08:00 |
|
 
|
2bb25dc18b
|
[Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-03 20:08:41 +08:00 |
|
amote-i
|
354ed6d66b
|
[NPU] [DOC] Refresh supported models and features on NPU (#37799)
|
2026-09-03 19:53:39 +08:00 |
|
 kkandwunhuang
|
dd091f43cd
|
[AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2026-09-03 18:44:39 +08:00 |
|
Cheng Wan
|
a11dba1a01
|
[Feature] Unified memory: support decode context parallelism for the trtllm_mla family (#37693)
|
2026-09-03 03:37:12 -07:00 |
|
 YC Yen-Ching TsengandChen
|
f59a4840c5
|
[AMD][CI] Correct MI355X Slurm exclude node (#37779)
Co-authored-by: Chen <bingxche@amd.com>
|
2026-09-03 18:32:44 +08:00 |
|
Duyi-Wang
|
429ac2d82c
|
[AMD] Fix DSv4 draft extend taking the target compression path during prefill (#37713)
|
2026-09-03 02:46:49 -07:00 |
|
 
|
27b7a2dc3b
|
[Kimi K3] Rework skipped-think fix as opt-in force_nonempty_content with streaming coverage (#34187)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2026-09-03 17:37:28 +08:00 |
|
 kkandwunhuang
|
a6001478f4
|
[AMD] Perf Kimi-K3 MoE optimization (#33838)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2026-09-03 02:28:28 -07:00 |
|
Xiaoyu Zhang
|
397aeca376
|
fix(benchmark): support Glm4MoeLite in fused MoE tuner (#37623)
|
2026-09-03 17:21:42 +08:00 |
|
 Xinyi SongandThomas Wang
|
7ed29eba80
|
[AMD] Fix FP4 indexer OOR (#37660)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
|
2026-09-03 01:49:46 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
e59a576f03
|
fix test/manual/test_forward_split_prefill.py UT due to many refactors and design changes (#36617)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
|
2026-09-03 01:48:55 -07:00 |
|
 
|
3bac084d4e
|
[Model] Add native IFM K2 Horizon serving support (#37654)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-03 16:39:43 +08:00 |
|
Yash Akhauri
|
02d9b3060a
|
[Docs] Update K2 Horizon MoE model names (#37723)
|
2026-09-02 23:32:44 -07:00 |
|
 YC Yen-Ching TsengandPhil Li
|
1fb85053e7
|
[AMD][Diffusion] Migrate FlyDSL fused norm kernels to the v0.3.0 stable API (#36349)
Co-authored-by: Phil Li <haicli@amd.com>
|
2026-09-02 23:03:08 -07:00 |
|
 xiaobochen-amdandZhang, Jiejing
|
030d7e7e9b
|
[ROCm] Define the DSA head-gate graph helpers on HIP (#37118)
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
|
2026-09-02 22:55:45 -07:00 |
|
Zaili Wang
|
4b0edb25f2
|
[CPU] Update base image to Ubuntu 26.04 (#35313)
|
2026-09-03 12:42:14 +08:00 |
|
Liangsheng Yin
|
54da74e83e
|
[Fix] Broadcast PP dynamic-chunk profiling failures so every rank disables together (#37675)
|
2026-09-02 21:09:04 -07:00 |
|
Yash Akhauri
|
98ef7d8ae6
|
docs: add K2 Horizon cookbook recipes and H200 results (#37655)
|
2026-09-03 11:52:05 +08:00 |
|
Liangsheng Yin
|
49e6e81830
|
[CI] Accept t64-suffixed apt packages in the install skip check (#37689)
|
2026-09-02 20:51:47 -07:00 |
|
hhhh1252023
|
81b0e4985a
|
Modify KUBE_JOB_NAME to fix the problem of the string being too long (#37572)
|
2026-09-03 11:43:36 +08:00 |
|
 Khoa PhamandClaude Opus 5
|
cf3173aeb9
|
[Perf] Walk the radix tree by offset instead of re-slicing token storage (ported from #36507) (#37324)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-03 11:13:03 +08:00 |
|
 Alex NailsandClaude Opus 5
|
57c26a84e0
|
[chore] Add .git-blame-ignore-revs for the black -> ruff-format reformat (#37210) (#37695)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-02 20:01:48 -07:00 |
|
James Liu
|
4229088a48
|
feat(kernels): generalize persistent CuTe JIT cache (#33911)
|
2026-09-02 19:59:04 -07:00 |
|
 Alex NailsandAlison Shao
|
28262c20df
|
[CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-02 19:46:08 -07:00 |
|
    
|
2641e427be
|
Xpu/weekly simple model enablement 2026 08 30 (#37193)
Co-authored-by: dayanandav <dayananda.vasantha.kumar@intel.com>
Co-authored-by: Girijala, Pavan Sivaram <pavan.sivaram.girijala@intel.com>
Co-authored-by: Cui, Lily <lily.cui@intel.com>
Co-authored-by: Juan Muneton <juan.muneton.gallego@intel.com>
Co-authored-by: Gao, Pengfei <pengfei.gao@intel.com>
|
2026-09-03 09:35:59 +08:00 |
|
Liangsheng Yin
|
a522c8a4b6
|
[misc] Extract PP dynamic chunk sizing into a DynamicChunkSizer scheduler component (#37674)
|
2026-09-02 18:35:43 -07:00 |
|
Mick
|
0dd66def7c
|
[chore] harden checkpoint quantization metadata parsing (#36922)
|
2026-09-03 09:27:22 +08:00 |
|
  
|
fbf909b460
|
[Fix] Alpha-channel images and tool-result media ordering (port of #36507) (#37320)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-09-02 18:23:06 -07:00 |
|
 Lee NauandYangmin Li
|
6e41f1ad29
|
[Fix] Preserve FP32 in SM107 MXFP8 fallback (#37489)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
|
2026-09-02 17:55:14 -07:00 |
|
 Shiyan DengandLianmin Zheng
|
80e8302d03
|
Add SGLANG_CRASH_ON_JIT_COMPILE to forbid on-the-fly JIT compilation (#36615)
Signed-off-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-09-02 17:26:25 -07:00 |
|
   
|
87d60a2229
|
Improve CUDA graph and speculative execution output handling (#37329)
Co-authored-by: jiayisuse <jiayisuse@fb.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
Co-authored-by: Hao Zhang <zhisbug@users.noreply.github.com>
Co-authored-by: Yichao Fu <yichaofu@meta.com>
|
2026-09-02 17:25:27 -07:00 |
|
Alison Shao
|
db1eb48651
|
[CI] Graceful teardown for the PD and HiSparse server fixtures (#37485)
|
2026-09-02 17:22:18 -07:00 |
|
 
|
ff04a00d73
|
Reduce tokenizer overhead and offload CUDA VMM publication (#37330)
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
|
2026-09-02 17:21:08 -07:00 |
|
Oguz Ulgen
|
f15748d965
|
[bench] Support real-traffic replay with early-stop-aware steady-state metrics in bench_one_batch_server (#37469)
|
2026-09-02 17:20:25 -07:00 |
|
Liangsheng Yin
|
5c46ce37f5
|
[Fix] Apply the attention-CP broadcast result in PP dynamic-chunk profiling (#37669)
|
2026-09-02 17:20:07 -07:00 |
|