 Lianmin Zhengandcctry
|
c19a333944
|
[mm] Handle per-item embeddings in cache misses (#32498)
Co-authored-by: cctry <cctry@meta.com>
|
2026-07-27 17:16:17 -07:00 |
|
 Caio RochaandCheng Wan
|
5a46e16f01
|
[Fix] Enable graph capture and MSCCL++ for attention TP groups (#31629)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-07-27 15:56:41 -07:00 |
|
Mohammad Miadh Angkad
|
3005af0941
|
Fix compressed-tensors NVFP4 MoE W13 layout (#32430)
|
2026-07-27 14:56:42 -07:00 |
|
Kangrui Du
|
8a311d1c88
|
[diffusion] fix: preserve tensor stride when offloading rollout weights to pinned host memory (#32420)
|
2026-07-27 12:05:13 -07:00 |
|
James Liu
|
1da062f018
|
[Inkling] Add minimal DFLASH support (#31840)
|
2026-07-27 12:01:13 -07:00 |
|
Yuhao Yang
|
7cae831e41
|
Update mi35x ROCm image to k3-20260727 (#32559)
|
2026-07-27 10:52:41 -07:00 |
|
 Mohammad Miadh AngkadandXinyuan Tong
|
3ebb7c2d07
|
docs: point Kimi-K3 references to public branch (#32547)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-07-27 16:23:43 +00:00 |
|
      
|
7dafacca49
|
docs(cookbook): add the Kimi-K3 serving cookbook (#32542)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: thomawan <thomawan@amd.com>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-07-27 08:37:30 -07:00 |
|
 
|
8d6549bc40
|
[Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-27 21:31:04 +08:00 |
|
inkcherry
|
5656de2d9a
|
[PD] pool decode bootstrap HTTP sessions (#31543)
|
2026-07-27 21:07:23 +08:00 |
|
Peng Xingchen
|
db9143ee08
|
[NPU] Fix MTP IndexShare warm-up for attention DP and prefill CP (#32210)
|
2026-07-27 19:19:38 +08:00 |
|
Lianmin Zheng
|
34454c06b8
|
[Refactor] Tidy server_args.py section grouping and drop unused alias (#32496)
|
2026-07-27 04:09:09 -07:00 |
|
 
|
1d350aaad3
|
fix(reasoning): let --enable-strict-thinking works for DeepSeek-V4 (#32400)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-07-27 18:15:47 +08:00 |
|
Mick
|
08af5aea57
|
optimize: optimize EmbeddingGemma prefill performance (#32383)
|
2026-07-27 17:34:29 +08:00 |
|
 Jackey HuaandClaude Opus 5
|
9a0bd24bed
|
model: serve bare Qwen3Model backbone natively as an embedding model (#32457)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-27 15:49:58 +08:00 |
|
McZyWu
|
169fc1e20c
|
[NPU] Acc fix for afmoe model introduced by topk refactor. (#31280)
|
2026-07-27 15:40:49 +08:00 |
|
McZyWu
|
c0f47a06fc
|
[NPU] Determine the topk norm_type through scoring_func (#31393)
|
2026-07-27 15:12:38 +08:00 |
|
Zheng Wengang
|
3d3ba4f746
|
[BugFix][EPD] Fix Mooncake source-MR lifecycle for multi-TP /send (#32071)
|
2026-07-27 15:06:01 +08:00 |
|
Baizhou Zhang
|
082b2a10b6
|
Add local ZIP uploader for whl releases (#32489)
|
2026-07-27 00:03:59 -07:00 |
|
Sam Shleifer
|
5cc273a780
|
[feat] Opt-in flat response format for prompt top logprobs (#32078)
|
2026-07-26 23:44:34 -07:00 |
|
 Yiqi Yangandhzh0425
|
4ea17169b0
|
[UnifiedTree] fix: drop prefetched host refill under an un-backed-up parent (#31902)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-07-27 13:53:02 +08:00 |
|
Ethan (Yusheng) Su
|
ee1736f39a
|
[LoRA] Support LoRA under the breakable/full prefill CUDA graph (#30988)
|
2026-07-26 22:10:03 -07:00 |
|
Xinyu Jiang
|
c6a6200a1a
|
[AMD] Fix pip setup in AMD Miles nightly builds (#32469)
|
2026-07-27 12:32:40 +08:00 |
|
 DAI0818andybyang
|
2abb1d2c37
|
fix(hisparse): correct DSA KV memory budget (#31992)
Co-authored-by: ybyang <ybyang7@iflytek.com>
|
2026-07-27 11:19:32 +08:00 |
|
Mick
|
abb8f4b5e3
|
model: support EmbeddingGemma (#32375)
|
2026-07-27 10:40:47 +08:00 |
|
icarus_zh
|
a358374ae9
|
[NPU][Fix Issue]: Send expert weights contiguous tensor across cards during EPLB rebalance (#32001)
|
2026-07-27 09:20:28 +08:00 |
|
Lianmin Zheng
|
3863612023
|
Add oulgen to CI_PERMISSIONS.json (#32453)
|
2026-07-26 15:04:11 -07:00 |
|
shadowxz109
|
e14068d161
|
[NPU]Add Ascend transfer version compatibility. (#31189)
|
2026-07-26 21:07:59 +08:00 |
|
icarus_zh
|
e8a635a412
|
Load initial expert location metadata on CPU (#32435)
|
2026-07-26 20:44:25 +08:00 |
|
ming_wang
|
a76b74cbe0
|
add fill_draft_extend_prepare_buffers_native for NPU (#32427)
|
2026-07-26 20:42:14 +08:00 |
|
xdtbynd
|
78d7928296
|
[UT][NPU] add NPU attention unit tests for ascend_backend and ascend_dsv4_backend (#32294)
|
2026-07-26 19:59:44 +08:00 |
|
Wang, FangYuan
|
61057bda6c
|
[BugFix] Prevent TBO crash when return_logprob is enabled (#32180)
|
2026-07-26 00:07:54 -07:00 |
|
jacky.cheng
|
833e1bc601
|
[Fix][AMD] Qwen3.5 MoE: disable global-slot shared-expert fusion under per-rank EP backends (MoRI + dp-attention init crash) (#31793)
|
2026-07-25 23:59:47 -07:00 |
|
silencejade
|
72e415dfc8
|
[Bugfix] Fix prefill suspension caused by delayed negotiate_should_allow_prefill invocation (#32389)
|
2026-07-26 14:59:26 +08:00 |
|
YC Yen-Ching Tseng
|
1d0cd2e473
|
[AMD] Nightly Test Coverage - Minimax-M3-MXFP8 Accuracy Test (#30613)
|
2026-07-25 23:52:22 -07:00 |
|
 ormandjandMohammad Miadh Angkad
|
2cbddb842d
|
[DSV4/SM120] Allow fused MHC opt-in with standalone TileLang pre disabled (#30954)
Signed-off-by: David Orman <ormandj@corenode.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-07-26 09:02:54 +08:00 |
|
 
|
55c4853487
|
[comm] Enable multi-node custom-AR v2 on a single NVLink clique (#32339)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-07-25 17:25:18 -07:00 |
|
 
|
2c63a2f12b
|
Fix --hicache-size allocating ~2x host memory on hybrid SWA (#32373)
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-07-25 17:19:44 -07:00 |
|
 Lianmin ZhengandAlec S
|
9989077f24
|
Use native batched llguidance mask generation (#32412)
Co-authored-by: Alec S <10566873+alecsolder@users.noreply.github.com>
|
2026-07-25 16:36:32 -07:00 |
|
 Lianmin ZhengandXingyu Liu
|
fae84ac0f9
|
Fix token count localization for replicated attention-TP forwards (#32411)
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
|
2026-07-25 16:36:15 -07:00 |
|
Lianmin Zheng
|
5f330004bd
|
Fix flaky test_sampling_mask: mask length can legitimately be top_k + 1 (#32410)
|
2026-07-25 16:05:57 -07:00 |
|
Liangsheng Yin
|
3da1071d56
|
[Spec] Hold the grammar bitmask in one GrammarMask type across all decode paths (#32409)
|
2026-07-25 15:27:43 -07:00 |
|
Lianmin Zheng
|
d3cf4dfbaa
|
Update audio container test time estimate (#32408)
Signed-off-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-07-25 14:45:08 -07:00 |
|
Jialin Ouyang
|
cd145f840f
|
Radix Cache Split: Spin off TreeCore (#29901)
|
2026-07-25 14:31:59 -07:00 |
|
 Kangrui DuandYihao Wang
|
a23f6ea090
|
[Diffusion] offload rollout weights to pinned host memory (#32032)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
|
2026-07-25 13:35:53 -07:00 |
|
YAMY
|
91f386a5b2
|
fix(disagg): support pipeline-parallel hybrid-linear transfer (#32270)
|
2026-07-25 13:34:38 -07:00 |
|
 Cheng WanandShu Wang
|
659d349b61
|
[core/loader] Add presharded load format (#24256)
Co-authored-by: Shu Wang <shuwanguc@google.com>
|
2026-07-25 13:03:39 -07:00 |
|
Mohammad Miadh Angkad
|
9791fc7090
|
Add configurable FlashInfer autotune skips (#31389)
|
2026-07-25 11:17:58 -07:00 |
|
Mohammad Miadh Angkad
|
953c587adf
|
[Docs] Add Qwen3.6 35B NVFP4 to cookbook (#31413)
|
2026-07-25 11:16:53 -07:00 |
|
Ke Bao
|
69a3c54c70
|
Fix SWA admission livelock on cached-prefix resumes (#32379)
|
2026-07-25 22:37:35 +08:00 |
|