Commit Graph
16750 Commits
Author SHA1 Message Date
f7cb328eb7 [AMD] [GLM5] Skip DSA decode indexer when kv_len <= index_topk (dense k-only fast path) (#31324)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-08-16 16:30:54 -07:00
Liangsheng Yin 5e73c89b34 [Spec] Simplify compute_spec_v2_logprobs signature and skip identity gathers (#35058) 2026-08-16 16:01:26 -07:00
Yuwei AnandClaude Opus 5 a508d60295 [BCG][6/N] Allow prefill breakable CUDA graph for the Kimi archs (#34245)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 15:57:14 -07:00
Lianmin Zheng b7eccd642f Increase post-capture decode memory reserve (#34996) 2026-08-16 15:31:36 -07:00
Lianmin Zheng f61f584347 Add explicit EPLB balancedness reporting modes (#34998) 2026-08-16 15:31:11 -07:00
Liangsheng Yin 77cadf6b98 [Spec] Point multi-layer eagle's last shared-read runner at the draft runner (#35057) 2026-08-16 15:28:12 -07:00
Lianmin ZhengandJialin Ouyang c6ebcf39ee [VLM] Avoid synchronizing multimodal placeholder counts (#34995)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-08-16 15:15:49 -07:00
Lianmin ZhengandYonghao Zhuang 4c51248427 Support unified SWA page mapping in attention metadata (#35000)
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
2026-08-16 15:14:50 -07:00
Lianmin ZhengandYe Qi 32e6fb4fdc [Frontend] Apply request header overrides to chat completions (#35001)
Co-authored-by: Ye (Charlotte) Qi <ye.charlotte.qi@gmail.com>
2026-08-16 15:08:47 -07:00
Lianmin ZhengandLu Fang e49557b8da Support model-defined prefill input embedding width (#35002)
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
2026-08-16 15:08:19 -07:00
EthanandQAQEthan 5534380d46 [Spec] Support logprobs with DSpark speculative decoding (#34696)
Co-authored-by: QAQEthan <QAQEthan@users.noreply.github.com>
2026-08-16 15:05:13 -07:00
Lianmin Zheng 67e12131df Build Rust extensions on demand in source checkouts (#34994) 2026-08-16 14:58:06 -07:00
Lianmin Zheng 0e231d365a Clean up playground scripts and add PR babysitter launcher (#35018) 2026-08-16 14:43:57 -07:00
Liangsheng Yin bae353ba55 [misc] Rename shared-read boundary to shared-read ends and fix wrapper delegation (#34982) 2026-08-16 14:36:31 -07:00
Ke Bao ace7314173 Add bit-exact guard for extra_buffer_lazy (#35030) 2026-08-16 23:34:46 +08:00
Mick d3589a7251 [diffusion] CI: tighten NVIDIA perf baselines (#35016) 2026-08-16 20:54:50 +08:00
Xiaoyu Zhang 41abbb0d32 [diffusion] Accelerate Cosmos3 T2I QKNorm+RoPE (#34932) 2026-08-16 20:15:32 +08:00
Xiaoyu Zhang 095ec6c997 [diffusion][kernel] Accelerate Sana BCG with bit-exact conv post-processing (#34928) 2026-08-16 20:05:57 +08:00
Mick 3d3194f6c3 vlm: cache kimi-k3 per-image processor artifacts (#34404) 2026-08-16 19:51:13 +08:00
Mick 968b355f12 vlm: streamline vision sdpa reshapes (#34991) 2026-08-16 19:05:53 +08:00
Xiaoyu Zhang 0761d3f3a4 [diffusion] Accelerate lossless Ideogram norm post-processing (#34931) 2026-08-16 17:20:09 +08:00
Xiaoyu Zhang b752f1e533 [diffusion] Enable breakable CUDA graphs for LTX-2.3 (#34929) 2026-08-16 17:18:43 +08:00
Lianmin Zheng 6bb73082c8 Add skill for babysitting PR CI (#35015) 2026-08-16 01:23:32 -07:00
Mohammad Miadh AngkadandMohammad Angkad 6ab4b99bc2 [Quantization] Fix GPTQ scheme attachment broken by LinearBase.scheme default (#34962)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-08-16 00:48:02 -07:00
Mick 2ee0d38a85 [diffusion] chore: refresh docs, retire stale knobs, and fix nightly attribution (#34663) 2026-08-16 15:41:08 +08:00
Mick a54de989c8 [diffusion] chore: speed up minimax-h3 vae decode on 2×h100 (#34817) 2026-08-16 15:38:21 +08:00
cctry 8922bb98e2 refactor(hicache): flatten L2 transfer execution (#34793)
GB300 test fails unrelated
2026-08-16 00:33:34 -07:00
Chenzhou LiandXiaoyu Zhang 56a759cffc [JIT Kernel] Migrate moe_topk_softmax from AOT to JIT (#34509)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-08-16 15:02:57 +08:00
LinyuanLi 0da87024d3 [NPU] Add mxfp4-w4a8 MOE Quantization Support for NPU (#30318) 2026-08-16 14:03:17 +08:00
Zhaoyi Liandjacky.cheng 24ab8f9ed9 [AMD] Qwen3.5: guard attn layers against empty DP-attention batch (#34474)
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-08-15 21:28:13 -07:00
66de161976 [Fix][AMD] MoRI EP: drop record_stream in TBO dispatch/combine (HSA out-of-resources) (#32746)
Co-authored-by: billishyahao <bill.he@amd.com>
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
2026-08-15 21:23:48 -07:00
wangwenmingaa 4654b927eb [HiCache] Optimize LogicalHostPool free-list release (#33998) 2026-08-16 12:17:33 +08:00
jacky.cheng 6314e9e4f5 [AMD][Fix] Qwen3.5: guard zero-grid launch in fused_qk_gemma_rmsnorm(_with_gate) (HIP invalid configuration on idle DP rank) (#31794) 2026-08-15 21:02:54 -07:00
Ke Bao 0f706c33d2 Fix swa eviction frontier for bigram keys (#34870) 2026-08-16 11:52:14 +08:00
Mick e9fe58139f [diffusion] refactor: unify component residency controls (#34736) 2026-08-16 11:24:48 +08:00
Spandan Tiwari f68517f644 [AMD][Quantization][Bugfix] Fix bug related to fp8 max on gfx95x for per-token-group quant (ROCm) (#30900) 2026-08-15 19:54:16 -07:00
Raiden MakotoandRaiden-Makoto 4c0e85524d [AMD] [GLM5] Enable dense-MHA short-context prefill fallback on gfx950 (#30808)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-08-15 19:30:14 -07:00
Mick d269a28b47 [diffusion] refactor: route minimax h3 vae attention through native backends (#34949) 2026-08-16 10:07:26 +08:00
Mick 4f9da62547 [diffusion] chore: use native hunyuan3d paint and delight models (#34980) 2026-08-16 10:03:48 +08:00
Mick d106e8b23a [diffusion] chore: use native ernie prompt enhancer (#34951) 2026-08-16 09:59:51 +08:00
Mick 19e3bd6391 [diffusion] chore: use native qwen3-vl vision encoder (#34945) 2026-08-16 09:58:46 +08:00
Mick eb6b773149 [diffusion] chore: use native qwen2.5-vl generation (#34896) 2026-08-16 09:57:39 +08:00
gilfordtingandClaude Fable 5 4a6dc267e1 [Spec] Support mamba-radix-cache-strategy extra_buffer_lazy with DFLASH (#34763)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 16:36:59 -07:00
Mohammad Miadh AngkadandMohammad Angkad 3802a725ac [CI] Pin the allowed_media_domains supplied-instance reads in the step-12 ratchet (#34961)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-08-15 16:35:19 -07:00
d22c4cc177 [AMD] perf(sgl-kernel): default block_quota=16 for MLA page_first KV gather… (#30024)
Co-authored-by: Niko Ma <nima@amd.com>
Co-authored-by: figo <fizhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
2026-08-15 16:05:40 -07:00
Liangsheng Yin 0f7aaceda5 [misc] Rename the WAR read-done fastpath to shared-read-done (#34916) 2026-08-15 15:02:02 -07:00
Thomas Wang 4d0c5a89af [AMD] Add concat_and_cast_mha_k_pad_kernel to support 12-head and enable K3 aiter prefill kernel (#34837) 2026-08-15 14:39:30 -07:00
Ke Bao dd458f3212 Fix dsv4 kl test timeout (#34963) 2026-08-16 01:57:40 +08:00
cctryandYilong Zhao e5b3a48751 Add --http2-max-concurrent-streams server arg (#34796)
Co-authored-by: Yilong Zhao <74357408+happierpig@users.noreply.github.com>
2026-08-15 10:34:49 -07:00
huangtingwei cef8a32b9d update codeowner (#34866) 2026-08-16 00:37:18 +08:00