 vikram singh shekhawatandClaude Sonnet 4.5
|
f7a404e9c3
|
Fix rope config compatibility and VL/transformers-fallback weight loading (#31575)
Co-authored-by: Claude Sonnet 4.5 (1M context) <noreply@anthropic.com>
|
2026-08-17 14:53:59 +08:00 |
|
Liangsheng Yin
|
0099107e8b
|
Revert "[AMD] [GLM5] Fuse shared-expert append into aiter grouped-topk (skip per-layer append kernel)" (#35105)
|
2026-08-16 23:49:49 -07:00 |
|
 ziang663andZhangheng
|
43226af812
|
fix(hicache): limit load-back pending to write-back (#34519)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-08-17 14:23:54 +08:00 |
|
MingxuZh
|
eafbe2cb6f
|
[CI] Install sgl-eval in xeon (CPU) Docker image (#34818)
|
2026-08-17 14:00:44 +08:00 |
|
 Jimmy ShongandClaude Fable 5
|
07a28ec5cf
|
docs: fix Qwen3.8-27B mamba ratio calculator for speculative decoding (#35064)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-16 22:54:57 -07:00 |
|
 Lianmin ZhengandJialin Ouyang
|
9be3044b9c
|
[Engine] Freeze GC after server warmup (#34999)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
|
2026-08-16 22:41:36 -07:00 |
|
Wang, FangYuan
|
eb61cb2823
|
[AMD] Support prefill context parallel two batch overlap for DeepSeek V4 (#33480)
|
2026-08-16 22:40:29 -07:00 |
|
HAI
|
4b06f917ca
|
Upd: code owners (#35094)
|
2026-08-17 13:27:54 +08:00 |
|
Kaixi
|
f3225bceb3
|
[DSV4] Emit TMA-aligned UE8M0 scales for FP8 einsum (#34277)
|
2026-08-16 22:06:41 -07:00 |
|
 Khoa PhamandClaude Opus 5
|
0fb040cbeb
|
[DCP]Localize HiCache DCP indices once per transfer, not per layer (#34889)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-16 21:54:51 -07:00 |
|
Enrique Shockwave
|
92b1d382c7
|
[Fix] Correct dense FP8 Marlin bias ordering (#35020)
|
2026-08-17 03:43:36 +00:00 |
|
  
|
b6d7602914
|
[CPU] Add support for Gemma4 on Xeon (#22498)
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
Co-authored-by: Haotong Zou <haotong.zou@intel.com>
|
2026-08-17 10:52:26 +08:00 |
|
TobyMint
|
3adc70bb5e
|
[MoE] Add H20 fp8_w8a8 tuned configs for Qwen3.8 (triton 3.7.1) + fix Qwen3_5MoeForCausalLM tuning (#34795)
|
2026-08-16 19:41:03 -07:00 |
|
 zijiexiaandClaude Opus 5
|
f019f0b064
|
[Docs] Feature MiniMax-H3 in the popular-models banner (#35068)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-17 02:09:26 +00:00 |
|
hanwlax
|
0e500feae6
|
fix tpot by adjusting the sliding max-prefill-size window size (#34856)
|
2026-08-17 09:50:48 +08:00 |
|
 
|
8e0499bd50
|
[AMD] [GLM5] Fuse shared-expert append into aiter grouped-topk (skip per-layer append kernel) (#31323)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-08-16 18:22:26 -07:00 |
|
Mick
|
0e178c3d22
|
[diffusion] chore: reuse srt siglip vision model (#34988)
|
2026-08-17 09:16:17 +08:00 |
|
Xiaoyu Zhang
|
0aa09ab40d
|
[diffusion] Reuse bit-exact modulation fast path for LTX-2.3 (#34930)
|
2026-08-17 09:04:10 +08:00 |
|
 
|
d91c3682b0
|
[AMD][CI] Add GPT-OSS perf benchmarks to the ROCm 7.2 nightly (#34645)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
|
2026-08-16 16:35:09 -07:00 |
|
 
|
f7cb328eb7
|
[AMD] [GLM5] Skip DSA decode indexer when kv_len <= index_topk (dense k-only fast path) (#31324)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-08-16 16:30:54 -07:00 |
|
Liangsheng Yin
|
5e73c89b34
|
[Spec] Simplify compute_spec_v2_logprobs signature and skip identity gathers (#35058)
|
2026-08-16 16:01:26 -07:00 |
|
 Yuwei AnandClaude Opus 5
|
a508d60295
|
[BCG][6/N] Allow prefill breakable CUDA graph for the Kimi archs (#34245)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-16 15:57:14 -07:00 |
|
Lianmin Zheng
|
b7eccd642f
|
Increase post-capture decode memory reserve (#34996)
|
2026-08-16 15:31:36 -07:00 |
|
Lianmin Zheng
|
f61f584347
|
Add explicit EPLB balancedness reporting modes (#34998)
|
2026-08-16 15:31:11 -07:00 |
|
Liangsheng Yin
|
77cadf6b98
|
[Spec] Point multi-layer eagle's last shared-read runner at the draft runner (#35057)
|
2026-08-16 15:28:12 -07:00 |
|
 Lianmin ZhengandJialin Ouyang
|
c6ebcf39ee
|
[VLM] Avoid synchronizing multimodal placeholder counts (#34995)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
|
2026-08-16 15:15:49 -07:00 |
|
 Lianmin ZhengandYonghao Zhuang
|
4c51248427
|
Support unified SWA page mapping in attention metadata (#35000)
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
|
2026-08-16 15:14:50 -07:00 |
|
 Lianmin ZhengandYe Qi
|
32e6fb4fdc
|
[Frontend] Apply request header overrides to chat completions (#35001)
Co-authored-by: Ye (Charlotte) Qi <ye.charlotte.qi@gmail.com>
|
2026-08-16 15:08:47 -07:00 |
|
 Lianmin ZhengandLu Fang
|
e49557b8da
|
Support model-defined prefill input embedding width (#35002)
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
|
2026-08-16 15:08:19 -07:00 |
|
 EthanandQAQEthan
|
5534380d46
|
[Spec] Support logprobs with DSpark speculative decoding (#34696)
Co-authored-by: QAQEthan <QAQEthan@users.noreply.github.com>
|
2026-08-16 15:05:13 -07:00 |
|
Lianmin Zheng
|
67e12131df
|
Build Rust extensions on demand in source checkouts (#34994)
|
2026-08-16 14:58:06 -07:00 |
|
Lianmin Zheng
|
0e231d365a
|
Clean up playground scripts and add PR babysitter launcher (#35018)
|
2026-08-16 14:43:57 -07:00 |
|
Liangsheng Yin
|
bae353ba55
|
[misc] Rename shared-read boundary to shared-read ends and fix wrapper delegation (#34982)
|
2026-08-16 14:36:31 -07:00 |
|
Ke Bao
|
ace7314173
|
Add bit-exact guard for extra_buffer_lazy (#35030)
|
2026-08-16 23:34:46 +08:00 |
|
Mick
|
d3589a7251
|
[diffusion] CI: tighten NVIDIA perf baselines (#35016)
|
2026-08-16 20:54:50 +08:00 |
|
Xiaoyu Zhang
|
41abbb0d32
|
[diffusion] Accelerate Cosmos3 T2I QKNorm+RoPE (#34932)
|
2026-08-16 20:15:32 +08:00 |
|
Xiaoyu Zhang
|
095ec6c997
|
[diffusion][kernel] Accelerate Sana BCG with bit-exact conv post-processing (#34928)
|
2026-08-16 20:05:57 +08:00 |
|
Mick
|
3d3194f6c3
|
vlm: cache kimi-k3 per-image processor artifacts (#34404)
|
2026-08-16 19:51:13 +08:00 |
|
Mick
|
968b355f12
|
vlm: streamline vision sdpa reshapes (#34991)
|
2026-08-16 19:05:53 +08:00 |
|
Xiaoyu Zhang
|
0761d3f3a4
|
[diffusion] Accelerate lossless Ideogram norm post-processing (#34931)
|
2026-08-16 17:20:09 +08:00 |
|
Xiaoyu Zhang
|
b752f1e533
|
[diffusion] Enable breakable CUDA graphs for LTX-2.3 (#34929)
|
2026-08-16 17:18:43 +08:00 |
|
Lianmin Zheng
|
6bb73082c8
|
Add skill for babysitting PR CI (#35015)
|
2026-08-16 01:23:32 -07:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
6ab4b99bc2
|
[Quantization] Fix GPTQ scheme attachment broken by LinearBase.scheme default (#34962)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-08-16 00:48:02 -07:00 |
|
Mick
|
2ee0d38a85
|
[diffusion] chore: refresh docs, retire stale knobs, and fix nightly attribution (#34663)
|
2026-08-16 15:41:08 +08:00 |
|
Mick
|
a54de989c8
|
[diffusion] chore: speed up minimax-h3 vae decode on 2×h100 (#34817)
|
2026-08-16 15:38:21 +08:00 |
|
cctry
|
8922bb98e2
|
refactor(hicache): flatten L2 transfer execution (#34793)
GB300 test fails unrelated
|
2026-08-16 00:33:34 -07:00 |
|
 Chenzhou LiandXiaoyu Zhang
|
56a759cffc
|
[JIT Kernel] Migrate moe_topk_softmax from AOT to JIT (#34509)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-16 15:02:57 +08:00 |
|
LinyuanLi
|
0da87024d3
|
[NPU] Add mxfp4-w4a8 MOE Quantization Support for NPU (#30318)
|
2026-08-16 14:03:17 +08:00 |
|
 Zhaoyi Liandjacky.cheng
|
24ab8f9ed9
|
[AMD] Qwen3.5: guard attn layers against empty DP-attention batch (#34474)
Co-authored-by: jacky.cheng <yichiche@amd.com>
|
2026-08-15 21:28:13 -07:00 |
|
 
|
66de161976
|
[Fix][AMD] MoRI EP: drop record_stream in TBO dispatch/combine (HSA out-of-resources) (#32746)
Co-authored-by: billishyahao <bill.he@amd.com>
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
|
2026-08-15 21:23:48 -07:00 |
|