Bingxu Chen
|
632919e498
|
[AMD] Fix DeepSeek-R1-MXFP4 accuracy with AITER FP8 (#37762)
|
2026-09-21 10:51:37 -07:00 |
|
ronnie_zheng
|
e6931ca889
|
[Diffusion] migrate the whole _register_configs from registry.py to the model own config file (#40475)
|
2026-09-21 20:45:53 +03:00 |
|
cctry
|
7a6191c4b9
|
Preallocate HiCache MHA staging before post-capture KV sizing (#40256)
|
2026-09-21 10:44:29 -07:00 |
|
cctry
|
7ad55e4386
|
[HiCache] TMA-staged host<->device KV transfer kernel (sm_90+) (#40278)
|
2026-09-21 10:38:23 -07:00 |
|
William Hu
|
0cb37c018c
|
[KDA] Enable ReplaySSM for GLM-5.3 Flash (#40517)
|
2026-09-21 10:34:08 -07:00 |
|
 Eric.Chin.AMDandThomas Wang
|
3c71bb018a
|
[AMD] Enable GLM DSA prefill top-k to the v2 kernel (#37889)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-09-21 10:31:07 -07:00 |
|
  
|
5a6a1bb883
|
[mxfp8-kv] Skip writes to the reserved CUDA-graph padding slot (#35351)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Sam Shleifer <sam@thinkingmachines.ai>
|
2026-09-22 01:21:25 +08:00 |
|
 Kan WuandShangming Cai
|
008470abd8
|
[sgl-router] Bound streaming lifetimes and release guards on idle disconnect (#40391)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-09-21 10:07:07 -07:00 |
|
Liangsheng Yin
|
800613a74b
|
[Test] Split the serving perf tests by topic into basic_perf/ and route their thresholds through a kit (#40505)
|
2026-09-21 10:05:57 -07:00 |
|
Sage
|
14e9c40a72
|
[Observability] Expose python/rust frontend identity in /server_info (#39993)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
|
2026-09-21 23:32:52 +08:00 |
|
 
|
50ec9702d0
|
[diffusion] docs: update ComfyUI sections, trimmed examples, and the RTX 5090 DiT-resident recipe (1.42x) for Qwen-Image-2.1 cookbook (#40573)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-21 21:23:11 +08:00 |
|
iridiumine
|
69d1e5cfe0
|
[Docs][NPU] Add MiMo-V2.5-Pro FP4 DFlash best practice on Ascend NPU (#40577)
|
2026-09-21 20:32:52 +08:00 |
|
amote-i
|
b410010087
|
[NPU] [DOC] Add kimi k3 cookbook for 950PR/DT Series (#40575)
|
2026-09-21 20:18:03 +08:00 |
|
Kan Wu
|
0f6761b54f
|
[sgl-router] Add SGLang-compatible DeepSeek V4.1 Flash rendering (#40532)
|
2026-09-21 18:35:09 +08:00 |
|
 
|
0abb251a20
|
[sgl-router] Match DeepSeek V4 rendering to SGLang (#40530)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-09-21 17:56:41 +08:00 |
|
 Kan WuandClaude Fable 5.1
|
2016f5e7a1
|
[sgl-router] Add Kimi-K3 rendering with SGLang parity (#40390)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-21 17:38:49 +08:00 |
|
Zhaoyi Li
|
b86a30afba
|
[AMD][DI][CI] Add a SPUR cluster profile to AMD DI CI (#40113)
|
2026-09-21 01:56:46 -07:00 |
|
hanwlax
|
8faa2d6731
|
[NPU][Diffusion] Disable loading latency checks in Ascend fixtures (#40544)
|
2026-09-21 16:56:35 +08:00 |
|
 
|
630b1ef322
|
[sgl-router] Fix reorg admission proxy test build after BucketResolver::new (#40537)
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-21 16:42:29 +08:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
8d08dfdab7
|
[Fix] Add gigachat35 to the tool-call and reasoning parser name lists (#40554)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-21 16:34:43 +08:00 |
|
 Kan WuandClaude Fable 5.1
|
a9871012ac
|
[sgl-router] refactor - generalized admission policy definitions (#40271)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-21 16:26:30 +08:00 |
|
Shangming Cai
|
70b5b03e78
|
[NPU][CI] Fix paths-filter negation that makes every PR run the NPU tier (#40549)
|
2026-09-21 16:07:27 +08:00 |
|
ashwini rathi
|
c2f860af1c
|
[ci][xpu] Record device time in the multimodal_gen perf lane (#39956)
|
2026-09-21 15:39:29 +08:00 |
|
    
|
b63f8416b3
|
[Feature] Gigachat 3.5 support (#29189)
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: Viacheslav Barinov <vvadbarinov@sberbank.ru>
Co-authored-by: Viacheslav <viacheslav.teh@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-09-21 14:37:55 +08:00 |
|
kpjeeja
|
b54d5b7c7b
|
disaggregation: Fix FakeKVSender queue accumulation (#28652)
Signed-off-by: KP, Jeeja <jeeja.kp@intel.com>
|
2026-09-21 14:27:05 +08:00 |
|
skyler-apdx
|
f5f3c38aad
|
[Fix] Preserve YaRN scaling when extending rotary caches (#38786)
|
2026-09-21 14:14:05 +08:00 |
|
 AMRUTHA MandMa Mingfei
|
d20cd9d77f
|
[XPU]Enable HiSparse hierarchical sparse KV cache on Intel XPU (#32792)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-09-21 14:07:49 +08:00 |
|
 Kan WuandClaude Fable 5.1
|
fcb080bd40
|
[sgl-router] refactor - cache-aware policy (#40366)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-21 13:53:57 +08:00 |
|
Lianmin Zheng
|
11ecdbf39f
|
Clean up startup logging and streamline log audits (#40526)
|
2026-09-20 22:28:35 -07:00 |
|
ashwini rathi
|
292e3ccc0c
|
[ci][xpu] Re-seed the wan2_1_t2v_1.3b perf baseline on Arc Pro B60 (#39955)
|
2026-09-21 13:04:35 +08:00 |
|
 ZiruandNiu Ziru
|
1da8ac10e1
|
Update to the cookbook for XPU-supported models (#33649)
Co-authored-by: Niu Ziru <niuziru@a4bf018d3341.jf.intel.com>
|
2026-09-20 20:26:09 -07:00 |
|
jianzhao-xu
|
62ba964848
|
Fix: post-load staging regression breaks offload meta/sharded_gpu modes (#38779)
|
2026-09-21 11:19:07 +08:00 |
|
Xueshen Liu
|
ab03a8e7eb
|
[Perf] Fork-safe import: no CUDA context at import time, lighter argument parsing (#40201)
|
2026-09-21 10:48:35 +08:00 |
|
zhaozx-cn
|
176dbcb85d
|
[npu]add chunk gdn kernel and unify ssm state layout for ascend gdn backend (#36187)
|
2026-09-21 09:51:19 +08:00 |
|
 MickandMick Qian
|
6ad78f2281
|
[diffusion] docs: add verified DGX Spark recipe for Qwen-Image 2.1 (#40487)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-21 08:49:35 +08:00 |
|
WenhaoZhang
|
b912db67ea
|
[diffusion] fix: keep Qwen-Image 2.1 prefix KV per layer under Cache-DiT (#40472)
|
2026-09-21 08:48:31 +08:00 |
|
 MickandMick Qian
|
501b7851e4
|
[diffusion] CI: guard E2E/loading latency with runner-aware baselines (#39206)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-21 08:43:22 +08:00 |
|
 MickandMick Qian
|
3a0324fb9b
|
[diffusion] optimization: reduce Qwen-Image 2.1 vae and graph warmup memory (#40481)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-21 08:42:02 +08:00 |
|
Liangsheng Yin
|
76a9065bef
|
[Fix] Raise on undelivered embeddings in send_with_url, fix broken tests (#40502)
|
2026-09-20 17:35:27 -07:00 |
|
 Kan WuandClaude Fable 5.1
|
4027740569
|
[sgl-router] refactor - session-aware policy (#40379)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-20 17:22:36 -07:00 |
|
 Kan WuandClaude Fable 5.1
|
aedda8377e
|
[sgl-router] refactor - layout BucketResolver, Bucket, EngineGroup and implement PowerOfTwo (#40241)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-20 16:57:07 -07:00 |
|
 Kan WuandClaude Fable 5.1
|
4a9dc5c4af
|
[sgl-router] refactor - move policy-required states under src/state (#40272)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-20 16:49:23 -07:00 |
|
Liangsheng Yin
|
acd20a516e
|
[CI] Give the kernel lane a 5090 suite and move kernel-only tests off the general lane (#40496)
|
2026-09-20 16:46:53 -07:00 |
|
 
|
42875bcd2a
|
fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global (#38932)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
|
2026-09-20 19:28:54 -04:00 |
|
 Mohammad Miadh Angkadandmmangkad
|
2fa6b94e34
|
[Perf] Fuse the glm5_next mHC attn->MLP boundary (#39200)
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
|
2026-09-20 16:20:37 -07:00 |
|
ollybbmonster
|
983e643854
|
[Feature] support bf16 MoE router and mxfp4 MoE for MiMo V2 (#40448)
|
2026-09-20 16:13:59 -07:00 |
|
Liangsheng Yin
|
2e2d8a2fda
|
[CI] Drive per-commit stage jobs from a runner table instead of copied job blocks (#40495)
|
2026-09-20 15:43:46 -07:00 |
|
 luoroger37andHank Han
|
d97aed2c90
|
Fix TopK v2 fallback when 16-block cluster capacity is zero (#40163)
Co-authored-by: Hank Han <hanhan7630@outlook.com>
|
2026-09-21 06:41:41 +08:00 |
|
 
|
f31a7bd45c
|
Use pinned memory for asynchronous sampling metadata transfers (#39777)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-09-20 15:31:31 -07:00 |
|
Harmya Bhatt
|
95521da18d
|
[DeepSeek-V4.1] Bound dense prefill indexer memory (#40217)
|
2026-09-20 14:43:19 -07:00 |
|