Liangsheng Yin
|
a5c2cc517c
|
[CI] Split the CI control labels into four axes and resolve them live (#40527)
|
2026-09-21 15:37:28 -07:00 |
|
Zhang, Jiejing
|
66f19f5c46
|
[AMD] Enable HiCache for GLM-5.2 MI355X throughput recipe (#40570)
|
2026-09-21 15:28:27 -07:00 |
|
 
|
2261c2e618
|
Add MiMo-V2.6 cookbook (#40622)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-09-21 14:58:53 -07:00 |
|
Faradawn Yang
|
f0940fe3a6
|
Update DeepSeek-V4 Pro for B200 FP4 agentic PD disaggregation (#40610)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
|
2026-09-21 12:10:03 -07:00 |
|
ronnie_zheng
|
f702a0be29
|
Revert "[Diffusion] migrate the whole _register_configs from registry.py to the model own config file" (#40611)
|
2026-09-21 21:05:19 +03:00 |
|
ronnie_zheng
|
e6931ca889
|
[Diffusion] migrate the whole _register_configs from registry.py to the model own config file (#40475)
|
2026-09-21 20:45:53 +03:00 |
|
 
|
50ec9702d0
|
[diffusion] docs: update ComfyUI sections, trimmed examples, and the RTX 5090 DiT-resident recipe (1.42x) for Qwen-Image-2.1 cookbook (#40573)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-21 21:23:11 +08:00 |
|
iridiumine
|
69d1e5cfe0
|
[Docs][NPU] Add MiMo-V2.5-Pro FP4 DFlash best practice on Ascend NPU (#40577)
|
2026-09-21 20:32:52 +08:00 |
|
amote-i
|
b410010087
|
[NPU] [DOC] Add kimi k3 cookbook for 950PR/DT Series (#40575)
|
2026-09-21 20:18:03 +08:00 |
|
    
|
b63f8416b3
|
[Feature] Gigachat 3.5 support (#29189)
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: Viacheslav Barinov <vvadbarinov@sberbank.ru>
Co-authored-by: Viacheslav <viacheslav.teh@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-09-21 14:37:55 +08:00 |
|
 ZiruandNiu Ziru
|
1da8ac10e1
|
Update to the cookbook for XPU-supported models (#33649)
Co-authored-by: Niu Ziru <niuziru@a4bf018d3341.jf.intel.com>
|
2026-09-20 20:26:09 -07:00 |
|
 MickandMick Qian
|
6ad78f2281
|
[diffusion] docs: add verified DGX Spark recipe for Qwen-Image 2.1 (#40487)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-21 08:49:35 +08:00 |
|
WenhaoZhang
|
b912db67ea
|
[diffusion] fix: keep Qwen-Image 2.1 prefix KV per layer under Cache-DiT (#40472)
|
2026-09-21 08:48:31 +08:00 |
|
 MickandMick Qian
|
6880a47955
|
[diffusion] docs: simplify Qwen-Image 2.1 cookbook (#40455)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-20 20:39:26 +08:00 |
|
amote-i
|
5c69e32abe
|
[NPU] [DOC] fix typos, heading levels and terminology in NPU docs (#40402)
|
2026-09-20 15:15:58 +08:00 |
|
 sglang-botandsglang-bot
|
c1a1eb5f66
|
docs: sync LMSYS SGLang blog cards (#40276)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-09-20 06:35:20 +00:00 |
|
 MickandMick Qian
|
031bff5dd3
|
[diffusion] chore: batch qwen-image 2.1 targets and document measured deployment recipes (#40408)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-20 13:59:36 +08:00 |
|
 
|
f9c2791460
|
[diffusion] model: support qwen-image-2.1 (#39983)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-20 09:46:09 +08:00 |
|
![copilot-swe-agent[bot]](/assets/img/avatar_default.png) DefTruthandcopilot-swe-agent[bot]
|
f1fbbd17bb
|
[diffusion] chore: update Cache-DiT to 1.5.1 for DMD Calibrator, SVDQuant DQ, etc (#40104)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
|
2026-09-19 11:14:33 +08:00 |
|
 hmalgewattaandMick Qian
|
c475ac5eaf
|
[diffusion] feat: out of tree platform support (#37547)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-19 09:35:58 +08:00 |
|
Sam (Kesen Li)
|
d346b214fb
|
feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding (#36340)
|
2026-09-18 14:50:46 -07:00 |
|
ChangLiu0709
|
45938a24ae
|
[AMD] GLM-5.2 MI355X MXFP4: bump image to 20260916, use HIP Top-K (#40148)
|
2026-09-18 23:21:58 +08:00 |
|
 MickandMick Qian
|
9784d5f979
|
[diffusion] doc: sync CFG and tracing documentation (#39883)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-18 18:03:40 +08:00 |
|
       
|
a6cf05817f
|
dsv4.1: remaining model and runtime integration (#38798)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-09-18 02:55:30 -07:00 |
|
Ziang Li
|
c46bf5e990
|
[MoE] Disable FlashInfer fused finalize by default for numerical accuracy (#40105)
|
2026-09-18 00:15:57 -07:00 |
|
amote-i
|
6952538980
|
[NPU] [DOC] delete unsupported api --disable-hybrid-swa-memory for npu (#40058)
|
2026-09-18 14:31:05 +08:00 |
|
Xinyuan Tong
|
4dbba37965
|
Verify the Ling-3.0-flash-VL FP4 lane on H200 and disable shared-expert fusion in quant recipes (#39419)
|
2026-09-17 23:20:18 -07:00 |
|
 Chi McIsaacandMick Qian
|
6215aecd51
|
[diffusion] feat: add metrics support (#19084)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-18 13:22:03 +08:00 |
|
Khoa Pham
|
b98a2d1096
|
[Docs] GLM-5.3-Flash cookbook: temporarily remove the DCP option (#40036)
|
2026-09-17 16:17:25 -07:00 |
|
Arseniy Mironov
|
a1b4ec02ae
|
[NPU][Diffusion] FA MXFP8 and modelslim w4a4f8 and w8a8f8 support for Wan2.2 and FLUX (#39438)
|
2026-09-17 13:13:05 +03:00 |
|
 yl3469andShuwen Wang
|
1a90ae6727
|
Add Agentic-Aware Tail-Optimized LRU eviction to the unified radix cache (#34012)
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
|
2026-09-17 17:52:59 +08:00 |
|
 faceless voidandronnie_zheng
|
44bd359082
|
[NPU][Diffusion] Optimize SenseNova-U1 batched generation (#39382)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-09-17 10:08:20 +03:00 |
|
 MickandMick Qian
|
c525ed8f02
|
[diffusion] fix: preserve explicit attention backends during autotune (#39882)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-17 13:28:55 +08:00 |
|
 Liangsheng YinandBBuf
|
464fffbec8
|
dsv4.1: chat encoding and tool parsing (#39665)
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-16 20:49:42 -07:00 |
|
Thomas Wang
|
f60652a43e
|
[AMD] Update deepseek-v4 PDI and cache policy setting for agentic workload (#39702)
|
2026-09-15 21:16:30 -07:00 |
|
 jacky.chengandChangLiu0709
|
a9bb4d7d45
|
[AMD] Align Qwen3.5 MI355X HiCache cookbook with kernel / page_first (#39572)
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
|
2026-09-16 11:22:28 +08:00 |
|
amote-i
|
575163ff32
|
[NPU] [DOC] delete unsupported models in npu docs (#39555)
|
2026-09-15 14:55:17 +08:00 |
|
amote-i
|
bdf8886ad3
|
[NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389)
|
2026-09-15 10:33:14 +08:00 |
|
 
|
a25f213bc4
|
[diffusion] docs: give the RTX 5090 its own H3 recipe, measured on a physical desktop (#39373)
Co-authored-by: Mick Qian <mickqian@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-15 09:15:13 +08:00 |
|
ChangLiu0709
|
242d8a70c0
|
[AMD] GLM-5.2 MI355X MXFP4: bump image to 20260913, enable TOPK_V2 (#39406)
|
2026-09-14 21:50:52 +08:00 |
|
Theresa Shan
|
95140a7b0c
|
[docs] DeepSeek-V4: MI355X PD disaggregation recipes for all three strategies (#39396)
|
2026-09-14 02:11:51 -07:00 |
|
jacky.cheng
|
2f5cc8e33e
|
[AMD] Align Qwen3.5 MI355X cookbook with AttnFP8-V2 and HiCache direct / page_first_direct (#39358)
|
2026-09-14 12:46:17 +08:00 |
|
  
|
42b5af8c62
|
[diffusion] docs: refresh the VDN-H3 on b200 numbers in cookbook (#39244)
Co-authored-by: haochengxi <xihc@berkeley.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
|
2026-09-14 11:50:40 +08:00 |
|
 
|
6220f45d8e
|
[PD] Add /v1/responses support to the HTTP PD router (#36141)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-13 22:48:13 +08:00 |
|
 Shangming CaiandXinyuan Tong
|
14a131ad5b
|
[PD][OpenAI] Gate /v1/responses persistence behind --enable-response-store, default off (#39122)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-13 21:20:47 +08:00 |
|
    
|
cebca698e2
|
[Qwen3.8] Enable NVIDIA NVFP4 on DGX Spark with file-backed PLE and PDL router fix (#39126)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: rdxa <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Manrique <nanomlm@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
|
2026-09-13 16:23:41 +08:00 |
|
Thomas Wang
|
ec5fba5777
|
[AMD] Add dspark config and agentic workload section for deepseek-v4 model (#39252)
|
2026-09-12 19:32:01 -07:00 |
|
 Xiaoyu ZhangandMick Qian
|
23bc4c6ed9
|
[Diffusion] Return Qwen-Image-Layered outputs and preserve CFG2 rounding (#38549)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-13 09:15:47 +08:00 |
|
Zhang, Jiejing
|
288627e400
|
[AMD] Document GLM-5.2 MXFP4 recipe update on MI355X (#39230)
|
2026-09-12 15:57:15 -07:00 |
|
Xinyuan Tong
|
b5a2aebc7e
|
[Docs] GLM-5.3-Flash cookbook: fixed MTP 5/1/6, EP1 + flashinfer_trtllm on Blackwell (#39213)
|
2026-09-12 12:56:57 -07:00 |
|