Kevin Mi
|
8ac19cc19f
|
[AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbook (#39066)
PR Test (Arm64) / check-changes (push) Successful in 10s
PR Test (NPU) / set-image-config (push) Successful in 1s
PR Test (NPU) / Recommend tests from coverage (push) Skipped
PR Test (NPU) / check-changes (push) Successful in 12s
PR Test (sgl-router) / gate (push) Successful in 8s
PR Test (Xeon) / check-changes (push) Successful in 9s
PR Test (XPU) / check-changes (push) Successful in 16s
pr-test-arm64.yml / pr-gate (push) Successful in 3s
PR Test (Arm64) / pr-gate (push) Successful in 3s
pr-test-npu.yml / pr-gate (push) Successful in 2s
PR Test (NPU) / pr-gate (push) Successful in 2s
PR Test (sgl-router) / tier-1 — lint (push) Failing after 33s
PR Test (sgl-router) / tier-2 — build + test (push) Skipped
PR Test (sgl-router) / tier-3 — docker (placeholder) (push) Skipped
PR Test (sgl-router) / tier-3 — k8s integration (push) Skipped
PR Test (sgl-router) / tier-3 — e2e (push) Skipped
pr-test-xpu.yml / pr-gate (push) Successful in 2s
PR Test (XPU) / pr-gate (push) Successful in 2s
pr-test-xeon.yml / pr-gate (push) Successful in 3s
PR Test (Xeon) / pr-gate (push) Successful in 3s
PR Test (sgl-router) / finish (push) Successful in 1s
Lint / lint (push) Failing after 2m51s
PR Test (NPU) / base-a-test-1-npu-a2 (push) Canceled after 0s
PR Test (NPU) / base-b-test-1-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-4-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-8-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (0) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (1) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-4-npu-a3 (0) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-4-npu-a3 (1) (push) Canceled after 0s
PR Test (NPU) / base-c-test-acc-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-acc-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-perf-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-perf-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / Analyze failure report (push) Canceled after 0s
PR Test (NPU) / setup-covstub (push) Canceled after 0s
PR Test (NPU) / pr-test-npu-finish (push) Canceled after 0s
pr-test-npu.yml / run (${{ fromJson(inputs.partitions).arr }}) (push) Canceled after 0s
PR Test (XPU) / finish (push) Canceled after 0s
PR Test (Arm64) / build-test (push) Canceled after 0s
PR Test (XPU) / stage-a-test-1-gpu-xpu (push) Canceled after 0s
PR Test (XPU) / multimodal-gen-test-1-gpu-xpu (push) Canceled after 0s
PR Test (Xeon) / build-test (gnr, gnr, xeon-gnr, stage-a-tp-test-cpu-intel) (push) Canceled after 0s
PR Test (Xeon) / build-test (spr1, 0, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Canceled after 0s
PR Test (Xeon) / build-test (spr2, 1, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Canceled after 0s
PR Test (Xeon) / build-test (spr3, 2, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Canceled after 0s
|
2026-09-22 06:35:12 +00:00 |
|
 Brayden ZhongandXinyuan Tong
|
a0781f2714
|
[Docs] GLM-5.3/5.3-Flash cookbooks: enable reasoning/tool-call parsers by default via auto (#40497)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-22 13:54:08 +08:00 |
|
 Guangda LiuandGuangda Liu
|
04c0913434
|
[HiSparse] Add MHA hisparse support for MiniMax M3 (#31446)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
|
2026-09-22 13:28:03 +08:00 |
|
Zhang, Jiejing
|
66f19f5c46
|
[AMD] Enable HiCache for GLM-5.2 MI355X throughput recipe (#40570)
|
2026-09-21 15:28:27 -07:00 |
|
 
|
2261c2e618
|
Add MiMo-V2.6 cookbook (#40622)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-09-21 14:58:53 -07:00 |
|
Faradawn Yang
|
f0940fe3a6
|
Update DeepSeek-V4 Pro for B200 FP4 agentic PD disaggregation (#40610)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
|
2026-09-21 12:10:03 -07:00 |
|
amote-i
|
b410010087
|
[NPU] [DOC] Add kimi k3 cookbook for 950PR/DT Series (#40575)
|
2026-09-21 20:18:03 +08:00 |
|
 ZiruandNiu Ziru
|
1da8ac10e1
|
Update to the cookbook for XPU-supported models (#33649)
Co-authored-by: Niu Ziru <niuziru@a4bf018d3341.jf.intel.com>
|
2026-09-20 20:26:09 -07:00 |
|
ChangLiu0709
|
45938a24ae
|
[AMD] GLM-5.2 MI355X MXFP4: bump image to 20260916, use HIP Top-K (#40148)
|
2026-09-18 23:21:58 +08:00 |
|
Xinyuan Tong
|
4dbba37965
|
Verify the Ling-3.0-flash-VL FP4 lane on H200 and disable shared-expert fusion in quant recipes (#39419)
|
2026-09-17 23:20:18 -07:00 |
|
Khoa Pham
|
b98a2d1096
|
[Docs] GLM-5.3-Flash cookbook: temporarily remove the DCP option (#40036)
|
2026-09-17 16:17:25 -07:00 |
|
Thomas Wang
|
f60652a43e
|
[AMD] Update deepseek-v4 PDI and cache policy setting for agentic workload (#39702)
|
2026-09-15 21:16:30 -07:00 |
|
 jacky.chengandChangLiu0709
|
a9bb4d7d45
|
[AMD] Align Qwen3.5 MI355X HiCache cookbook with kernel / page_first (#39572)
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
|
2026-09-16 11:22:28 +08:00 |
|
amote-i
|
bdf8886ad3
|
[NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389)
|
2026-09-15 10:33:14 +08:00 |
|
ChangLiu0709
|
242d8a70c0
|
[AMD] GLM-5.2 MI355X MXFP4: bump image to 20260913, enable TOPK_V2 (#39406)
|
2026-09-14 21:50:52 +08:00 |
|
Theresa Shan
|
95140a7b0c
|
[docs] DeepSeek-V4: MI355X PD disaggregation recipes for all three strategies (#39396)
|
2026-09-14 02:11:51 -07:00 |
|
jacky.cheng
|
2f5cc8e33e
|
[AMD] Align Qwen3.5 MI355X cookbook with AttnFP8-V2 and HiCache direct / page_first_direct (#39358)
|
2026-09-14 12:46:17 +08:00 |
|
Thomas Wang
|
ec5fba5777
|
[AMD] Add dspark config and agentic workload section for deepseek-v4 model (#39252)
|
2026-09-12 19:32:01 -07:00 |
|
Zhang, Jiejing
|
288627e400
|
[AMD] Document GLM-5.2 MXFP4 recipe update on MI355X (#39230)
|
2026-09-12 15:57:15 -07:00 |
|
Xinyuan Tong
|
b5a2aebc7e
|
[Docs] GLM-5.3-Flash cookbook: fixed MTP 5/1/6, EP1 + flashinfer_trtllm on Blackwell (#39213)
|
2026-09-12 12:56:57 -07:00 |
|
ChangLiu0709
|
7bc4eb3740
|
[AMD] Update MI355X MXFP4 HiCache defaults and quick-reduce quantization for Qwen3.5 cookbook (#39104)
|
2026-09-12 00:59:32 -07:00 |
|
 Khoa PhamandClaude Fable 5.1
|
0d08668821
|
[Cookbook] Kimi-K3: keep DCP under HiCache L1+L2 with DSPARK (#39190)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-11 23:14:24 -07:00 |
|
Kevin Mi
|
5f3606c7b2
|
[Cookbook][AMD] Kimi-K3 MI350X/MI355X: pin a ROCm image with the DSPARK graph-capture fix, add measured cell numbers (#39029)
|
2026-09-11 16:10:44 -07:00 |
|
  
|
df6424967a
|
[docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Jiminator <jimmysh341@gmail.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
|
2026-09-11 11:58:30 -07:00 |
|
Zhang, Jiejing
|
d7c284b894
|
[AMD] Use the triton DSA backend for GLM-5.2 MXFP4 on MI355X (#39106)
|
2026-09-11 11:48:50 -07:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
52c191da52
|
[Deps] Retire the CUDA 12 lane (#38404)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-10 16:58:09 -07:00 |
|
 Yuhao YangandClaude Code
|
a37ded1693
|
[Cookbook] DeepSeek-V4.1: add the HiCache L2 knob to the Playground (#38844)
Co-authored-by: Claude Code <noreply@anthropic.com>
|
2026-09-10 18:20:34 +08:00 |
|
 zijiexiaandClaude Opus 5
|
5caafd2118
|
Fix the DeepSeek-V4.1 reasoning example and mark the B300 cells verified (#38839)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-10 01:49:44 -07:00 |
|
Xinyuan Tong
|
9a2f17f41d
|
Add INT4 and FP4 lanes to the Ling-3.0-flash-VL cookbook (#38527)
|
2026-09-10 15:23:50 +08:00 |
|
 Brayden ZhongandBrayden Zhong
|
c0b790cf7f
|
Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-09-10 15:12:01 +08:00 |
|
 zijiexiaandClaude Opus 5
|
69777c4d36
|
Add DeepSeek-V4.1 Flash cookbook (#38802)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-09 23:15:30 -07:00 |
|
Thomas Wang
|
c415f977b8
|
[AMD] Update v4 args for agentic workload (#38677)
|
2026-09-09 21:41:05 -07:00 |
|
William Hu
|
0084030179
|
Add Opt-In for GLM-5.3 Flash breakable prefill CUDA graphs (#38522)
|
2026-09-09 17:02:08 -07:00 |
|
Baizhou Zhang
|
e54ff1efb9
|
[CP V1 Deprecation 5/5] Update prefill CP documentation (#36230)
|
2026-09-08 19:27:49 -07:00 |
|
Cheng Wan
|
db272201a2
|
[Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement (#38375)
|
2026-09-08 16:42:12 -07:00 |
|
Xinyuan Tong
|
afe90a8bc9
|
Point Ling-3.0-flash-VL cookbook install section at the model image (#38539)
|
2026-09-08 10:51:33 -07:00 |
|
Xinyuan Tong
|
482e9f257b
|
Add Ling-3.0-flash-VL cookbook (#38434)
|
2026-09-08 23:00:24 +08:00 |
|
 Kedar PotdarandPo-Han Huang
|
f4bbf12423
|
docs(cookbook): Qwen3.5 FP8 on B200/B300 — trtllm-gen MoE + symm mem (#38374)
Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
|
2026-09-08 09:12:49 +08:00 |
|
 
|
f4b75b5c36
|
docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-07 15:33:03 -07:00 |
|
 zijiexiaandClaude Opus 5
|
e4008de757
|
Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-07 21:30:17 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) faceless voidandgithub-actions[bot]
|
30d0eb2ca9
|
[NPU] Adapt DFlash2 speculative decoding to Ascend NPUs (#35629)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-09-07 09:08:39 +08:00 |
|
 
|
92a4d8b5ee
|
Clean logging under --weight-loader-prefetch-checkpoints (#33930)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-09-04 20:05:53 -07:00 |
|
 zijiexiaandClaude Opus 5
|
3b64169f9d
|
[Cookbook] Kimi-K3: add measured B300 1x8 Unified 8k/1k speed numbers (#37878)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-04 17:44:00 -07:00 |
|
Faradawn Yang
|
c8ba8996c4
|
Update DeepSeek-V4 Pro for B200 FP4 agentic HiCache DSpark (#38026)
|
2026-09-04 11:09:24 -07:00 |
|
Thomas Wang
|
225129fe44
|
[AMD] Update v4 amd cookbook 0903 (#37829)
|
2026-09-03 22:11:13 -07:00 |
|
 Jimmy ShongandClaude Fable 5.1
|
2da5802bfa
|
[Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-03 11:43:01 -07:00 |
|
 kkandwunhuang
|
dd091f43cd
|
[AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2026-09-03 18:44:39 +08:00 |
|
Yash Akhauri
|
02d9b3060a
|
[Docs] Update K2 Horizon MoE model names (#37723)
|
2026-09-02 23:32:44 -07:00 |
|
Yash Akhauri
|
98ef7d8ae6
|
docs: add K2 Horizon cookbook recipes and H200 results (#37655)
|
2026-09-03 11:52:05 +08:00 |
|
Faradawn Yang
|
9c70d22721
|
Update GLM-5.2 NVFP4 B200/B300 for AgentX HiCache (#35368)
|
2026-09-02 14:39:09 -07:00 |
|