Commit Graph
7240 Commits
Author SHA1 Message Date
R0CKSTAR 02521420b3 [MPS] Support sglang.check_env (#20753)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-03-25 20:59:25 -07:00
gjsheuandgengjinsong d9e96153de [NPU] Support Hybrid KV Cache for Ascend backend (#18032)
Co-authored-by: gengjinsong <gengjinsong@huawei.com>
2026-03-26 11:27:36 +08:00
Simo Lin b835309f0c Reland: compute M-RoPE positions for preprocessed VL inputs (#21244) 2026-03-25 20:12:43 -07:00
DarkSharpness bb29893689 [Fix] Try to fix nvcc compilation error (#21246) 2026-03-26 10:59:36 +08:00
Aurick Qiao a34e9ed64a Add adjusted_filter_batch (#21260) 2026-03-26 10:59:05 +08:00
Aurick Qiao 53c1d8e963 Fix customized_info offset truncation (#21262) 2026-03-26 10:57:51 +08:00
Sam Shleifer 1100b9865c Fix MxInt4 MoE returning wrong output variable (#21348) 2026-03-26 10:57:09 +08:00
Xiaoyu Zhang 6f2b51ade1 [Diffusion] Optimize diffusion Triton rotary embedding by processing multiple heads per token (#21387) 2026-03-26 08:59:25 +08:00
Hubert Lu 7c7b2a8c97 [Bugfix] Lazy-import CuteDSL KDA kernel to fix AMD/ROCm startup crash (#21428) 2026-03-25 16:37:26 -07:00
Liangsheng Yin 75682f1d2f Remove noisy streaming backlog warning log (#21432) 2026-03-25 16:25:16 -07:00
Liangsheng Yin 4dd4e06f1d [CI] Fix resource leak when setUpClass fails (#21338) 2026-03-25 16:22:44 -07:00
Xiaoyu Zhang 68f7f00174 [Diffusion] Speed up Qwen select01 Triton modulation kernels (#21318) 2026-03-25 20:48:39 +08:00
Mick 04eb72801f [diffusion] CI: add performance tracking job to nightly (#21091) 2026-03-25 19:01:33 +08:00
Xiaoyu Zhang 689e9ef05c [Diffusion] Add AKO4ALL kernel optimization skill (#21323) 2026-03-25 18:46:21 +08:00
Xiaoyu Zhang e4ad10520b [diffusion] Skip automatic Wan/MOVA DiT layerwise offload on high-end GPUs (#21248) 2026-03-25 18:45:30 +08:00
DarkSharpness 3d2a61cbf6 [Chore] Clean up JIT compilation flags (#21022) 2026-03-25 18:08:40 +08:00
Liangsheng Yin 4480e6c237 [CI] Add retry loop to killall_sglang GPU cleanup verification (#21393) 2026-03-25 02:16:20 -07:00
YC Yen-Ching Tseng c494e47843 [AMD] Fix stage-b-test-small-1-gpu-amd (test_tool_choice.py) (#19868) 2026-03-25 01:10:21 -07:00
5297a3cb46 [CI] Rewrite killall_sglang as Python with CI/local dual mode (#21331)
Co-authored-by: Alison Shao <alison.shao@mac.lan>
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-03-24 23:54:01 -07:00
Mick 6cc5717e8a [diffusion] doc: update quantization.md (#21356) 2026-03-25 14:48:38 +08:00
17e41cfb21 Fix RDMA device mapping for non-zero GPU indices in disaggregation tests (#21303)
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-03-24 22:56:57 -07:00
Duyi-Wang 61a902ce88 [AMD][MoRI] Auto-select dispatch quantization type from MoE weight dtype. (#21040) 2026-03-24 22:53:57 -07:00
kkandwunhuang 86e2622097 [AMD] Add mha fp8-kv support (#21253)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-03-24 22:38:02 -07:00
Baizhou Zhang 2b75fed0dd Workaround of DSA performance drop on B200 + DP (#21337) 2026-03-24 22:21:07 -07:00
Ke Bao 92492896a5 Fix disaggregation test bootstrap port conflict (#21271) 2026-03-24 21:14:41 -07:00
Ke Bao c1d930c028 Increase flush cache timeout in hicache CI (#21305) 2026-03-24 19:00:59 -07:00
Yuan Luoandluoyuan.luo f273ba1ccc [KDA] Support CuTeDSL KDA decode kernel (#21203)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-03-25 09:47:09 +08:00
DarkSharpness dfc15b78b0 [misc] clean up kernel API (#21325) 2026-03-25 09:10:23 +08:00
281fe10b5e [diffusion] quant: support nvfp4 for Flux.2 (#20137)
Co-authored-by: zcnrex <zcnrex@gmail.com>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Yikang Cai <dcai@catalyst-fleet1.cs.cmu.edu>
Co-authored-by: CHEN Xi <78632976+RubiaCx@users.noreply.github.com>
Co-authored-by: RubiaCx <1084281732@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-03-25 08:28:25 +08:00
Liangsheng Yin 37420dce0b [CI] Enable failfast (-f) by default in run_suite.py (#21330) 2026-03-24 17:04:42 -07:00
Baizhou Zhang 1046dbe038 [Fix] Fix trtllm fp4 moe kernel not found error (#21343) 2026-03-24 16:38:05 -07:00
Mohammad Miadh Angkadandelvischenv bbe25b2412 Use FlashInfer tinygemm for GPT-OSS MoE router on SM90+ (#20755)
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
2026-03-24 15:00:18 -07:00
c4db64c16b Add Lychee Doc Links Check to Local and CI (#19742)
Co-authored-by: Zijie Xia <zijie_xia@icloud.com>
Co-authored-by: Zijie Xia <zijiexia@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-03-24 13:48:26 -07:00
a32e0d57e7 [LoRA][III] Add LoRA support for MoE layers and enable TP (#14105)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-24 13:14:14 -07:00
a3ed2e4d29 [diffusion][CI] Add CI for MOVA model inference (#20430)
Co-authored-by: Luo <139519292+0-693@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-03-24 21:28:16 +03:00
YC Yen-Ching Tsengandbingxche 71f5ae3f9a [AMD] Fix AMD Nightly Test - Transformers 5.3.0 incompatibility and gemma2-27b kv issue (#21193)
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
2026-03-24 10:41:44 -07:00
Elizaveta MartirosianandElizaveta Martirosian 9f4d8ac99f [Diffusion][NPU] Add support for Hunyuan3D (#20352)
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
2026-03-24 16:18:49 +03:00
shadowxz109 1b4933d45d [NPU][ModelSlim] adapt w2 quant layer for Minimax2.5 (#20905) 2026-03-24 20:57:18 +08:00
Aleksi VesantoandMick eefb504f84 [diffusion] model: Fix FLUX.1 output correctness (#21041)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-03-24 15:17:33 +03:00
Mohammad Miadh Angkad 4fbb311234 [Fix][Eval] Keep --dataset-path scoped to longbench_v2 (#21156) 2026-03-24 02:25:11 -07:00
Thomas Wang 855d15adf6 [AMD] Tilelang sparse fwd for dsv32 mi355/mi300 (#19945) 2026-03-24 02:01:39 -07:00
ShunkangzandShunkang dac148167c Enable the qwen3 test (#21195)
Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>
2026-03-23 23:40:59 -07:00
Xiaoyu Zhang 69f02e36e8 [diffusion] Fix torch.zeros typo in causal wan (#21250) 2026-03-24 14:39:16 +08:00
Xiaoyu Zhang d9f97b2115 Refine diffusion skills and align JIT kernel docs with the new CI flow (#21283) 2026-03-24 14:38:36 +08:00
Cheng Wan c01ee848b0 Revert "fix: use consistent time denominator for throughput metrics in bench_one_batch_server" (#21276) 2026-03-23 22:14:54 -07:00
6491728797 [Perf] Overlap NSA-CP key all-gather with query computation for DeepSeek-V3.2 (#20438)
Co-authored-by: Shurui Jia <18817781975@163.com>
Co-authored-by: Baidu-AIAK <baiduaiak~123>
2026-03-23 21:31:48 -07:00
Lianmin Zheng 260abe1fb1 Refactor JIT kernel CI to use run_suite.py registration system (#21239) 2026-03-23 21:17:27 -07:00
0986bed8e2 [HiCache][HybridModel]: Support mamba state offloading & HybridCacheController (#20457)
Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
2026-03-23 20:02:50 -07:00
Ratish PandXiaoyu Zhang 2b1d3c935e [diffusion] fix Z-Image SP sharding for portrait and padded resolutions (#21042)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2026-03-24 10:15:33 +08:00
Yuxuan Zhang fcaad42b00 [Bug Fix] GLM-V / GLM-OCR: field detection for transformers 5.x and MTP omission fix (#21134) 2026-03-23 13:19:48 -07:00