18694 Commits
Author SHA1 Message Date
Aurick Qiao a34e9ed64a Add adjusted_filter_batch (#21260) 2026-03-26 10:59:05 +08:00
Aurick Qiao 53c1d8e963 Fix customized_info offset truncation (#21262) 2026-03-26 10:57:51 +08:00
Sam Shleifer 1100b9865c Fix MxInt4 MoE returning wrong output variable (#21348) 2026-03-26 10:57:09 +08:00
Yanda Cheng 662635e7a7 fix(sgl-kernel): align wheel METADATA/WHEEL with +cu filename (#21437) 2026-03-25 19:44:50 -07:00
MichaelandBingxu Chen 80389fec00 [AMD] Fix AMD CI: mark /sglang-checkout as git safe.directory in container (#21423)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-03-26 10:18:19 +08:00
Xiaoyu Zhang 6f2b51ade1 [Diffusion] Optimize diffusion Triton rotary embedding by processing multiple heads per token (#21387) 2026-03-26 08:59:25 +08:00
Baizhou Zhang dbe871efdd Rollback flashmla to older version [1/2] (#21430) 2026-03-25 17:49:54 -07:00
Hubert Lu 7c7b2a8c97 [Bugfix] Lazy-import CuteDSL KDA kernel to fix AMD/ROCm startup crash (#21428) 2026-03-25 16:37:26 -07:00
Liangsheng Yin 75682f1d2f Remove noisy streaming backlog warning log (#21432) 2026-03-25 16:25:16 -07:00
Liangsheng Yin 4dd4e06f1d [CI] Fix resource leak when setUpClass fails (#21338) 2026-03-25 16:22:44 -07:00
Liangsheng Yin d5c5683d2b [CI] Use ETag conditional requests in wait-for-jobs and add CI infra to check-changes (#21345) 2026-03-25 14:40:07 -07:00
Minglei Zhu a12fea21ed perf(sgl-kernel): expose get_scheduler_metadata for FA3 decode optimization (#21103) 2026-03-25 13:17:27 -07:00
Baizhou Zhang e90cba715c Revert "[Bugfix] Disable ci for .md files" (#21420) 2026-03-25 12:00:20 -07:00
Артем Савкин f1f13f8966 [Bugfix] Disable ci for .md files (#21410) 2026-03-25 10:19:44 -07:00
Nave Assaf 77872a8d55 Update Nemotron Example docs to include Super v3 and Nano 4B (#21416)
Signed-off-by: Nave Assaf <nassaf@nvidia.com>
2026-03-25 12:03:19 -04:00
Xiaoyu Zhang 68f7f00174 [Diffusion] Speed up Qwen select01 Triton modulation kernels (#21318) 2026-03-25 20:48:39 +08:00
Mick 04eb72801f [diffusion] CI: add performance tracking job to nightly (#21091) 2026-03-25 19:01:33 +08:00
Xiaoyu Zhang 689e9ef05c [Diffusion] Add AKO4ALL kernel optimization skill (#21323) 2026-03-25 18:46:21 +08:00
Xiaoyu Zhang e4ad10520b [diffusion] Skip automatic Wan/MOVA DiT layerwise offload on high-end GPUs (#21248) 2026-03-25 18:45:30 +08:00
DarkSharpness 3d2a61cbf6 [Chore] Clean up JIT compilation flags (#21022) 2026-03-25 18:08:40 +08:00
Liangsheng Yin 4480e6c237 [CI] Add retry loop to killall_sglang GPU cleanup verification (#21393) 2026-03-25 02:16:20 -07:00
YC Yen-Ching Tseng c494e47843 [AMD] Fix stage-b-test-small-1-gpu-amd (test_tool_choice.py) (#19868) 2026-03-25 01:10:21 -07:00
Mick 6425df5c8a [diffusion] doc: consolidate documentation (#21373) 2026-03-25 16:01:32 +08:00
Baizhou Zhang f5c225eeba [CI] Fix TestQwen35WithHiCache (#21371) 2026-03-25 00:04:59 -07:00
amote-i 2d583799eb Update ascend docs (#20846) 2026-03-25 09:58:44 +03:00
5297a3cb46 [CI] Rewrite killall_sglang as Python with CI/local dual mode (#21331)
Co-authored-by: Alison Shao <alison.shao@mac.lan>
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-03-24 23:54:01 -07:00
Mick 6cc5717e8a [diffusion] doc: update quantization.md (#21356) 2026-03-25 14:48:38 +08:00
Shangming Cai 6b0f1e3b43 Update skip condition for TestQwen35PPAccuracy (#21370) 2026-03-25 14:28:42 +08:00
17e41cfb21 Fix RDMA device mapping for non-zero GPU indices in disaggregation tests (#21303)
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-03-24 22:56:57 -07:00
Duyi-Wang 61a902ce88 [AMD][MoRI] Auto-select dispatch quantization type from MoE weight dtype. (#21040) 2026-03-24 22:53:57 -07:00
kkandwunhuang 86e2622097 [AMD] Add mha fp8-kv support (#21253)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-03-24 22:38:02 -07:00
Baizhou Zhang 2b75fed0dd Workaround of DSA performance drop on B200 + DP (#21337) 2026-03-24 22:21:07 -07:00
Liangsheng Yin d937d01fe6 [CI] Fix cancel workflow: use bypass-maintenance label (#21364) 2026-03-24 21:43:00 -07:00
Liangsheng Yin c769398748 [CI] Add include_high_priority checkbox to cancel workflow (#21363) 2026-03-24 21:30:44 -07:00
Ke Bao 92492896a5 Fix disaggregation test bootstrap port conflict (#21271) 2026-03-24 21:14:41 -07:00
Michaelandbingxche f81dec090b [AMD] Fix CI: correct stage-b job dependency names in AMD CI workflows (#21357)
Co-authored-by: bingxche <bingxche@amd.com>
2026-03-25 11:48:47 +08:00
Ke Bao c1d930c028 Increase flush cache timeout in hicache CI (#21305) 2026-03-24 19:00:59 -07:00
Liangsheng Yin cd634abee4 [CI] Reduce session correctness test to 30 turns to fix flakiness (#21349) 2026-03-24 18:56:01 -07:00
Yuan Luoandluoyuan.luo f273ba1ccc [KDA] Support CuTeDSL KDA decode kernel (#21203)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-03-25 09:47:09 +08:00
DarkSharpness dfc15b78b0 [misc] clean up kernel API (#21325) 2026-03-25 09:10:23 +08:00
281fe10b5e [diffusion] quant: support nvfp4 for Flux.2 (#20137)
Co-authored-by: zcnrex <zcnrex@gmail.com>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Yikang Cai <dcai@catalyst-fleet1.cs.cmu.edu>
Co-authored-by: CHEN Xi <78632976+RubiaCx@users.noreply.github.com>
Co-authored-by: RubiaCx <1084281732@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-03-25 08:28:25 +08:00
Liangsheng Yin 37420dce0b [CI] Enable failfast (-f) by default in run_suite.py (#21330) 2026-03-24 17:04:42 -07:00
Liangsheng Yin 3937a07b17 [CI] Add cross-job fast-fail health check (Layer 3) (#21341) 2026-03-24 16:58:56 -07:00
Baizhou Zhang 1046dbe038 [Fix] Fix trtllm fp4 moe kernel not found error (#21343) 2026-03-24 16:38:05 -07:00
Liangsheng Yin 30896bd927 [CI] Remove test partition assignments from CI summary (#21344) 2026-03-24 16:27:47 -07:00
Mohammad Miadh Angkadandelvischenv bbe25b2412 Use FlashInfer tinygemm for GPT-OSS MoE router on SM90+ (#20755)
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
2026-03-24 15:00:18 -07:00
Liangsheng Yin 31c35f1c22 [CI] Skip multimodal CI for doc-only changes (#21334) 2026-03-24 14:07:25 -07:00
Michael 6cb1c2d53d [AMD] Add 4-GPU test suite for MI325 runners (#20294) 2026-03-24 14:04:49 -07:00
c4db64c16b Add Lychee Doc Links Check to Local and CI (#19742)
Co-authored-by: Zijie Xia <zijie_xia@icloud.com>
Co-authored-by: Zijie Xia <zijiexia@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-03-24 13:48:26 -07:00
a32e0d57e7 [LoRA][III] Add LoRA support for MoE layers and enable TP (#14105)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-24 13:14:14 -07:00