Commit Graph
11249 Commits
Author SHA1 Message Date
Ke Bao 9f409d0749 [CI] Adjust CI server launch timeout (#22045) 2026-04-03 22:38:07 +08:00
Xiaoyu Zhang ee9d922f5a Revert "[Kernel] Fuse temperature + softmax in sampling for decode speedup" (#22046) 2026-04-03 21:32:08 +08:00
Baizhou Zhang 97adf8a290 [misc] Add hint for kernel release trigger (#22036) 2026-04-03 03:31:44 -07:00
Baizhou ZhangandClaude Opus 4.6 98ac40192b [Workflow] Fix kernel release build failures for aarch64 and wheel renaming (#22018)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 03:23:03 -07:00
Mick 838f815e9f [diffusion] CI: temporarily disable accuracy ci (#22031) 2026-04-03 17:39:29 +08:00
56ac9c9932 [Fix] Add _MOE_TP to graph_capture for MoE models with ep>1 (#21907)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-03 02:33:16 -07:00
Duyi-Wang ac593fed90 [AMD][Dockerfile] Support build-arg AITER_COMMIT for rocm.Dockerfile (#21949) 2026-04-03 01:54:28 -07:00
Khoa PhamandClaude Opus 4.6 cd75d54fc5 [Bugfix] Fix CUDA graph replay issues in trtllm_mla draft_extend (#21987)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 01:45:13 -07:00
shuwenn 4f84ce5807 [CI] ci: add test_http_server_auth.py to CI (#21866) 2026-04-03 16:32:18 +08:00
monkeyLoveding 658a2813d8 [NPU] Update CI Dependency (#21578) 2026-04-03 16:22:11 +08:00
Michael d07d0a15ce [AMD] Add MiniMax-M2.5 nightly perf benchmarks for MI30x and MI35x (#21524) 2026-04-03 01:01:03 -07:00
Thomas Wang 7431db7392 [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend (#21511) 2026-04-03 00:58:23 -07:00
Kelon ad0516d9c1 [NPU] optimize glm4.7 (#19246) 2026-04-03 15:44:07 +08:00
Shangming Cai d82097a0df [PD] Tiny register info field cleanup for mooncake backend (#22016) 2026-04-03 15:13:44 +08:00
Ricardo-M-L 24f52e66d3 fix: remove duplicate words in comments (#22007) 2026-04-03 00:05:39 -07:00
Liangsheng Yin 4cc970290d [CI] Fix duplicate job names that bypass branch protection (#22001) 2026-04-02 23:59:35 -07:00
Yuzhen Zhou 6b876a7710 [ROCM][RL] Shuffle Weight In-Place to Preserve Parameter Attributes (#21825) 2026-04-02 23:43:55 -07:00
Baizhou Zhang 75de479680 [Misc] Update CI permission (#22014) 2026-04-02 23:37:05 -07:00
4d097047f2 [PD]: Add support for HiSparse to directly transfer the cache from Prefill to Decode DRAM. (#21591)
Co-authored-by: Tingwei Huang <huangtingwei9988@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-04-02 23:06:12 -07:00
Baizhou ZhangandClaude Opus 4.6 5c082c307a [Workflow] Fix kernel release jobs skipped on push events (#22011)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 23:03:13 -07:00
Baizhou Zhang 0a709cfe02 [Workflow] Avoid triggering nightly tests in kernel bump workflow (#22010) 2026-04-02 22:40:33 -07:00
sglang-botandsglang-bot 2c4fb88929 chore: bump sgl-kernel version to 0.4.1 (#21447)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-04-02 22:31:59 -07:00
kkandwunhuang 5bcbc9757c [AMD] Resolve the performance degression when launch server with "--enable-aiter-allreduce-fusion" (#21947)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-04-02 22:10:24 -07:00
DarkSharpness d1b7c3907d [Parallel State Refactor 2/n] Unify code path of AMD deterministic all reduce (#20871) 2026-04-03 12:33:17 +08:00
amote-i 81efcc353a [NPU] Optimized the wording in the npu docs (#21998) 2026-04-03 11:51:40 +08:00
Baizhou Zhang efa7b2d5d3 Revert "[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine)" (#22002) 2026-04-02 20:42:13 -07:00
lviy 5f0df1e2ad [Bugfix] Fix incorrect dp-attention parallel info in bench_one_batch (#21519) 2026-04-02 20:13:53 -07:00
Yuhao Yang 69e89a1fcc [VLM] Enable per-image MM splitting by default and remove MULTI_IMAGES modality (#21899) 2026-04-03 11:04:41 +08:00
narutolhy 8897ac58f0 [PP] qwen3 vl skip layer id for pp (#19135) 2026-04-03 10:51:53 +08:00
Mook 991f3aa5b3 [Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+) (#19652) 2026-04-03 10:48:15 +08:00
Khoa Pham 2b5aed94f5 Remove maxItems=1 restriction when tool_choice is specified (#20208) 2026-04-03 02:35:24 +00:00
Thomasandzhangshuai 0539c62bc1 [Diffusion][NPU] Add support for MOVA (#21633)
Co-authored-by: zhangshuai (S) <z00836796@china.huawei.com>
2026-04-03 05:33:14 +03:00
Kangyan-ZhouandClaude Opus 4.6 1f97714f9b [CI] Add timeouts to Slack upload urlopen and WebClient (#21903)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 19:30:55 -07:00
Xiaoyu Zhang 89affff290 Skip broken AutoModel mapping entries when resolving Llava submodules (#21892) 2026-04-03 09:04:26 +08:00
Baizhou ZhangandClaude Opus 4.6 29d8e959d7 [CI] Remove stale Ascend suite entries from test/srt/run_suite.py (#21978)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 16:47:19 -07:00
Adarsh Shirawalmath 34ddf135fd [Feature] Stronger transformers modeling backend with TP, PP, MoE, VLMs, and torch compile (#19163) 2026-04-02 16:02:33 -07:00
oriandR0CKSTAR 939cf398a9 [MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine) (#17985)
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
2026-04-02 15:04:31 -07:00
Ethan (Yusheng) Su 566b4a4f1c [4/n] Support gpt oss 20b lora (#21570) 2026-04-02 12:57:38 -07:00
Lianmin Zheng fe38410c3e Remove logging for subprocess watchdog start (#21968) 2026-04-02 11:30:33 -07:00
Feng Su 8732b2e9c6 [CI] [Tracing] Add ci for tracing and fix bugs (#21740) 2026-04-02 10:50:50 -07:00
Mick 2278a321ca [diffusion] chore: fix stage profiler for multi-stage denoising (#21955) 2026-04-03 01:16:38 +08:00
DarkSharpness df94cdcebb [Parallel State Refactor 1/n] Remove stream of PyNCCL (#20866) 2026-04-03 00:47:50 +08:00
Ke Bao b21db86e2f [CI] Fix gpu deps import in cpu test (#21950) 2026-04-03 00:06:31 +08:00
Todobe 083304ca44 [NPU] Support GLM-4.7-Flash on NPU (#21408) 2026-04-02 17:44:50 +08:00
Liangsheng YinandDarkSharpness 9d9537fbd3 Migrate ngram corpus from torch cpp_extension to TVM FFI jit_kernel (#21920)
Co-authored-by: DarkSharpness <2040703891@qq.com>
2026-04-02 02:18:11 -07:00
Qiaolin Yu b684b0b72f Fix spec v2 + logprob when max_num_token is set (#20799) 2026-04-02 01:55:16 -07:00
foraxeandyunzhi e55a35fbcd test: add manual init test for mooncake transfer engine (#21842)
Co-authored-by: yunzhi <ningyunxiao.nyx@antgroup.com>
2026-04-02 16:01:10 +08:00
Baizhou Zhang c7d03a6215 Revert "Rollback flashmla to older version [1/2]" (#21922) 2026-04-02 00:27:02 -07:00
Baizhou Zhang fbc1f92453 [DSA] Set trtllm kernels as nsa default for Blackwell (#21914) 2026-04-02 00:22:27 -07:00
Yilong Zhao f30df723bf scheduler: add prefill-only update in merge batch (#21840) 2026-04-01 23:33:06 -07:00