Commit Graph
11824 Commits
Author SHA1 Message Date
Mick fc9de157f9 [diffusion] feat: support overlay model materialization (#21600) 2026-03-28 23:02:38 +08:00
Yuan Luo ee15c104ef [CI] hot-fix ci lint (#21608) 2026-03-28 21:32:39 +08:00
Aditya Sharma 627e162335 [diffusion] fix: fix Flux2-Klein prompt tokenization length to 512 and add regression coverage (#21407) 2026-03-28 17:28:02 +08:00
Baizhou Zhang edd4d54023 [Clean] Remove deprecated environs (#21536) 2026-03-28 00:35:44 -07:00
Jacob0226andClaude Opus 4.6 7078e385ea [AMD] Add GLM-4.7-FP8 accuracy CI test for MI35x (#21534)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 00:28:56 -07:00
Baizhou Zhang 6ef4318ec0 [CI] Move v32 cp test to deepep running suite (#21585) 2026-03-27 22:49:06 -07:00
Liangsheng Yin 402628e560 Patch transformers is_base_mistral in CI to avoid HF 429 rate limiting (#21586) 2026-03-27 22:19:36 -07:00
Kangyan-ZhouandClaude Opus 4.6 33cca495ae [CI] Replace upload/download-artifact with job outputs in release-docker workflow (#21579)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-27 22:12:55 -07:00
Jianyingandjianyingzhu daf02bde33 Fix Piecewise CUDA Graph crash with -enable-mixed-chunk (#20441)
Co-authored-by: jianyingzhu <joeyzhu@nvidia.com>
2026-03-27 21:56:21 -07:00
Liangsheng Yin 19b1f75186 Fix HFRunner hang when subprocess dies during init (#21582) 2026-03-27 21:22:42 -07:00
Yuhao Yang 5ef56682b8 reduce CPU peak memory in multimodal tensor hashing (#21123) 2026-03-28 11:09:16 +08:00
Adarsh ShirawalmathandLianmin Zheng 588320262e Update CODEOWNERS for transformers.py and docs (#21555)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-03-27 20:07:53 -07:00
Fengyuan YuandFengyuan Yu 9fa7b974fd [diffusion] chore: remove redundant identity preprocess_text functions(#20633)
Co-authored-by: Fengyuan Yu <15fengyuan@gmail.com>
2026-03-28 10:07:30 +08:00
Eitan TurokandMick e570ca96f6 [diffusion] refactor: Unify TeaCacheParams and WanTeaCacheParams (#20706)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-03-28 09:51:44 +08:00
Mick f0c68fbefd [diffusion] UX: aggregate expected dtype-cast logs during weight loading (#21552) 2026-03-28 09:50:40 +08:00
Trevor Morris 7160b6cb76 [NVIDIA] Enable automatic NUMA configuration (#19452) 2026-03-27 18:44:13 -07:00
Lianmin Zheng 83997080a6 docs: flesh out MAINTAINER.md oncall lists and link GitHub profiles (#21575) 2026-03-27 17:39:16 -07:00
Vladislav NosivskoyandLianmin Zheng c37200f5e4 Scope streaming backlog coalescing to incremental_streaming_output mode (#21037)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-03-27 17:29:54 -07:00
Qiaolin Yu a27651d5e0 Remove sync when enabling return_logprob (#20972) 2026-03-27 16:36:28 -07:00
zhangxiaolei e2b8463c80 [fix] qwen3.5 fuse_moe_triton_tune bug (#20232) 2026-03-27 19:23:24 -04:00
Ethan (Yusheng) SuandBaizhou Zhang 6d48719e31 [1/n] lora support - Auto detect lora target modules (#21439)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-27 16:08:36 -07:00
narutolhy 9b29131961 fix tp capture in vit cuda graph (#17255) 2026-03-27 22:38:18 +00:00
Baizhou Zhang ec29bbb286 Split workflow for releasing runtime docker (#21563) 2026-03-27 15:05:52 -07:00
Qiaolin Yu 4a41aec844 Fix flaky test_pp_single_node (#21564) 2026-03-27 14:33:46 -07:00
Muqi Liandgemini-code-assist[bot] 38ad251738 feat: add gc_threshold arg (#21481)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-27 13:42:46 -07:00
Baizhou Zhang 4e905febd2 [CI] Relax several thresholds in flaky CIs (#21562) 2026-03-27 13:16:49 -07:00
Lianmin Zheng 9a91323c9f test: point DSV3 int8 MLA CI models to lmsys Hugging Face org (#21561) 2026-03-27 13:04:01 -07:00
huangtingwei d864622a68 [Hicache & JIT_kernel] Support page first layout & mla jit kernel (#18311) 2026-03-27 08:54:36 -07:00
Bi Xue 30397e0a1e [rl][sgl] fix tensor mismatch after pause (#21514) 2026-03-27 23:02:30 +08:00
yang1002378395-cmykandMick 279e7738c5 [diffusion] fix: return None instead of raising RuntimeError when no model info found (#21319)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-03-27 22:42:39 +08:00
Xiaoyu Zhang 9238bd08a2 [CI] Register missing jit_kernel test files (#21547) 2026-03-27 19:39:08 +08:00
Bingxu Chen 6047d2c690 [AMD] Fix AMD CI monitor GitHub API rate limit exhaustion (#21527) 2026-03-27 02:55:56 -07:00
yang1002378395-cmykand阳虎 f83b1b73a8 [diffusion] feat: add --strict-ports option for predictable port assignment (#21320)
Co-authored-by: 阳虎 <yanghu@yanghudeMacBook-Pro.local>
2026-03-27 16:40:50 +08:00
YC Yen-Ching Tseng 448b528720 [AMD] Adjust AMD 4gpu partitions (#21533) 2026-03-27 15:59:26 +08:00
zwang86 5fc5c18bed fix(security): replace unsafe pickle.loads with SafeUnpickler for CVE-2026-3989 (#20904) 2026-03-27 00:43:41 -07:00
Khoa Pham 8d4fca5908 [Security] 1/N: Bind ZMQ sockets to localhost to prevent unauthenticated remote access (#21435) 2026-03-26 23:33:49 -07:00
Xiaoyu Zhang d633ab7349 [Diffusion] Add qknorm rope fuse kernel (#21440) 2026-03-27 14:27:08 +08:00
Xiaoyu Zhang e8d46f145c Opt jit qknorm_across_heads cuda kernel (#21503) 2026-03-27 13:30:46 +08:00
Johnsonms 8a56a7b04d [jit_kernel] Migrate cast (downcast_fp8) from sgl-kernel AOT to JIT (#19103) 2026-03-27 13:21:44 +08:00
JohnsonmsandXiaoyu Zhang c531be455e [jit_kernel] Add fused_qknorm_rope JIT kernel (#19059)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2026-03-27 13:21:28 +08:00
Baizhou Zhang 0138129d3c [CI] Fix nemotron nvfp4 test estimated time (#21516) 2026-03-26 21:53:09 -07:00
Mohammad Miadh Angkad eaf392b9cc Remove redundant DeepSeek V3 FP4 PCG test (#21485) 2026-03-26 21:52:47 -07:00
Shangming Cai 1487f80158 chore: bump mooncake version to 0.3.10 (#20942)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-03-27 10:35:31 +08:00
Mick d7c4c57ace [diffusion] refactor: move format-specific weight loading hooks (quant-related) to a dedicated file (#21366) 2026-03-27 09:58:49 +08:00
Liangsheng Yin e1ee68d0fc Release mm features on session close and support multiple /rerun-ut specs (#21501) 2026-03-26 18:31:29 -07:00
Aurick Qiao c2b3e42ad6 Fix sessions with mm inputs (#21269) 2026-03-26 17:38:23 -07:00
Liangsheng Yin 8a4cdcd538 Simplify flush_cache: reject concurrent requests, remove client-side retry (#21490) 2026-03-26 16:31:04 -07:00
Liangsheng Yin 9dc266adb4 Fix concurrent /rerun-ut posting duplicate workflow URLs (#21495) 2026-03-26 16:26:00 -07:00
Liangsheng Yin c580ddd19d Fix benchmark generating empty prompts when random_input_len is small (#21492) 2026-03-26 16:24:35 -07:00
Baizhou Zhang a93065679b Revert "bugfix for weight loading for qwen3-next" (#21496) 2026-03-26 16:17:18 -07:00