26 Commits
Author SHA1 Message Date
7aab39a18b [Diffusion] SGLang backend for GLM Image AR. Step 1 - Separate server (#25381)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: yuefeng Wu <33725817+ChefWu551@users.noreply.github.com>
Co-authored-by: wuyuefeng <wuyuefeng@noreply.gitcode.com>
2026-07-09 15:54:50 +03:00
Makcum888eandronnie_zheng 3afc80d781 [diffusion] Fix multi image input for GLM-Image (#26311)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-27 12:33:20 +03:00
Makcum888e 9060509214 [NPU] fix CI (#26390) 2026-05-27 09:45:42 +03:00
Makcum888eandronnie_zheng 0801cc05ed [Diffusion][NPU] Disaggregation diffusion stages support for NPU (#25895)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-25 13:51:25 +03:00
Makcum888e 3b62604cec [Diffusion] Support parallelism for GLM-Image (#25645) 2026-05-19 17:27:21 +03:00
Makcum888eandronnie_zheng 39c720d1b9 [Diffusion][NPU][CI] update perf numbers (#23056)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-04-20 19:34:11 +03:00
Makcum888e e353630b57 [Diffusion] [NPU] Fix multimodal gen CI (#22879) 2026-04-17 04:09:44 +03:00
Makcum888e f4b0e9c64a [diffusion] [NPU] support ring attention on NPU with FA (#21383) 2026-03-30 20:10:55 +03:00
Makcum888e 666b5e4852 [NPU] Update torch and torch_npu version for NPU (#20013) 2026-03-17 21:25:02 +03:00
05950853bc [Diffusion] [NPU] Add CI tests for FLUX (#19001)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-02 20:40:22 +03:00
Makcum888e b1249ac909 [Diffusion] [NPU] [CI] fix CI performance (#19486) 2026-02-27 18:23:02 +03:00
Makcum888e 0217e82a08 [diffusion] Clean code (#19325) 2026-02-25 21:16:03 +03:00
Makcum888e 9bce3b040c [diffusion] [NPU] Update perf baselines (#19227) 2026-02-24 21:15:16 +03:00
Makcum888e d07e8aa4a3 [Diffusion] [NPU] Enable profiler on NPU (#17807) 2026-02-19 15:33:51 +03:00
Makcum888e 2aa0db7d9c [Diffusion] [NPU] Fix CI run (#18921) 2026-02-17 16:54:19 +03:00
Makcum888e 5f81ec1ad5 [Diffusion] Fix get model name when model local path end with "/" (#18918) 2026-02-17 13:19:54 +03:00
Makcum888eandgemini-code-assist[bot] 14c95d255c [Diffusion] [NPU] [Doc] Add NPU documentation for sglang-diffusion (#18894)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-17 10:12:20 +03:00
Makcum888e 49cbb469b4 [NPU] [CI] Enable run multimodal NPU CI when changes only in multimodal_gen (#18523) 2026-02-10 14:53:43 +03:00
00248d85c7 [diffusion] platform: support WAN/FLUX/Qwen-Image/Qwen-Image-edit on Ascend (#13662)
Co-authored-by: dhx98 <haox.dai@gmail.com>
Co-authored-by: DHX98 <haoxiand@andrew.cmu.edu>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: DHX98 <DHX98@noreply.gitcode.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-02-08 10:45:30 +08:00
Makcum888e bba6e38ff8 [NPU] Split pyproject npu from pyproject other (#17641) 2026-01-26 09:45:44 -08:00
Makcum888e 64d809937a revert row from https://github.com/sgl-project/sglang/pull/17584/ (#17701) 2026-01-25 12:17:47 +03:00
Makcum888e d1042e0d62 [Refactore] [CI] Remove redundant CI test runs step 2 (#17584) 2026-01-24 23:39:48 -08:00
Makcum888e 3e968ab369 [Refactor] [CI] Remove redundant CI test runs (#17217) 2026-01-16 09:52:06 -08:00
c2e56dadb2 [Ascend] torch_npu.npu_mrope for MRotaryEmbedding (#10907)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 21:45:54 +08:00
Makcum888e 8e2ac2e628 [NPU] fix pp_size>1 (#12195) 2025-10-30 11:18:36 +08:00
2aaf22c46c Optimization for AscendPagedTokenToKVPoolAllocator (#8293)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: VDV1985 <vladdv85@mail.ru>
2025-08-11 23:06:39 -07:00