Ying Sheng
90bdc3192b
Update sponsorship details in README.md ( #21658 )
2026-03-29 21:42:59 -07:00
db5d9eb8ce
[diffusion] CI: fix dashboard chart (nightly) display issues ( #21653 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-30 12:02:01 +08:00
Feng Su and zhangxiaolei123456
9b4dd27478
[Fix] Fix Qwen3.5 MoE model loading and Mamba cache sharding in PP mode ( #21448 )
...
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com >
2026-03-30 11:57:26 +08:00
Liangsheng Yin and wan4ch
c06ca1526c
Fix circular reference in CustomTestCase.__init_subclass__ ( #21650 )
...
Co-authored-by: wan4ch <wan4ch@gmail.com >
2026-03-29 20:38:12 -07:00
Lianmin Zheng
afb32d7622
README: coding agent sponsorship for long-term contributors ( #21642 )
2026-03-29 16:02:51 -07:00
Lianmin Zheng and Claude Opus 4.6
9f7792415a
Clean up TokenizerManager: remove dead code and improve rid validation ( #21639 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-29 15:12:49 -07:00
Lianmin Zheng and Claude Opus 4.6
f3970b17ef
[Cleanup] Remove unused BatchMultimodalOutput and BatchMultimodalDecodeReq ( #21640 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-29 14:54:25 -07:00
Lianmin Zheng and Claude Opus 4.6
1d9c8e8c9e
Simplify routed experts test and move base64 encoding to tokenizer manager ( #21634 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-29 12:44:01 -07:00
Mohammad Miadh Angkad
2acdda1d85
[Fix] Remove redundant allreduce fusion block and skip TP=1 ( #20621 )
2026-03-29 12:30:40 -07:00
wili
bda94fc779
[Fix] SGLANG_USE_CUDA_IPC_TRANSPORT=1 and SGLANG_ENABLE_MM_SPLITTING=1 do not work at the same time. ( #19915 )
2026-03-30 01:15:26 +08:00
saatwiknagpal
d2440dcf58
[VLM] perf: optimize CUDA IPC for multimodal transfer by caching IPC pool handles ( #21418 )
2026-03-30 00:20:38 +08:00
wili
5bb9ca0e63
[Feature] Optimizations for JPEG input on NVIDIA GPU ( #19749 )
2026-03-30 00:06:14 +08:00
Bi Xue
42c46e6334
[sgl] disable piecewise cuda graph when a model doesn't have layers ( #21565 )
2026-03-29 23:04:20 +08:00
Hanlin Bi
aa9177152e
fix cuda graph capturing error in sm120 mxfp8 triton path ( #19835 )
2026-03-29 01:59:24 -07:00
Liangsheng Yin
fec9961a1f
Clean up _wait_for_scheduler_ready implementation ( #21626 )
2026-03-29 01:02:33 -07:00
shuwenn
c34593f951
[HiCache] fix: graceful shutdown of pending async tasks in bench_mix.py ( #20276 )
2026-03-29 00:46:32 -07:00
d2fa8d67ba
Wrap IPv6 addresses in gRPC, bench_serving, and log messages ( #21236 )
...
Co-authored-by: hnyls2002 <lsyincs@gmail.com >
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
2026-03-29 00:36:31 -07:00
shuwenn
18074e25dc
fix: scheduler launch hang when non-current rank dies ( #20287 )
2026-03-29 00:28:45 -07:00
22e4733ab9
Add subprocess liveness monitor to detect scheduler crashes ( #18582 )
...
Co-authored-by: 继优 <jiyou.ljy@alibaba-inc.com >
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com >
2026-03-29 00:09:13 -07:00
Junrong Lin
35f5a0ff35
[CI] Lossen test_return_routed_experts threshold ( #21270 )
2026-03-28 22:04:53 -07:00
Kangyan-Zhou and Claude Opus 4.6
9d64a82173
feat(ci): add GB300 nightly benchmark test suites ( #21487 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-28 21:54:03 -07:00
Shangming Cai
166e9090ee
[CI] Skip flaky elastic EP test ( #21619 )
2026-03-29 12:50:40 +08:00
Lianmin Zheng and Claude Opus 4.6
ba6b501f3a
Clean up detokenizer and remove dead multimodal_gen code ( #21588 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-28 21:44:40 -07:00
Xiaoyu Zhang and gemini-code-assist[bot]
516cff97a3
[Diffusion] Align diffusion benchmark skill presets with nightly comparison cases ( #21616 )
...
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-29 12:12:17 +08:00
Yuan Luo and luoyuan.luo
343a7ac652
[GDN] Fuse GDN kkt + solve_tril into one kernel ( #21411 )
...
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com >
2026-03-29 12:02:07 +08:00
jacky.cheng and HaiShaw
c86f6c2831
[AMD] Add peft>=0.18.0 to diffusion_hip deps for transformers 5.x compat for AMD diffusion model ( #21442 )
...
Co-authored-by: HaiShaw <hixiao@gmail.com >
2026-03-28 20:28:05 -07:00
Yuhao Yang
4e69f14b95
fix bench_serving sglang backend to support image dataset ( #21294 )
2026-03-29 10:02:11 +08:00
eigen and Avery Huang
3ab9afd653
fix: piecewise_cuda_graph get correct qo_indptr ( #21452 )
...
Co-authored-by: Avery Huang <averyh@nvidia.com >
2026-03-28 15:57:29 -07:00
Shu Wang
efebcab43e
Support skip-softmax attention ( #19089 )
2026-03-28 15:55:48 -07:00
Артем Савкин
40ff652862
Skip ci for .md files ( #21482 )
2026-03-28 13:00:17 -07:00
Артем Савкин and Tamir Baydasov
27071e0a43
[NPU] Update quantization&CI documentation ( #21100 )
...
Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com >
2026-03-28 21:42:21 +03:00
Xinyuan Tong
ced69c9f84
feat: enable CUDA graph and timestamp for the whisper model( #21190 )
2026-03-29 01:46:03 +08:00
Yuhao Yang
57cf4790ca
[VLM] Optimize ShmPointerMMData for multi-pickle safety and deferred unwrap ( #21465 )
2026-03-28 23:11:12 +08:00
Mick
fc9de157f9
[diffusion] feat: support overlay model materialization ( #21600 )
2026-03-28 23:02:38 +08:00
Yuan Luo
ee15c104ef
[CI] hot-fix ci lint ( #21608 )
2026-03-28 21:32:39 +08:00
Aditya Sharma
627e162335
[diffusion] fix: fix Flux2-Klein prompt tokenization length to 512 and add regression coverage ( #21407 )
2026-03-28 17:28:02 +08:00
Baizhou Zhang
edd4d54023
[Clean] Remove deprecated environs ( #21536 )
2026-03-28 00:35:44 -07:00
Jacob0226 and Claude Opus 4.6
7078e385ea
[AMD] Add GLM-4.7-FP8 accuracy CI test for MI35x ( #21534 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-03-28 00:28:56 -07:00
Baizhou Zhang
6ef4318ec0
[CI] Move v32 cp test to deepep running suite ( #21585 )
2026-03-27 22:49:06 -07:00
Liangsheng Yin
402628e560
Patch transformers is_base_mistral in CI to avoid HF 429 rate limiting ( #21586 )
2026-03-27 22:19:36 -07:00
Kangyan-Zhou and Claude Opus 4.6
33cca495ae
[CI] Replace upload/download-artifact with job outputs in release-docker workflow ( #21579 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-27 22:12:55 -07:00
Jianying and jianyingzhu
daf02bde33
Fix Piecewise CUDA Graph crash with -enable-mixed-chunk ( #20441 )
...
Co-authored-by: jianyingzhu <joeyzhu@nvidia.com >
2026-03-27 21:56:21 -07:00
Liangsheng Yin
19b1f75186
Fix HFRunner hang when subprocess dies during init ( #21582 )
2026-03-27 21:22:42 -07:00
Yuhao Yang
5ef56682b8
reduce CPU peak memory in multimodal tensor hashing ( #21123 )
2026-03-28 11:09:16 +08:00
Adarsh Shirawalmath and Lianmin Zheng
588320262e
Update CODEOWNERS for transformers.py and docs ( #21555 )
...
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
2026-03-27 20:07:53 -07:00
Fengyuan Yu and Fengyuan Yu
9fa7b974fd
[diffusion] chore: remove redundant identity preprocess_text functions( #20633 )
...
Co-authored-by: Fengyuan Yu <15fengyuan@gmail.com >
2026-03-28 10:07:30 +08:00
Eitan Turok and Mick
e570ca96f6
[diffusion] refactor: Unify TeaCacheParams and WanTeaCacheParams ( #20706 )
...
Co-authored-by: Mick <mickjagger19@icloud.com >
2026-03-28 09:51:44 +08:00
Mick
f0c68fbefd
[diffusion] UX: aggregate expected dtype-cast logs during weight loading ( #21552 )
2026-03-28 09:50:40 +08:00
Trevor Morris
7160b6cb76
[NVIDIA] Enable automatic NUMA configuration ( #19452 )
2026-03-27 18:44:13 -07:00
Lianmin Zheng
83997080a6
docs: flesh out MAINTAINER.md oncall lists and link GitHub profiles ( #21575 )
2026-03-27 17:39:16 -07:00