18694 Commits
Author SHA1 Message Date
57ffc55fb6 feat: [1/2] [DeepEP] Fuse shared expert into MoE dispatch under EP (#20089)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: AichenF <aichenf@nvidia.com>
2026-04-09 01:48:28 -07:00
amote-i 7965573eb4 fix issues for npu docs (#22307) 2026-04-09 16:27:34 +08:00
Liwansi 8ec0934f8f [NPU]add Qwen3-32b and Qwen3-8b low latency md (#22429) 2026-04-09 16:18:34 +08:00
19bbaeb3ee [HiSparse]: Add HiSpares-DSA Model's nightly CI (#22425)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-04-09 01:00:55 -07:00
Sundara Raman Ramachandran a64905a7b8 [CICD] [prefill-only] Consolidate prefill-only model E2E tests (#22405) 2026-04-09 00:54:34 -07:00
Mick 9709192ce9 [diffusion] feat: support FLUX.2-small-decoder (#22414) 2026-04-09 15:53:14 +08:00
Liangsheng YinandKe Bao 8ff01d6841 [Test] Add CPU unit tests for MemoryPoolConfigurator (#22420)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-04-09 00:39:19 -07:00
Liangsheng Yin de441ac6bb [core] Introduce MemoryPoolConfigurator class hierarchy (#22389) 2026-04-09 15:29:19 +08:00
Evgueni Petrov b9c316917b fix AttributeError: 'LazyValue' object has no attribute 'keys' in eplb_manager.py for qwen3 moe (#21822) 2026-04-09 00:13:29 -07:00
Michael ef6bfc1197 [AMD] Add GLM-5.1-FP8 nightly accuracy and performance benchmarks for MI30x and MI35x (#22336) 2026-04-08 22:57:43 -07:00
Nicolas Castet e379befbac Add symmetric debug mode to print stack trace of comm ops with unregistered tensors (#18569) 2026-04-08 22:34:58 -07:00
Bingxu Chen 6b96f8341d [AMD] Fix multimodal diffusion test crash on ROCm by falling back to SDPA (#22335) 2026-04-08 22:32:49 -07:00
Xiaoyu Zhang 30b738d3a6 [SKILL] add torch profiler analysis workflow (#22353) 2026-04-09 12:53:48 +08:00
Liangsheng Yin edfddda192 Move runai model loader test to nightly suite (#22418) 2026-04-08 21:39:32 -07:00
Mick 355fcbcc17 [diffusion] fix: fix cache dit refresh none mask (#22374) 2026-04-09 11:58:24 +08:00
jsheng_LinkedinandClaude Opus 4.6 6838a23226 [Feature] Add token embedding overrides for sparse embedding replacement (#20960)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 20:51:36 -07:00
Kurkur a69be2e866 [Feature] Support eagle3 for qwen3-vl (#22230) 2026-04-09 11:45:36 +08:00
Lianmin Zheng ddc8ef1038 Lazy import flash_attention_v4 to avoid loading flash_attn.cute at startup (#22306) 2026-04-08 20:40:25 -07:00
Khoa Pham f127d67823 [Spec][Ngram] Misc enhance support for multiple SAMs (#22294) 2026-04-08 19:56:23 -07:00
tfhddd c431b11d8b [CI] Use UV to improve pip install speed (#22029) 2026-04-09 09:18:32 +08:00
Kangrui DuandMikukuOvO 1b7c33a5b7 [diffusion] rl: revamp rollout Log-Prob support with SDE/CPS for RL post-training (#21204)
Co-authored-by: MikukuOvO <mikukuovo@gmail.com>
2026-04-09 09:00:00 +08:00
Liangsheng Yin 2c4e113dd7 [CI] Fast-fail on lint check failure in check-stage-health (#22400) 2026-04-08 17:17:07 -07:00
Kangyan-ZhouandClaude Opus 4.6 46c2b77627 [CI] Add GLM-5.1 nightly tests and update Qwen3.5 model (#22399)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 17:04:57 -07:00
Alison ShaoandAlison Shao cf27b11498 [CI] Increase stage-c-test-4-gpu-b200 partitions from 4 to 5 (#22395)
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
2026-04-08 16:36:27 -07:00
Alison ShaoandAlison Shao 614dd1d4f5 [CI] Add alexnails to CI_PERMISSIONS.json (#22391)
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
2026-04-08 16:22:05 -07:00
Kangyan-ZhouandClaude Opus 4.6 cc8ea08b8b [CI] Replace upload/download-artifact with job outputs in release-docker-runtime (#22388)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 16:13:33 -07:00
Liangsheng Yin 1e3f6ebea6 [core] Extract pool sizing logic to pool_configurator.py (#22384) 2026-04-08 16:13:21 -07:00
e41647f52b [CI] Add pre-commit hook to validate test/registered/ files have CI registry (#22308)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
2026-04-08 15:59:15 -07:00
Baizhou Zhang 5e1a9834f9 Update ci permission (#22387) 2026-04-08 15:42:06 -07:00
Baizhou ZhangandClaude Opus 4.6 4e5b8cb041 Fix get_version_tag.py to handle dot-separated post versions (#22385)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 15:18:22 -07:00
sglang-botandsglang-bot df3275bd6c chore: bump flashinfer version to 0.6.7.post3 (#22382)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-04-08 14:49:45 -07:00
c89afaea7c Fix hybrid_linear_attn_backend crash with ngram speculation (#20739)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 12:52:07 -07:00
YAMY c26b8b4a4b [GDN] Remove FlashInfer GDN decode + no_buffer guard and default to FlashInfer on SM100+ (#21861) 2026-04-08 11:59:15 -07:00
Kurt Shuster db30a63a13 [sgl-kernel] support > 1024 experts in moe_align_block_size kernel (#21610) 2026-04-08 11:45:13 -07:00
Mick 4ac6fa0d87 [diffusion] fix: fix loading multiple ckpts with different precision for a same module (#22360) 2026-04-09 02:44:19 +08:00
Yihao Wang a5ed507a16 [refactor] [asr] add transcription adapter for extensible ASR models support (#22181) 2026-04-09 01:19:37 +08:00
Yihao Wang ae8da14ea3 [fix] [whisper] ensure inputs are moved to the correct device before processing. (#22293) 2026-04-08 23:45:42 +08:00
Xiaoyu ZhangandMick b5b2dbe05f [Diffusion] Add diffusion NVFP4 scaled-mm correctness test (#22127)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-04-08 22:07:24 +08:00
Xiaoyu Zhang ea119adc90 Refactor auto benchmark unit tests and fix CI bug (#22270) 2026-04-08 21:54:41 +08:00
zhaozx-cn 33c9cc8994 [NPU] fix qwen3.5 video processor (#22266) 2026-04-08 21:13:29 +08:00
Alex NailsandClaude Opus 4.6 931dbceadc [CI] Set RUNAI_STREAMER_MEMORY_LIMIT=0 for stage-b-test-1-gpu-small (#22346)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 02:23:35 -07:00
Alison ShaoandAlison Shao 2ad5e6df12 [CI] Relax gpt-oss 4GPU accuracy threshold from 0.60 to 0.58 (#22237)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
2026-04-08 02:20:23 -07:00
Fergus 413913763f fix: wrap _import_static_state in inference_mode to fix resume on Blackwell (#21035) 2026-04-08 02:03:39 -07:00
Vladislav Nosivskoyandhzh0425 79c82c5c42 [HiCache] Fix write_backup return type when parent not backed up (#22185)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-08 16:42:57 +08:00
Sundara Raman Ramachandran 712c8c5051 [Score API] Add SequenceClassification Model support (#22118) 2026-04-08 01:30:58 -07:00
Baizhou Zhang 213af1d4f7 Add CI tests for GLM-5 (#22285) 2026-04-08 01:05:36 -07:00
HuangJi c3c13dd5e3 [diffusion] fix: make warmup image initialization rank-safe (#21817) 2026-04-08 15:51:09 +08:00
Bingxu Chenandbingxche de0cfed159 [AMD] Fix DLPack Error in Aiter flydsl GEMM by Detaching MoE Gate Weight (#22262)
Co-authored-by: bingxche <binxche@amd.com>
2026-04-07 23:42:10 -07:00
Артем Савкин cd373667cd [Bugfix] [NPU] Qwen3.5 with quantization fix (#21692) 2026-04-08 09:15:48 +03:00
Michael db60a620db [AMD] Add GLM-5-FP8 nightly performance benchmarks for MI30x and MI35x (#21710) 2026-04-07 22:43:14 -07:00