Commit Graph
11962 Commits
Author SHA1 Message Date
jhchouuu f7e840682c [AMD][MoRI] bump MoRI to v1.1.1 (#23642) 2026-04-24 13:12:20 -07:00
Jia GuoandClaude Opus 4.6 587fd15bd2 perf: eliminate attention DtoD copy by passing pre-allocated output to FA (#21985)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-24 12:05:16 -07:00
6d03861476 support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
2026-04-24 12:03:24 -07:00
Lianmin Zheng 6344b546c8 Deprecate --collect-tokens-histogram, auto-collect with --enable-metrics (#23595) 2026-04-24 12:00:16 -07:00
Mick 05696527ea [diffusion] feat: support LoRA for LTX2.3 (#23649) 2026-04-25 01:52:41 +08:00
baa0aa670f [HiCache & HybridModel] 3FS backend support DSA & mamba model (#23241)
Co-authored-by: 墨已 <kangyifei.kyf@alibaba-inc.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-25 00:48:01 +08:00
Kangrui Du 92d262f710 [diffusion] RL: add per-step rollout options for SDE and trajectory capture (#23151) 2026-04-24 23:26:16 +08:00
Siju Samuel bca3dd958a [Intel GPU] Enable pipeline parallelism on XPU (#23645) 2026-04-24 19:52:44 +08:00
Yuwei AnandClaude Opus 4.6 60bbb800db [Experimental] Breakable Piecewise Cuda Graph (#22218)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-24 04:33:05 -07:00
Mick b3b03369a5 [diffusion] fix: unify LTX-2.3 HQ codepath gates for all LTX-2.3 variants (#23624) 2026-04-24 17:44:08 +08:00
YC Yen-Ching Tsengandbingxche b060a5ccfd [AMD] Fix nightly version tag selection (#23644)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-24 17:39:47 +08:00
Shangming Cai b8d883398d Revert "[Intel GPU] Enable pipeline parallelism on XPU" (#23641) 2026-04-24 17:36:35 +08:00
Ziang LiandBrayden Zhong 1758856762 [CI] Fix mxfp8 TrtllmGenMoe test (#23125)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-04-24 09:02:11 +00:00
fzyzcjy 92bb5c6bbe Update pro fp8 checkpoint in DeepSeek V4 cookbook (#23634) 2026-04-24 15:58:04 +08:00
fzyzcjy 3a620cb761 Again update DeepSeek V4 cookbook (#23622) 2026-04-24 15:12:35 +08:00
zijiexia 1a37e57fb1 [codex] docs: note H200 DeepSeek-V4 checkpoint (#23628) 2026-04-24 00:06:30 -07:00
Hubert Lu 4cb0c4e1f3 [AMD] Fix memory access fault when --page-size > 1 with speculative decoding on AMD GPUs (#23596) 2026-04-23 23:56:36 -07:00
Mick cd1fa7506a [diffusion] model: support LTX2.3 high quality pipeline (#23366) 2026-04-24 14:18:20 +08:00
fzyzcjy 734e1e2965 Further update Deepseek V4 docs (#23617) 2026-04-24 13:23:50 +08:00
492883c8ca Add DeepSeek V4 cookbook (#23605)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-04-23 22:10:28 -07:00
YC Yen-Ching Tsengandbingxche 30909cbeeb [AMD] upd local registry address (#23607)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-24 12:07:09 +08:00
Shaojun Zhou 59724e90a9 model: support Moss-VL (#23454) 2026-04-24 11:14:29 +08:00
Siju SamuelandShangming Cai bf98eb3ab7 [Intel GPU] Enable pipeline parallelism on XPU (#23472)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-04-24 10:41:51 +08:00
b35213be11 [MUSA][16/N] Add MUSA backend support for layers and DeepSeek models (V2/V3/R1) (#22774)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-23 18:59:51 -07:00
Zaili Wang cbc2bee547 [Intel CPU/XPU] SGL doc updates (#23547)
merge this one as doc change only.
2026-04-24 09:23:27 +08:00
Ma Mingfei 23e4d381f0 [CPU] remove RECORD_FUNCTION (#23528) 2026-04-24 09:18:29 +08:00
R0CKSTAR 87e50f20f6 [Apple Silicon][MLX] Cache seq_lens-derived tensors in BatchedDecodeContext (#23470)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-04-23 18:12:26 -07:00
MARATRIXandAlex Nails 74c2e5bacd [MUSA][8/N] Port CUDA kernels that are compatible with MUSA (#17946)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-04-23 18:04:58 -07:00
Mick c0166355ae [diffusion] CI: minor refactor CI (#23576) 2026-04-24 08:48:31 +08:00
Cheng Wan d9c72bdd2b Skip unselected experts in flashinfer_trtllm (#23493) 2026-04-23 17:30:19 -07:00
Cheng WanandClaude Opus 4.7 000a2525e1 Move expert_mask_gpu from FusedMoE layer to StandardDispatcher (#23585)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 17:17:27 -07:00
Lianmin Zheng 95d021b523 Pre-set SWA cache location in CudaGraphRunner (#23552) 2026-04-23 16:51:29 -07:00
Lianmin Zheng bb962b0046 Fix MoE no_combine: skip router weight in down projection (#23545) 2026-04-23 16:47:58 -07:00
Sundara Raman Ramachandran cf88fdcc9c Expose child process PIDs from Engine for health check support (#23320) 2026-04-23 16:44:49 -07:00
Sahithi Chigurupati 9891572c4a [CI] Export GB200 nightly logs to S3 (#23502) 2026-04-23 15:35:10 -07:00
sglang-botandsglang-bot f3b88e080a chore: bump flashinfer version to 0.6.8.post1 (#23281)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-04-23 15:23:03 -07:00
Jia Guo 6428392b6f ci: fix cu129 wheel tagging + pipefail-abort in install script (follow-up to #23497) (#23587) 2026-04-23 14:52:58 -07:00
Byron Hsu 17210350fd [PD+DP] Allow PrefillDelayer in disaggregated-prefill mode (#23588) 2026-04-23 14:51:16 -07:00
Kangyan-ZhouandClaude Opus 4.6 2882a136bf [CI] Consolidate Docker release workflows into reusable workflow (#22541)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-23 14:06:01 -07:00
Alex NailsandClaude Opus 4.7 579bd0b152 [bug fix] has_fp8_weights_in_checkpoint: handle HF repo IDs, not just local paths (#23542)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 12:56:49 -07:00
zijiexiaandgemini-code-assist[bot] 9b2f7f8a91 docs: split MI300X and MI325X options in GLM-5.1 generator (#23540)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-23 12:01:58 -07:00
WangHao-hw 80125febb1 [BUGFIX]Fix Ascend backend pre-allocated range in NPU Graph Mode. (#22778) 2026-04-24 01:23:35 +08:00
Jinghong Liandronnie_zheng c6872fc8fb Fix: fallback to torch API when NVML memory query is not supported (#23426)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-04-23 19:26:04 +03:00
Jie Hao 86ed0680d7 feat: add OpenTelemetry tracing to DiffGenerator (#21254) 2026-04-23 09:25:23 -07:00
Arseniy MironovandNapkin-AI 76e4c5a1f8 [Diffusion][NPU][Bugfix] Ascend_fa crashes when sequence parallelism is used. (#23572)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
2026-04-23 19:21:30 +03:00
Baichuanandliubaichuan 54e21bb3a5 [fix] Fix dynamic chunking profiling crash on GLM-5 models (#23060)
Co-authored-by: liubaichuan <liubaichuan@infini-ai.com>
2026-04-23 19:30:57 +08:00
Xinyuan Tong 4868e367f8 docs: add Hunyuan 3 Preview cookbook (#23532) 2026-04-23 02:44:47 -07:00
Xinyi SongandHaiShaw cd459af4e2 [AMD] Use bpreshuffle FP8 blockscale GEMM to replace ABScale GEMM (#23319)
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-04-23 01:51:30 -07:00
Bingxu ChenandYC Yen-Ching Tseng fd88a1c562 [AMD] skip deterministic inference for MLA FP8 test (#23382)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-04-23 00:43:23 -07:00
Goalina 1340d8b3bd [CI][NPU]use rsproxy.cn mirror to speed up Rust toolchain installation on NPU runners (#23514) 2026-04-23 09:52:14 +03:00