Commit Graph
12076 Commits
Author SHA1 Message Date
Kangyan-ZhouandClaude Opus 4.7 feec1ac7f9 ci: clean up stale-CUDA mooncake variant in install_extra_deps (#23960)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 19:32:38 -07:00
Hubert Lu e5da200d0a [AMD] Fix Aiter RMSNorm layout handling (#23974) 2026-04-28 19:28:46 -07:00
Chandrakant KhandelwalandMa Mingfei 0ac23cffac Add intel_xpu as backend for GptOssForCausalLM, enabled for bf16 models with torch native backend (#12771)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-29 10:21:52 +08:00
iridiumineandiridiumine 08699bb1b2 [NPU] Fix DeepEP LL dispatch BF16 flag and skip triton kernel on NPU for Qwen3.5 (#23815)
Co-authored-by: iridiumine <iridiumine@users.noreply.github.com>
2026-04-29 10:19:24 +08:00
MingxuZhandMa Mingfei 2d27c38f13 Update XPU Docker runtime stack & hf_home config (#23820)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-29 10:03:11 +08:00
Xiaoyu Zhang 13afe8acdf [codex] Enable Qwen3-Next MoE all-reduce fusion (#23619) 2026-04-29 09:11:35 +08:00
Lianmin Zheng 14b4e6fa69 Support --model as alias for --model-path in CLI (#23894) 2026-04-28 17:23:27 -07:00
shuwenn 233048212a [HiCache][SPEC] fix: normalize storage prefetch key (#23631) 2026-04-28 15:53:20 -07:00
zijiexia 387c932dfc [Docs] update Docker image for Nemotron 3 Nano Omni (#23968) 2026-04-28 15:08:34 -07:00
Khoa Pham ddcacaf1bd Fix failing test_nvidia_nemotron_3_nano by fixing test_grouped_topk (#23874) 2026-04-28 15:03:58 -07:00
Alex NailsandClaude Opus 4.7 345fecc547 fix(bench): wire request_func in bench_long_context ContextWorkloadGenerator (#23898)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 14:45:51 -07:00
Liangsheng Yin cf0061da43 [Spec] Fix spec_accept_rate and unify accept/draft naming (#23530) 2026-04-28 14:40:04 -07:00
shuwennandClaude Opus 4.6 3e1c5e1b74 [HiCache][SPEC] fix: empty key after page alignment in match_prefix (#23387)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-28 14:06:55 -07:00
Lewisand百麒 9814cc89ce [Fix] NVFP4 qwen3.5 quant error fix by add packed_modules_mapping (#23471)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
2026-04-28 13:36:09 -07:00
914ef7c7f3 Fix multimodal /v1/embeddings Jinja chat template handling (#20835)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-04-28 13:05:45 -07:00
Qingfu Wen dc1eac4903 [MUSA][Diffusion] Fix fa3 API on MT MUSA (#23646) 2026-04-28 13:01:35 -07:00
Khoa PhamandClaude Opus 4.7 826f2d0620 chore(codeowners): add @kpham-sgl as owner for gemma4 files (#23916)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 11:43:32 -07:00
AlonKejzman 66ea0aee7f tokenizer: Add fastokens support (#23753) 2026-04-28 11:43:10 -07:00
zijiexia ad785a2299 [Docs] add Nemotron 3 Nano Omni cookbook (#23907) 2026-04-28 10:24:40 -07:00
Xinyuan Tong e458a9248f docs: enable MiMo V2.5 MTP cookbook path (#23945) 2026-04-28 10:22:19 -07:00
Xinyuan Tong 3fce8f2009 [Docs] add cookbook for Ling-2.6 family (#23947) 2026-04-29 00:42:04 +08:00
jacky.cheng d95715ec65 [AMD] Fix CI test_diffusion_generation[flux_2_image_t2i_2_gpus] (#23944) 2026-04-28 23:06:29 +08:00
Yuhao Yang 4e1ef6b3cf [Docs] Add single-node H200 DeepSeek-V4-Pro low-latency recipe (#23943) 2026-04-28 23:03:26 +08:00
Mick 144038fbae [diffusion] chore: change default seed to 42 (#23836) 2026-04-28 20:39:23 +08:00
Muqi Li 69a71219cb feat: tiny improve fp8_gemm tune usage (#23912) 2026-04-28 07:47:46 -04:00
Xiaoyu Zhang 7824903417 [SKILL] Sync SGLang skill docs (#23921) 2026-04-28 17:05:36 +08:00
Yinzuo Jiang 71160e4ddb feat(observability): add OpenTelemetry tracing for pipeline parallelism (#23169)
Signed-off-by: Yinzuo Jiang <jiangyinzuo@foxmail.com>
2026-04-28 17:05:23 +08:00
Xun Sun 9a53ab3d6d [6/N] (Elastic EP) Recover failed ranks (#15771) 2026-04-28 00:44:26 -07:00
b8a2dcd300 fix: resolve tensor file overwrite between target and draft models (#21694)
Co-authored-by: jiangguangya <jiangguangya@baidu.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-04-27 23:40:49 -07:00
zijiexia 7ce2352323 [docs] fix sglang serve --model-path in cookbooks (#23905) 2026-04-27 21:54:21 -07:00
Xiaoyu Zhang 6fbad22feb Remove smoke wording from tests and comments (#23355) 2026-04-28 12:05:27 +08:00
JoyFuture 1a55646dcd [Feature] Xiaomi MiMo-V2.5-Pro day0 support (#23808) 2026-04-28 11:43:29 +08:00
Baizhou Zhang c1d1412333 [Flashinfer] Integrate flashinfer router gemm for sm103 (#23285) 2026-04-27 20:37:56 -07:00
3066ba8167 fix(hicache): add retry logic for MooncakeStore warmup (#17195)
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-04-27 18:50:43 -07:00
Baizhou Zhang 3177fa7951 Enable DeepGemm warmup in DeepSeek-V4 cookbook (#23883) 2026-04-27 18:41:12 -07:00
cen121212 b1e1fe8eee 【NPU】【bugfix】accuracy fix when enable both nsa cp and prefixcache (#23268) 2026-04-28 09:08:28 +08:00
看海的人 9ffc0cc67e [NPU] Support GLM-4.5V (#22961) 2026-04-28 09:08:19 +08:00
PiteXChenandZhiqiang Xie 4a04a9818e 【hicache】Optimize HiCache prefetch logic: adapt to remaining available memory when memory is insufficient (#16370)
Signed-off-by: CLFutureX <chenyongqyl@163.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-04-27 17:18:35 -07:00
Byron Hsuandroot 27659ea8c8 [PD+Pause] Remove redundant post processing (#23886)
Co-authored-by: root <root@slurm-h200-206-011.slurm-compute.tenant-slurm.svc.cluster.local>
2026-04-27 17:01:26 -07:00
cb0429f253 [Disagg] Finalize routed_experts_output in process_batch_result_disagg_prefill (#23885)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 16:40:05 -07:00
Pai Liu 7b9ff79f93 docs: update Python prerequisite to 3.10 (#23801) 2026-04-27 15:36:38 -07:00
Alison Shao b73c44b545 test: relax TestMLADeepseekV3.test_gsm8k threshold 0.62 -> 0.60 (#23879) 2026-04-27 15:27:15 -07:00
Bi Xue 41181b6238 [sgl] copy mm_input in piecewise cuda graph when eagle3 is on (#23613) 2026-04-27 13:35:19 -07:00
Kurt ShusterandEthan Su f34c20af86 [VLM] Fix Kimi-K2.5 CPU path: rename grid_thws -> image_grid_thw (#23501)
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
2026-04-27 20:34:28 +00:00
Vladislav Nosivskoy 28ee08c172 [HiCache] Add synchronization for context parallelism (#20460)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
2026-04-28 02:13:07 +08:00
Xinyuan Tong f34222da1b [Docs] add cookbook for MiMo-V2.5 family (#23851) 2026-04-28 01:38:41 +08:00
Xinyuan Tong 96b0c64c88 Add docs_new code owner (#23855) 2026-04-27 10:31:35 -07:00
Jonah BernardandJonah Bernard 4fc5ebf0b7 [Chore] Remove deadcode in prefill delayer (#23389)
Co-authored-by: Jonah Bernard <96398205+Jonahcb@users.noreply.github.com>
2026-04-27 09:59:36 -07:00
1874. 47b8eadbc4 [Docs] Update Ascend NPU GGUF quantization documentation (#23845) 2026-04-27 17:30:24 +03:00
ranjiewen f2b84b90ac [npu]fix: qwen3-next w8a8 precision bugs (#21698) 2026-04-27 18:14:33 +08:00