Commit Graph
12041 Commits
Author SHA1 Message Date
cen121212 b1e1fe8eee 【NPU】【bugfix】accuracy fix when enable both nsa cp and prefixcache (#23268) 2026-04-28 09:08:28 +08:00
看海的人 9ffc0cc67e [NPU] Support GLM-4.5V (#22961) 2026-04-28 09:08:19 +08:00
PiteXChenandZhiqiang Xie 4a04a9818e 【hicache】Optimize HiCache prefetch logic: adapt to remaining available memory when memory is insufficient (#16370)
Signed-off-by: CLFutureX <chenyongqyl@163.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-04-27 17:18:35 -07:00
Byron Hsuandroot 27659ea8c8 [PD+Pause] Remove redundant post processing (#23886)
Co-authored-by: root <root@slurm-h200-206-011.slurm-compute.tenant-slurm.svc.cluster.local>
2026-04-27 17:01:26 -07:00
cb0429f253 [Disagg] Finalize routed_experts_output in process_batch_result_disagg_prefill (#23885)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 16:40:05 -07:00
Pai Liu 7b9ff79f93 docs: update Python prerequisite to 3.10 (#23801) 2026-04-27 15:36:38 -07:00
Alison Shao b73c44b545 test: relax TestMLADeepseekV3.test_gsm8k threshold 0.62 -> 0.60 (#23879) 2026-04-27 15:27:15 -07:00
Bi Xue 41181b6238 [sgl] copy mm_input in piecewise cuda graph when eagle3 is on (#23613) 2026-04-27 13:35:19 -07:00
Kurt ShusterandEthan Su f34c20af86 [VLM] Fix Kimi-K2.5 CPU path: rename grid_thws -> image_grid_thw (#23501)
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
2026-04-27 20:34:28 +00:00
Vladislav Nosivskoy 28ee08c172 [HiCache] Add synchronization for context parallelism (#20460)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
2026-04-28 02:13:07 +08:00
Xinyuan Tong f34222da1b [Docs] add cookbook for MiMo-V2.5 family (#23851) 2026-04-28 01:38:41 +08:00
Xinyuan Tong 96b0c64c88 Add docs_new code owner (#23855) 2026-04-27 10:31:35 -07:00
Jonah BernardandJonah Bernard 4fc5ebf0b7 [Chore] Remove deadcode in prefill delayer (#23389)
Co-authored-by: Jonah Bernard <96398205+Jonahcb@users.noreply.github.com>
2026-04-27 09:59:36 -07:00
1874. 47b8eadbc4 [Docs] Update Ascend NPU GGUF quantization documentation (#23845) 2026-04-27 17:30:24 +03:00
ranjiewen f2b84b90ac [npu]fix: qwen3-next w8a8 precision bugs (#21698) 2026-04-27 18:14:33 +08:00
Lianmin Zheng 8536d4b402 Clean up noisy startup warnings from third-party deps (#23669) 2026-04-27 03:10:46 -07:00
amote-i 06725ecf0d [NPU] [DOC] Add support new models doc for NPU (#23824) 2026-04-27 17:13:23 +08:00
Shenxiu LiuandXinyuan Tong a3fc982ba7 [Whisper] Automatic language detection via structured generation (#22997)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-04-27 15:54:41 +08:00
zijiexia c2ec64f243 docs: verify GB300 Pro DeepSeek V4 recipes (#23817) 2026-04-27 00:21:26 -07:00
Xiaoyu Zhang 5f47cae1a0 add H100 configs for GLM-4.7-Flash (#23719) 2026-04-27 15:07:39 +08:00
Yujing 9ec4aa1b5c [Doc]Add msprobe doc in docs_new path (#23712) 2026-04-27 10:06:50 +03:00
Colin Z d49561b8ae [AMD] Fix Kimi-K2.6 Quark MXFP4 loading prefix and packed module mapping (#23408) 2026-04-26 23:56:15 -07:00
Praneth Paruchuri b7113cadb1 [Bug Fix] Reject pp_max_micro_batch_size=0 to prevent silent deadlock on generate() (#23799) 2026-04-27 13:36:04 +08:00
Xinyuan Tong e5198386bd Upgrade transformers from 5.5.4 to 5.6.0 (#23525) 2026-04-26 22:33:54 -07:00
Zheng Wengang 91825b8808 [FEAT][EPD] support encoder real health (#23343) 2026-04-27 13:21:28 +08:00
AMD-yanfeiwang 5141d8ae21 [AMD]fix: use CUDA event for targeted draft-to-verify sync in EAGLE overlap (#21940) 2026-04-26 21:58:34 -07:00
Bingxu Chen d84470079d [AMD] Fix Grok-2 nightly: avoid multimodal misdetection from auto-populated vision_config (#23383) 2026-04-26 21:54:36 -07:00
zijiexia c74b345e3a [docs] Update FA4 support SWA (#23793) 2026-04-26 20:40:36 -07:00
Jia GuoandClaude Opus 4.7 bead2e3470 perf: optimize PCG inductor path for FP8 models (redo of #21734) (#23227)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 20:34:27 -07:00
85376a6119 refactor(moe): centralize post-experts all-reduce skip predicate (#23748)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 20:29:59 -07:00
sglang-botandsglang-bot da175b964d chore: update CI test est_time values (#23785)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-04-26 20:17:50 -07:00
iridiumine 32c3513816 [NPU] Support MTP for Qwen3.5 (#20918) 2026-04-27 10:44:17 +08:00
Baizhou ZhangandClaude Opus 4.7 977830e91e ci(deepseek-v4): add b300/grace-blackwell dev-branch build options (#23778)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 16:44:16 -07:00
Kangyan-ZhouandClaude Opus 4.7 35591c7d51 fix(lora): don't assert on non-LoRA lm_head adapter weights (#23433)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 12:10:07 -07:00
zijiexia 3d95ca7546 docs(DeepSeek-V4): mark gb200|big|low-latency verified (#23737) 2026-04-26 11:15:59 -07:00
Mick a392ae8879 [diffusion] feat: accelerate multiple-outputs generation (#23759) 2026-04-27 01:47:33 +08:00
10fd0faccd [CPU] Add Qwen3.5 model optimization for CPU (#19484)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-04-26 10:12:36 -07:00
Liwansi 7d49564431 [NPU]Fix support_triton bug (#23604) 2026-04-26 21:34:56 +08:00
Cheng WanandClaude Opus 4.7 c7878dbb6d [MoE] Deprecate act_and_mul_triton; fold filter_expert into JIT silu/gelu_and_mul (#23707)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 01:41:35 -07:00
Mick d49a0377de [diffusion] refactor: make timestep scheduler request-local (#23716) 2026-04-26 15:59:53 +08:00
9003f24e2b chore: bump sglang-kernel version to 0.4.1.post1 (#23733)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 23:23:49 -07:00
Kangyan-ZhouandClaude Opus 4.7 8efa177f1e [CI] release-pypi-nightly: install protoc before building wheel (#23750)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 22:37:29 -07:00
Kangyan-ZhouandClaude Opus 4.7 4be853e1c6 [CI] release-whl-kernel: strip +cu129 local version before PyPI upload (#23749)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 22:14:45 -07:00
Ethan (Yusheng) Su 3cfd1561df docs(DeepSeek-V4): add h200|big verified recipes + tune H200 Pro parameters (#23742) 2026-04-25 21:44:51 -07:00
Kangyan-ZhouandClaude Opus 4.7 282b47fcfd [CI] release-whl-kernel: clean root-owned build artifacts before checkout (#23747)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 21:39:41 -07:00
Yuhao Yang 049f1bf6fb docs(DeepSeek-V4): add GB200 platform to cookbook recipe (#23725) 2026-04-25 20:54:55 -07:00
ba4e9d2ac2 Apply should_use_dp_reduce_scatterv guard to remaining MoE models (follow-up to #23731) (#23732)
Co-authored-by: Byron Hsu <byronhsu@noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-04-25 20:36:16 -07:00
71029abd64 Fix Qwen3 MoE: also guard EP all-reduce with not use_reduce_scatter (follow-up to #23731) (#23734)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 20:35:52 -07:00
714173555c chore: bump sgl-kernel version to 0.4.1.post1 (#23720)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 17:13:02 -07:00
Byron HsuandByron Hsu 99b59b279c Fix Qwen3 MoE double-reduce when DP attention + EP + reduce_scatterv (#23729) (#23731)
Co-authored-by: Byron Hsu <byronhsu@noreply.github.com>
2026-04-25 15:28:28 -07:00