Commit Graph
7808 Commits
Author SHA1 Message Date
zijiexiaandClaude Opus 4.8 669be5448b [cuda graph] Enable prefill piecewise CUDA graph for Cohere2Vision (text path) (#28686)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 19:20:14 +00:00
Zhangheng 70883cb1b0 [UnifiedTree]: Rollback mamba hicache test to direct io backend (#28904) 2026-06-22 23:48:48 +08:00
Lianmin Zheng b28e990161 Migrate all ServerArgs fields to Annotated style, reduce add_cli_args by ~2400 lines (#28919) 2026-06-22 08:34:37 -07:00
Xiaoyu Zhang 6b2c730bf7 [codex] Fix DSA indexer in prefill piecewise CUDA graph (#28644) 2026-06-22 22:39:21 +08:00
Xiaoyu Zhang b43bd6824f [B300] Enable FlashInfer allreduce for Qwen3-VL MoE (#28786) 2026-06-22 22:38:45 +08:00
Thomas Wang cee1caaf47 [AMD] Fix nightly-8-gpu-mi35x-deepseek-v4-flash-rocm720 OOM issue (#28941) 2026-06-22 07:35:45 -07:00
Thomas Wang 04d952ea10 [AMD] deepseek-v4 clean env vars (#28920) 2026-06-22 07:32:43 -07:00
Lianmin Zheng ad9723af03 Clean up CUDA graph capture logs (#28937) 2026-06-22 06:15:26 -07:00
shihaozhou 1adb53f147 Fix CP page filtering by request-local position (#28718) 2026-06-22 20:47:29 +08:00
Yuwei An 2ce32366a0 [Fix][BCG][Spec] Restore EAGLE prefill plumbing dropped by #23906 (#28870) 2026-06-22 01:54:54 -07:00
Liangsheng Yinandthanhhao98 106d2930a6 [core] Gate the overlap WAR barrier on forward reads to recover decode throughput (#28363)
Co-authored-by: thanhhao98 <31717833+thanhhao98@users.noreply.github.com>
2026-06-22 00:40:56 -07:00
Lianmin Zheng 886b96621d Migrate more server args to annotated style (#28830) 2026-06-21 20:50:17 -07:00
cctryandcctry 0c065671c9 [Spec] Redo: split init_backends; account draft weights in --mem-fraction-static (#28855)
Co-authored-by: cctry <cctry@fb.com>
2026-06-21 20:45:16 -07:00
Trevor Morris c0bb04b67f [NVIDIA] Support NVFP4 MoE for DeepSeek-V4 (#25820) 2026-06-21 19:35:14 -07:00
6779ca8d7f Fix Qwen MoE precision issue with PP and all-reduce fusion (#28619)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-06-22 08:20:16 +08:00
Liangsheng Yin 4f5ff39bc9 [Spec] Enable FR-Spec in EAGLE draft-extend CUDA graph by sizing logits buffer from the draft head (#28856) 2026-06-21 15:35:57 -07:00
Liangsheng Yin e6722c751b [Feature] Add graceful scheduler shutdown; free hisparse host buffer on exit (#28779) 2026-06-21 15:08:10 -07:00
Liangsheng Yin 8e890391f5 [Spec] Support FlashInfer CUDA graph for EAGLE draft-extend (#28782) 2026-06-21 14:46:25 -07:00
Lianmin Zhengandhnyls2002 a4d0ff3def [misc] Make NaN-logit sanitization opt-in (default off) (#28829)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-06-21 14:31:35 -07:00
7f67965b4d [BugFix] NCCL deadlock in HiCache writing_check by making all_reduce unconditional (#26923)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-06-22 01:47:44 +08:00
iridiumine 5351800700 [Bugfix] Fix MTP acceptance regression on plan stream by moving int64 cast before plan stream context (#28410) 2026-06-22 01:26:58 +08:00
Mohammad Miadh Angkad 643ee748c6 [PP] Pass DSA topk through PP warmup proxy buffers (#28785) 2026-06-21 23:55:39 +08:00
Lianmin Zheng 7942d546d1 Revert "[Spec] Split init_backends; account draft weights in --mem-fraction-static" (#28841) 2026-06-21 07:52:35 -07:00
cctry 9691a29fe0 [Spec] Split init_backends; account draft weights in --mem-fraction-static (#28683) 2026-06-21 01:22:26 -07:00
cctry 6d4ca9bc54 Cap SWA pool sizing with chunk cache (#28755) 2026-06-21 01:06:59 -07:00
Lianmin Zheng c9488241e9 [Refactor] Auto-derive CLI args from dataclass fields to eliminate duplication (#28814) 2026-06-21 00:51:08 -07:00
Jairo David Campaña RoseroandXinyuan Tong b4dda8b3ce fix(anthropic): handle mid-conversation system messages (#26773)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-21 04:34:47 +00:00
Lianmin Zheng 54b9b9d0c9 Remove threading atexit monkey patch (#28812) 2026-06-20 18:30:50 -07:00
karverma-amd 2552b860a3 [AMD][bugfix] Place TBO cuda-graph num_token_non_padded buffer on model devices (#28337) 2026-06-20 18:06:22 -07:00
pure water 5b3eeaf504 [Fix] MM pool GPU alloc with base_gpu_id (#23377) 2026-06-21 08:51:06 +08:00
f42ec350b4 [mtp] add rejection sampling for speculative decoding (#26312)
Co-authored-by: lyc508653 <lyc508653@alibaba-inc.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: Huiqiang Jiang <30883354+iofu728@users.noreply.github.com>
Co-authored-by: Yi Zhang <25844240+yizhang2077@users.noreply.github.com>
Co-authored-by: Yizhong Cao <114661107+cao1zhg@users.noreply.github.com>
2026-06-20 15:10:42 -07:00
Lianmin Zheng fe428dd845 Clean up startup log noise (#28807) 2026-06-20 15:02:52 -07:00
Rita BrugarolasandClaude Opus 4.6 d6d06cdc17 [AMD] Fix no-op dtype cast in _topk_ids_logical_to_physical_dynamic on HIP (#28074)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-20 11:07:22 -07:00
shuwenn ff1fc1fbdf [mem_cache][5/N] refactor: extract host KV cache base layer into pool_host package (#27273) 2026-06-20 20:44:08 +08:00
kkandwunhuang 47cad39f34 [AMD] Optimize o_proj gemm and attn output rope performance (#28722)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-06-20 02:11:01 -07:00
Zhiyao JiangandXinyu Jiang 1115373668 [AMD] Fix garbled unquantized Qwen3-30B-A3B output on ROCm/aiter where the aiter CK fused-MoE falls back to Triton with pre-shuffled weights (#28244)
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
2026-06-20 01:25:01 -07:00
Lianmin ZhengandYinghai Lu 45d203fb08 Fix tokenizer state cleanup on dispatch failure (#28694)
Co-authored-by: Yinghai Lu <yinghai@meta.com>
2026-06-19 21:55:39 -07:00
Bi Xue 28e2096d1c [sgl] wire SGLANG_OPT_SWA_RELEASE_LEAF_LOCK_AFTER_WINDOW on Unified Cache. (#28161) 2026-06-20 12:06:52 +08:00
Jae B.andR0CKSTAR 2cbe1e6404 [Apple Silicon] [MLX] Fix MlxModelRunnerStub.initialize() signature desync with base (#28660)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-06-19 18:05:22 -07:00
Yanbin Jiang 6b945c16f4 [LoRA] Fix experimental fast-path multi-adapter correctness + flashinfer 0.6.12 compatibility (#28091) 2026-06-19 16:20:19 -07:00
Shu Wang c3bae61e16 [Bug] fix(DummyModelLoader): run post_load_weights before process_weights_after_loading (#28665) 2026-06-19 22:57:45 +00:00
3ed46f599f [core] Don't force seq_lens_cpu publication under piecewise CUDA graph (#28633)
Co-authored-by: jonnykong <jonnykong@fb.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-06-19 15:12:07 -07:00
Cheng Wan 6fdcb9934c fix(runner): size eager static buffers for prefill budget and MLP-sync autotune (#28677) 2026-06-19 13:23:26 -07:00
Cheng Wan 2aa7b58aa7 refactor(runner): reuse a prepared static buffer for every dummy run (#28740) 2026-06-19 13:16:07 -07:00
Cheng WanandClaude Opus 4.8 856b0dc74b refactor(runner): move kernel warmup into the shared runner lifecycle (warmup()) (#28739)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 13:13:13 -07:00
Cheng WanandClaude Opus 4.8 d705a91de1 refactor(runner): add EagerRunner, own the eager path, polymorphic dispatch (#28386)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 13:05:43 -07:00
c436a8161a [AMD] Enable HiSparse on ROCm (#26639)
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
Co-authored-by: HAI <hixiao@gmail.com>
2026-06-19 11:59:45 -07:00
Jaybe ca88b7f1d2 fix: remove manual rope parameters injection in PretrainedConfig (#23910) 2026-06-19 17:41:51 +00:00
Mohammad Miadh Angkad 88c261c3f3 Fix IndexCache PP topk handoff (#28532) 2026-06-19 23:39:09 +08:00
Oguz Ulgen 3af991fb3e [AMD] Make breakable CUDA graph run on ROCm/HIP (#28173) 2026-06-19 07:16:00 -07:00