Commit Graph
13617 Commits
Author SHA1 Message Date
Brayden ZhongandBrayden Zhong d381ec7997 [CI] Fix Nemotron nightly mixed precision checkpoints test (#27284)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-05 13:51:26 -07:00
Brayden ZhongandBrayden Zhong 3b62286fca Reland "Support NextN = 2/4 in DSV32" (#27166)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-05 13:43:28 -07:00
Liangsheng Yin 2c2a4f243a [Bug] Fix EAGLE draft CUDA-graph kv_indices under-allocation for topk > 1 (#27338) 2026-06-05 16:22:04 -04:00
zijiexiaandClaude Opus 4.8 632a3d480e docs: add Tencent Hunyuan and Poolside cards to autoregressive cookbook (#27400)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:17:13 -07:00
Khoa PhamandClaude Opus 4.8 bf172c492a Cookbook for QAT (#27396)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 11:18:11 -07:00
Shangming Cai 57909f731b [PD] Fix KV cache corruption on abort by notifying ongoing prefill (#27372)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-06-06 00:56:13 +08:00
Zhangheng 86b9bf5812 [CI]: Fix CI Stage for 3fs backend test (#27388) 2026-06-06 00:38:03 +08:00
Zhangheng e4a7388187 feat(agentic router, 1/N): Add LoadBasedPolicy (#26480) 2026-06-06 00:13:27 +08:00
xlyandMick d01cf27b7d [diffusion] model: support Ideogram 4 FP8 (#27279)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-05 23:54:48 +08:00
Michael_miaoande01448 e7f94d0d40 [Bug Fix] Fix activation.cuh JIT compilation failure on CUDA 13 due to template type/value mismatch (#26444)
Co-authored-by: e01448 <jwmao@birentech.com>
2026-06-05 22:59:24 +08:00
Zhangheng 6cfdc18585 HiCache: Fix Flaky CI For 3FS Backend (#27358) 2026-06-05 22:00:19 +08:00
Aurick Qiao c06802dc16 Fix customized_info incremental streaming (#27205) 2026-06-05 21:55:01 +08:00
Zhangheng faa6286946 [BugFix]: Fix HiMamba HiCache prefetch hang after L3 sidecar transfer (#27366) 2026-06-05 20:37:40 +08:00
XiaoTianandgongxiaotian e1955bf57a fix(pd): clear stale bootstrap_room when freeing metadata buffer slot (#27374)
Co-authored-by: gongxiaotian <gongxiaotian@didiglobal.com>
2026-06-05 19:46:35 +08:00
Wang, FangYuan 7f919edf00 [AMD] Support alt stream for Qwen3.5 on AMD platform (#25885) 2026-06-05 04:38:25 -07:00
R0CKSTAR 5d691a44f4 [diffusion] fix: fix LingBot World timestep error on MUSA (#27341)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2026-06-05 19:15:08 +08:00
Bingxu Chen 8c8281801d [AMD] update ROCm AITER commit (#27376) 2026-06-05 03:42:16 -07:00
McZyWu d8487bad06 Update best practice for qwen3-next-80b-a3b-instruct (#27353) 2026-06-05 17:02:38 +08:00
Mick 4ef081b903 [diffusion] optimize: optimize LingBot realtime transport and camera conditioning (#27297) 2026-06-05 16:00:48 +08:00
huangtingwei 00fefef16b [PD & HiSparse] Add DeepSeek V4 support for HiSparse direct Prefill-to-Decode DRAM (#24880) 2026-06-05 15:39:48 +08:00
66b932154f [AMD] fix(ci): run partition 3 of stage-c-test-large-8-gpu-amd (#27352)
Co-authored-by: bingxche <bingxche@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-05 15:06:36 +08:00
jianzhao-xu 4248695b07 [NPU] add GLM model best practice docs (#27032) 2026-06-05 14:27:19 +08:00
Zhangheng 4df1ccdadc [UnifiedTree]: Fix CP Reduce For L3 HiCache (#27330) 2026-06-05 14:03:54 +08:00
Qiaolin Yu bd47869ba4 [perf] parallelize create_flashmla_kv_indices over page-blocks (#27320) 2026-06-04 22:11:43 -07:00
6cbc035dc9 FrozenKVMTPVerifyInput: add _draft_preprocess_idle call for when all requests in the verify batch finish in the same iteration (#26859)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Harmya Bhatt <harmyacs@gmail.com>
2026-06-04 21:47:32 -07:00
liuxianglong17 aed0808e18 6-5 nightly failed test case fix (#27335) 2026-06-05 11:39:23 +08:00
zijiexiaandClaude Opus 4.8 c6c1f1a29a docs: sync legacy docs/-only updates into docs_new (Mintlify) (#27308)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 19:45:13 -07:00
Yuhao Yang 46c58b5c70 bench: fix MMMU VLM eval max_tokens for CoT prompt (#27327) 2026-06-05 10:28:16 +08:00
2c8357f794 [XPU] Enable Gemma 4 E2B / E4B / 31B/ 26B-A4B on Intel XPU (#23280)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: jmunetong <jmunetong@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: ckvermaAI <ckverma@habana.ai>
2026-06-05 10:05:07 +08:00
Kangyan-ZhouandClaude Opus 4.8 bcf89928b4 [router] Configure experimental sgl-router via CLI flags instead of a config file (#27073)
Signed-off-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 10:02:10 +08:00
Zhangheng 631db6c757 [UnifiedTree]: Sync sidecar component hits across TP ranks and make SWA prefetch all-or-nothing (#27264) 2026-06-05 09:23:54 +08:00
5af02c18ae [spec_v2] Enable trtllm_mha draft-extend CUDA graph with v2 semantics (#25002)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 17:50:12 -07:00
Cheng WanandClaude Opus 4.8 7dc7376697 fix(attn): delegate init_mha_chunk_metadata in HybridLinearAttnBackend (#27316)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 17:44:23 -07:00
Xinyuan Tong 7425bebb6c docs(cookbook): restore Gemma 4 transformers commit pin (#27321) 2026-06-04 17:43:36 -07:00
sglang-botandsglang-bot bba8de8cb4 docs: sync LMSYS SGLang blog cards (#27322)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-04 17:42:09 -07:00
zijiexiaandClaude Opus 4.8 088f70d0c0 ci: open the LMSYS blog-sync PR with the repo sglang-bot (#27318)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 17:35:33 -07:00
Cheng WanandClaude Opus 4.8 0aa72a9e76 Replace skip_attn_backend_init with a batch-carried attention plan marker (+ staleness re-plan) (#27193)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 17:13:18 -07:00
zijiexiaandClaude Opus 4.8 47377525cb ci: fix LMSYS blog sync to open a PR via gh and only run on main (#27179)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 16:47:46 -07:00
Lianmin Zheng 448d3afb76 fix(spec): complete CustomSpecAlgo duck-typing interface and guard against drift (#27300) 2026-06-04 15:52:46 -07:00
e76d36214b Changes for SM120 perf and usability for NVFP4 (#26496)
Co-authored-by: Martin Vit <martin@voipmonitor.org>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-04 15:29:25 -07:00
Bowen WangandXinyuan Tong 07f326c184 Fix multimodal synthetic benchmark prompt generation to exclude special tokens (#26864)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-04 22:27:43 +00:00
Faradawn Yangandzijiexia 0e4aa081ba Add --enable-symm-mem for Qwen3.5 (#27296)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-04 15:23:37 -07:00
Ziang Li 4cfebbb95f [FlashInfer v0.6.12] Support FlashInfer 4over6 NVFP4 (#25239) 2026-06-04 14:35:07 -07:00
5bf90ad988 Enable DeepGEMM PDL on by default (#23979)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-04 14:13:45 -07:00
Liangsheng Yin cd98d97037 Use level-1 (quiet) busy memory check in chunked-prefill and streaming tests (#27303) 2026-06-04 17:04:27 -04:00
Liangsheng Yin 687cfe9198 Enable runtime busy memory check for speculation topk>1 (#27228) 2026-06-04 16:46:20 -04:00
Ilia Iliev 88a9d513e0 [Quant] Support asymmetric weight quant in compressed-tensors WNA16 (#25292) 2026-06-04 20:15:47 +00:00
YC Yen-Ching Tseng 69623f4b11 [AMD] Guard aiter greedy_sample OOB token id (fixes VLM MMMU CI) (#27247) 2026-06-04 12:53:58 -07:00
Douglas YangandClaude Opus 4.8 75be922451 docs(cookbook): add Docker install option for Gemma 4 (#27287)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 12:45:32 -07:00
shuwenn efe24704b7 [HiCache] feat: truncate mamba prefetch length to available host KV size (#26945) 2026-06-05 03:01:37 +08:00