Commit Graph
9530 Commits
Author SHA1 Message Date
Peng Xingchen d0e974f40b [NPU] Use use_dsa to dispatch Ascend DSA attention (#28436) 2026-06-17 15:47:12 +08:00
b54f8432ad Batch EAGLE draft/draft-extend replay memcpys via grouped foreach copy (#28465)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-06-17 00:34:21 -07:00
Liangsheng Yin 3bc618485a [Perf] Make latest_output_ids H2D non-blocking in prepare_for_decode (#28491) 2026-06-16 23:46:35 -07:00
Baizhou Zhang 27291118b9 Upgrade sgl-deep-gemm to 0.1.3 (#28402) 2026-06-16 23:06:36 -07:00
Oxana Korzh c01f62e341 [bugfix] guard NVIDIA SM-capability checks with is_cuda() for AMD/ROCm (#28486) 2026-06-16 22:51:42 -07:00
Mick 0f5e14e1d9 [diffusion] fix: use Megatron-style tp for native encoders and dits (#28318) 2026-06-17 13:07:44 +08:00
Rohit Kumar Singhandgithub-actions[bot] 9371062ef3 Fix deep seek ocr2 image processing (#27884)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-17 12:59:18 +08:00
Chengze Fan 66ac385f52 fix(moe): MoRI EP init_mori_op missing BF16 dispatch branch (#28469)
Signed-off-by: Chengze Fan <fancz2002@gmail.com>
2026-06-16 21:02:06 -07:00
Brayden ZhongandBrayden Zhong b8a73bfba0 Call Flashinfer mm_fp8 for per-tensor FP8 GEMMs on SM100 (#28333)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-16 20:50:03 -07:00
Yufeng HeandYufeng He 6c8fdb5b62 [diffusion] fix: fix PicklingError with --backend diffusers on non-T2I models (#21472)
Co-authored-by: Yufeng He <40085740+universeplayer@users.noreply.github.com>
2026-06-17 11:14:25 +08:00
Zilin Zhu f06e2d3d1f Support asymmetric compressed-tensors MoE (#27690) 2026-06-16 19:53:33 -07:00
zhaozx-cnandgithub-actions[bot] 224b1dc775 [NPU]Replace ascend vision attn operator (#25768)
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-17 10:33:37 +08:00
Qeeweew 0ae4740bd1 fix: add missing clamp_limit for CompressedTensorsWNA16MoE (#27328) 2026-06-16 19:29:57 -07:00
Trevor MorrisandClaude Opus 4.7 9c53853ea3 Use pack topk ids triton kernel for flashinfer_trtllm_routed (#25702)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-16 19:27:14 -07:00
Yanbin Jiang 093908d4c0 [LoRA] Fix chunked SGMV (csgmv) CUDA graph segment replay (#28371) 2026-06-16 19:19:07 -07:00
37ef295c78 [AMD] Feat/dp moe reduce scatter (#28216)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Wang, FangYuan <39615225+At1a8@users.noreply.github.com>
2026-06-16 18:12:39 -07:00
jasonjk-park d86a7e7018 Custom spec algorithm can handle server args (#28162) 2026-06-16 17:13:52 -07:00
ashwini rathi 4f9b12c5dd [XPU] Guard tvm_ffi import in dsv4 compress modules under TYPE_CHECKING (#28426) 2026-06-16 15:52:12 -07:00
Hanming Lu a10eee3d80 [Tokenizer] Fix abort racing server crash when large amount of aborts (#28341) 2026-06-16 14:49:35 -07:00
huangtingwei 9b4432fe18 [HiCache]Asymmetric pool support direct backend (#28446) 2026-06-16 13:17:57 -07:00
Jason Mancuso c0a6c3ce66 Fix circular import when sglang.srt.model_executor.runner_backend is imported first (#28002) 2026-06-16 12:04:34 -07:00
799584e173 fix: get_processor fails when --tokenizer-path lacks model config.json (#25643)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-16 20:25:05 +03:00
feliang-gitandxutizhou 92b42c8d8a LPLB: linear-programming load balancer for MoE expert parallelism (#24515)
Co-authored-by: xutizhou <xutingz@nvidia.com>
2026-06-16 10:19:42 -07:00
Zhangheng 78b6a4fabf [UnifiedTree]: Clean up some unused dead code. (#28389) 2026-06-16 21:53:35 +08:00
265202cda2 fix(openai): validate assistant tool call arguments before chat template (#28035)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-06-16 12:21:23 +00:00
Yinghai LuandLianmin Zheng fcca4611fa [CAR] Let custom allreduce support VMM based allocation (#27593)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-16 04:49:05 -07:00
Wang, FangYuanandThomas Wang a362ba9da3 [AMD] Feat: Add prefill context parallel support for deepseek v4 unified kv attention (#27928)
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-06-16 02:00:51 -07:00
jacky.cheng 149fabcca7 [AMD] Fuse sigmoid + mul into single Triton kernel for shared expert gating (#27636) 2026-06-16 01:16:58 -07:00
YC Yen-Ching Tseng 102392df5b [AMD][Fix] Skip EPLB topk remap when global server args are unset (#28404) 2026-06-16 01:13:22 -07:00
Xiaoyu Zhang c5b9106c1a [perf] Use default torch compile mode for Wan2.2 T2V A14B (#28304) 2026-06-16 15:45:00 +08:00
Baizhou Zhang 77f327cb6e [2/n] [CP] Add context parallel strategy abstractions (#27313) 2026-06-16 00:20:04 -07:00
Zhangheng 6c908b3a3a [UnifiedTree]: Replace anonymous tuples with NamedTuples in UnifiedRadixCache (#28375) 2026-06-16 14:54:04 +08:00
Junlin Wuandronnie_zheng 2a8ea70059 [llm][npu][quant] Add W8A8 MXFP8 quantization support for Qwen3 Dense on Ascend NPU (#22352)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-16 09:45:18 +03:00
Michael 72d962be88 [AMD] Fix jit-kernel-unit-test-amd: activation.cuh ROCm build + per_token CUDA-only (R165) (#27947) 2026-06-15 23:44:26 -07:00
Liangsheng Yin 556cf54d47 [Perf] Avoid per-decode-step host sync in min_new_tokens penalty (#28397) 2026-06-15 23:32:28 -07:00
jy-song-hubandMick 637c9f780b [diffusion] fix: add precision consistency layer (#27088)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-16 14:21:01 +08:00
Qiaolin Yuandshuwenn e068355831 [spec decoding] supports step 0 in adaptive spec decoding (updating draft kv cache without draft decoding) (#23994)
Co-authored-by: shuwenn <2508695655@qq.com>
2026-06-15 22:21:26 -07:00
Thomas Wang 800aaefc9e [AMD] Annotate ATOM source for imported v4 unified attention kernels (#28392) 2026-06-15 22:12:12 -07:00
iridiumine 486ec150d0 [NPU] Add NPU fallback for fused Triton gating kernels (#28293) 2026-06-16 11:37:05 +08:00
Shu Wang 32685874f3 Reenable MNNVL backend for FlashInfer allreduce fusion (#23402) 2026-06-15 20:19:15 -07:00
shuwenn b23477af44 bench: infer tokenizer from serving model info (#28195) 2026-06-15 20:03:12 -07:00
zhangxiaolei 063ab89ac1 DeepSeek-V4 Online Compress support MTP (#26471) 2026-06-15 19:56:07 -07:00
huangtingwei b5bcd76a41 [HiCache & JIT Kernel] Refactoring HiCache Write-Back Kernel (#21631) 2026-06-15 19:44:27 -07:00
Mick a4a8a614b1 [diffusion] UX: suppress noisy diffusers torchao warning (#28317) 2026-06-16 10:32:21 +08:00
Mick 01e45762ba [diffusion] feat: use srt custom allreduce for tp groups (#28324) 2026-06-16 09:22:47 +08:00
YAMY b3be2e7402 [dsv4] Pad MLA decode q-heads to 64 (not full n_heads) for FlashMLA head64 kernel (#27954) 2026-06-15 17:18:10 -07:00
Liangsheng Yin 14f6348524 [Fix] Demote OpenAIServingResponses init failure log to one-line WARNING (#28349) 2026-06-15 16:58:33 -07:00
Hanming Lu 81dcb00673 [Spec v2] Use decode kernel for TRT-LLM MHA draft extend (#28241) 2026-06-15 16:57:08 -07:00
Kangyan-ZhouandClaude Fable 5 cad43d3212 [CI] Reclaim leaked /dev/shm segments on server startup (#28089)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-15 16:14:08 -07:00
Jia Guo 4ed698a491 fix(fa3): no NaN embeddings with fa_skip_kv_cache under piecewise CUDA graph (#27343) 2026-06-15 13:46:22 -07:00