Commit Graph
14908 Commits
Author SHA1 Message Date
Mohammad Miadh Angkad f32b4ecd26 [Docs] Use trtllm_mha for Qwen3.6 B300 (#29964) 2026-07-07 01:44:01 -07:00
Rita Brugarolas 9ddea8d9ef [AMD] [MORI-EP] Skip LocalExpertCount kernel in decode graph when not recording (#30302) 2026-07-07 01:07:07 -07:00
qinsir5522 5e9032c527 [NPU]Modify LoRA heading in ascend_npu_support_features.mdx to specify Qwen model limitations. (#30358) 2026-07-07 16:01:51 +08:00
linhu-nv 50aa97da45 Feat/flexkv main connector (#29701) 2026-07-07 15:35:52 +08:00
Bingxu ChenandYC Yen-Ching Tseng dabd4cfcfd [AMD] Cap DSV4 Flash max_total_num_tokens (#30313)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-07-07 15:33:05 +08:00
Alison Shao 99db3b0fa5 ci: run jit-kernel tests on scheduled full runs (#30306) 2026-07-07 00:14:39 -07:00
Wang, FangYuan 9a6f8e5992 [AMD] Fix DeepSeek V4 MTP accuracy issue (#30333) 2026-07-06 23:57:57 -07:00
669fd4b8a5 [PP] Fix start_layer_id with pp in get kv_buffer_shape (#29887)
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-07-07 14:54:38 +08:00
fefc1743a9 Cute-DSL FP8 MQA logits (#25220)
Co-authored-by: Mindy Li <11663212+limin2021@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-06 23:34:07 -07:00
ZeyuanChen2000 2d9f0b3317 [NPU] [DOC] Update arguments detail to NPU support features page (#30328) 2026-07-07 14:09:46 +08:00
Xiaoyu ZhangandClaude Opus 4.8 1da7d3a50b [MoE] Retire the AOT moe_fused_gate / kimi_k2_moe_fused_gate gate kernels (#26771) (#29997)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 13:53:17 +08:00
9bd02dc5b9 feat(sgl-kernel): add InfLLM v2 attention kernels (#29383)
Co-authored-by: Size Wang <paulgeorge13hhhhh@gmail.com>
Co-authored-by: lijiayi <lijiayi@modelbest.cn>
Co-authored-by: suhmily10 <suhmily@gmail.com>
Co-authored-by: Xiaoyue Xu <xiaoyue.xu.me@gmail.com>
Co-authored-by: hansjohn <74091612+hansjohn@users.noreply.github.com>
Co-authored-by: zhangyan <1762895426@qq.com>
2026-07-06 22:46:54 -07:00
DarkSharpness be70bfbdbb [DSA] Fold page-table into fused top-k v2 (decode): drop page_size=1 expansion (#30274) 2026-07-06 21:28:05 -07:00
36b449af19 [EPD] Optimize multimodal global cache with paged embedding pool (#28441)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
2026-07-07 12:05:52 +08:00
Zyann 541f9221da feat(metrics): add Prometheus metrics for the EPD encoder server (#27564) 2026-07-07 12:04:51 +08:00
Jialin Ouyang abafe0022e [Unified Radix Cache] Rename tree variables to cache in unittest (#30281) 2026-07-07 11:48:37 +08:00
loading66 998acf7df8 [DOCS][NPU]update npu support features (#30324) 2026-07-07 11:37:19 +08:00
16372b4c5f [Spec] Anchor GLM-5.2 MTP IndexShare topk on the draft-extend step (#29787)
Co-authored-by: kpham-sgl <264503018+kpham-sgl@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 20:36:48 -07:00
Siming Deng 4145e595cf [MLX] Add correctness tests for qwen2_moe and qwen3_moe (#29440) 2026-07-06 20:32:12 -07:00
Siming Dengandsiming-deng df06e03662 [MLX] Size the attention KV pool at the compute dtype for quantized models (#30097)
Co-authored-by: siming-deng <deng_siming@apple.com>
2026-07-06 20:31:59 -07:00
NOOB c3da0a2582 [MLX] Fix single-token chunked-prefill continuation misrouted as decode (#30181) 2026-07-06 20:31:33 -07:00
3a679459e5 [bench] Add agentic-trace multi-turn dataset to bench_serving (#29215)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:45:44 -07:00
e85ef54877 Support Cutedsl BF16 GEMM JIT kernel (#30117)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-07-06 19:16:16 -07:00
Brayden ZhongandBrayden Zhong 267ff1b5f9 Fix LTX2 RoPE JIT kernel CI (#30278)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-07 10:15:21 +08:00
laixinandPeng Zhang 6279805962 [DSv4] Loading Time Weight Dequant (#27867)
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
2026-07-07 09:54:34 +08:00
Yuwei An 3cbb7568bd [Experimental] Full Cuda Graph Support for Prefill (#27988) 2026-07-06 18:13:03 -07:00
Cheng Wan c861896721 [refactor] Resolve config declarations onto server_args at the end of __post_init__ (#30297) 2026-07-06 18:04:44 -07:00
sglang-botandsglang-bot cf4edda956 docs: sync LMSYS SGLang blog cards (#30311)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-07 01:04:21 +00:00
Rahul VijayaraghavanandMa Mingfei 4b5c612257 Skip redundant moe_sum_reduce for single-expert routing on XPU (#22660)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-07 09:03:16 +08:00
Rockdu 7047afafec [diffusion] fix: slice img_shapes per-sample in rollout response extractor (#29989) 2026-07-07 08:56:52 +08:00
Ma Mingfei 30fb0dd851 [CPU] add fused_qk_gemma_norm and refactor norm kernel implementation (#30216) 2026-07-07 08:52:59 +08:00
Mick 6c1fb8a937 [diffusion] fix: fix ragged-caption dynamic-batching accuracy bug in ernie-Image (#30241) 2026-07-07 08:41:51 +08:00
nvjullin b41552334d Fix disagg speculative decoding with NIXL connector (#30222) 2026-07-06 15:51:33 -07:00
Elizaveta MartirosianandElizaveta Martirosian 1c23954cb9 [NPU] Add new diffusion tests (#29331)
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
2026-07-07 00:40:26 +03:00
Aditya Kamat 093256aa4b [Mamba] Fix long-prefill accuracy drop in radix prefix-cache state restore (#29368) 2026-07-06 13:57:44 -07:00
Thomasandronnie_zheng ca73c77055 [Diffusion] cache cross-attn K/V across denoise steps for Helios (#29755)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-06 23:56:37 +03:00
Brayden ZhongandBrayden Zhong d8462f4961 Fixes for NVFP4 numerical accuracy for router GEMM output and wrong correction bias cast (#29783)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-06 13:53:05 -07:00
Lijuan Tang 1579a82d17 [MLX] Fix FakeOverlapScheduler test stub broken by forward_ct accounting (#30125) 2026-07-06 13:43:40 -07:00
Brayden ZhongandBrayden Zhong 8f40b5eb3f When attention TP for linear and full attention, use Flashinfer allreduce fusion (#29699)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-06 13:01:53 -07:00
sushil Dubey 52c6e27e7e [AMD][diffusion] fix: disable layernorm torch.compile decorator in eager mode on ROCm to avoid memory-access fault (#29673) 2026-07-06 11:15:57 -07:00
Junlin Wuandronnie_zheng 3abdbab9bb [llm][npu][quant] Add W4A8 MXFP quantization support for Qwen3 Dense on Ascend NPU (#23650)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-06 19:23:26 +03:00
hhhh1252023 1b481deade feat: sync npu nightly test improvements from Ascend testcases (#29403) 2026-07-06 22:41:15 +08:00
Thomas Wang 80decc78ec [AMD][DeepSeek V4] Set SGLANG_OPT_FLASHMLA_SPARSE_PREFILL to false on hip code path (#30237) 2026-07-06 01:25:55 -07:00
Chetan Kumar VermaandMa Mingfei b3ab56545b Add Accuracy Benchmark for OCR models (#25364)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-06 16:18:43 +08:00
chengcuipingandMohammad Miadh Angkad e2b55bdbab NUMA: probe numactl binding and fall back when --membind is rejected (#28401)
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-07-06 16:03:59 +08:00
zijiexiaandClaude Opus 4.8 7c9bb316cf docs(cookbook): total (input+output) throughput per GPU + percentile latency labels (#30214)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 00:39:48 -07:00
Mick 5f98f62a8a [diffusion] perf: tp-shard every text/image encoder across the full DiT replica (any parallelism) (#30086) 2026-07-06 14:48:07 +08:00
Jzz1943 de00b838c4 [diffusion] fix: pass progressive params through image API (#30148) 2026-07-06 14:45:45 +08:00
McZyWu 6bb2918938 Bugfix qwen prefix cache circumstances (#30047) 2026-07-06 14:39:30 +08:00
Lianmin Zheng 24c42c90be Clean up ServerArgs post-init dispatch (#30186) 2026-07-05 23:05:07 -07:00