 Alex NailsandClaude Opus 4.7
|
3d2e7cc601
|
[gRPC] Native server: launcher + HTTP + server args wiring (3/4) (#23508)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-07-07 14:57:25 -07:00 |
|
Qiaolin Yu
|
801571e949
|
[spec decoding] support rejection sampling in multi layer eagle (#30303)
|
2026-07-07 14:50:24 -07:00 |
|
Michael
|
60f502a4fd
|
[AMD] Register 2 hardware-agnostic 1-GPU PR tests for AMD CI (#30207)
|
2026-07-07 14:46:22 -07:00 |
|
Michael
|
090efa27a2
|
[AMD] Register 5 CI-verified 1-GPU kernel/attention unit tests for AMD PR CI (#30290)
|
2026-07-07 14:44:45 -07:00 |
|
Wang, FangYuan
|
40a68521c9
|
[AMD] Fix DeepSeekV4 server cutlass error (#30374)
|
2026-07-07 14:43:05 -07:00 |
|
 
|
bbc537035a
|
[DSA] Re-enable fused top-k v2 for MTP: clamp padded-row seq_lens to >= 0 (#30378)
Co-authored-by: ziyi.xu <ziyi.xu@radixark.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-07 13:44:01 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
6875df3378
|
[Cherry pick to release/v0.5.15] Fix NVILA weight loading (#30400)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-07 13:21:55 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
e2ea7aafad
|
[Cherry pick to release/v0.5.15] Fix NVFP4 online quantization (#30397)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-07 13:20:36 -07:00 |
|
 sglang-botandsglang-bot
|
d88644b430
|
docs: sync LMSYS SGLang blog cards (#30395)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-07 12:30:47 -07:00 |
|
zijiexia
|
0bf7ddb481
|
docs(install): add nightly install + docker tag guidance, and auto-bump version on release tag (#30308)
|
2026-07-07 12:10:05 -07:00 |
|
cctry
|
2ad9a243f5
|
Size KV pool after CUDA graph capture (opt-in) (#30157)
|
2026-07-07 12:05:01 -07:00 |
|
 Xiaoyu ZhangandZijie Xia
|
ead1e490b5
|
[Doc] Add LongCat 2.0 FP8 cookbook (#30320)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-07-07 11:48:13 -07:00 |
|
 Jzz1943andYihao Wang
|
11cea29c90
|
[diffusion][cache-dit] add dual-transformer Cache-DiT adapter specs (#30150)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
|
2026-07-07 08:10:52 -07:00 |
|
 
|
e339c83f82
|
[Model] Support LongCat 2.0 FP8 (#30275)
Co-authored-by: sunjiaqi11 <sunjiaqi11@meituan.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-07-07 19:51:12 +08:00 |
|
amote-i
|
cfd3fdc54f
|
[NPU] [DOC] Update features and mainstream models on ascend npu (#30370)
|
2026-07-07 19:03:43 +08:00 |
|
loading66
|
efdf02a38a
|
[NPU]Add support --pre-warm-nccl (#30312)
|
2026-07-07 17:17:29 +08:00 |
|
Jialin Ouyang
|
7fdc1cef17
|
[fix] Fix two trunk test regressions due to flexkv change (#29701) (#30372)
|
2026-07-07 16:58:45 +08:00 |
|
Baizhou Zhang
|
946804e042
|
Disable FA3 sparse mask kernels by default (#30356)
|
2026-07-07 01:46:44 -07:00 |
|
Mohammad Miadh Angkad
|
f32b4ecd26
|
[Docs] Use trtllm_mha for Qwen3.6 B300 (#29964)
|
2026-07-07 01:44:01 -07:00 |
|
Rita Brugarolas
|
9ddea8d9ef
|
[AMD] [MORI-EP] Skip LocalExpertCount kernel in decode graph when not recording (#30302)
|
2026-07-07 01:07:07 -07:00 |
|
qinsir5522
|
5e9032c527
|
[NPU]Modify LoRA heading in ascend_npu_support_features.mdx to specify Qwen model limitations. (#30358)
|
2026-07-07 16:01:51 +08:00 |
|
linhu-nv
|
50aa97da45
|
Feat/flexkv main connector (#29701)
|
2026-07-07 15:35:52 +08:00 |
|
 Bingxu ChenandYC Yen-Ching Tseng
|
dabd4cfcfd
|
[AMD] Cap DSV4 Flash max_total_num_tokens (#30313)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
|
2026-07-07 15:33:05 +08:00 |
|
Alison Shao
|
99db3b0fa5
|
ci: run jit-kernel tests on scheduled full runs (#30306)
|
2026-07-07 00:14:39 -07:00 |
|
Wang, FangYuan
|
9a6f8e5992
|
[AMD] Fix DeepSeek V4 MTP accuracy issue (#30333)
|
2026-07-06 23:57:57 -07:00 |
|
 
|
669fd4b8a5
|
[PP] Fix start_layer_id with pp in get kv_buffer_shape (#29887)
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-07-07 14:54:38 +08:00 |
|
 
|
fefc1743a9
|
Cute-DSL FP8 MQA logits (#25220)
Co-authored-by: Mindy Li <11663212+limin2021@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-06 23:34:07 -07:00 |
|
ZeyuanChen2000
|
2d9f0b3317
|
[NPU] [DOC] Update arguments detail to NPU support features page (#30328)
|
2026-07-07 14:09:46 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
1da7d3a50b
|
[MoE] Retire the AOT moe_fused_gate / kimi_k2_moe_fused_gate gate kernels (#26771) (#29997)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-07 13:53:17 +08:00 |
|
     
|
9bd02dc5b9
|
feat(sgl-kernel): add InfLLM v2 attention kernels (#29383)
Co-authored-by: Size Wang <paulgeorge13hhhhh@gmail.com>
Co-authored-by: lijiayi <lijiayi@modelbest.cn>
Co-authored-by: suhmily10 <suhmily@gmail.com>
Co-authored-by: Xiaoyue Xu <xiaoyue.xu.me@gmail.com>
Co-authored-by: hansjohn <74091612+hansjohn@users.noreply.github.com>
Co-authored-by: zhangyan <1762895426@qq.com>
|
2026-07-06 22:46:54 -07:00 |
|
DarkSharpness
|
be70bfbdbb
|
[DSA] Fold page-table into fused top-k v2 (decode): drop page_size=1 expansion (#30274)
|
2026-07-06 21:28:05 -07:00 |
|
 
|
36b449af19
|
[EPD] Optimize multimodal global cache with paged embedding pool (#28441)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
|
2026-07-07 12:05:52 +08:00 |
|
Zyann
|
541f9221da
|
feat(metrics): add Prometheus metrics for the EPD encoder server (#27564)
|
2026-07-07 12:04:51 +08:00 |
|
Jialin Ouyang
|
abafe0022e
|
[Unified Radix Cache] Rename tree variables to cache in unittest (#30281)
|
2026-07-07 11:48:37 +08:00 |
|
loading66
|
998acf7df8
|
[DOCS][NPU]update npu support features (#30324)
|
2026-07-07 11:37:19 +08:00 |
|
   
|
16372b4c5f
|
[Spec] Anchor GLM-5.2 MTP IndexShare topk on the draft-extend step (#29787)
Co-authored-by: kpham-sgl <264503018+kpham-sgl@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-06 20:36:48 -07:00 |
|
Siming Deng
|
4145e595cf
|
[MLX] Add correctness tests for qwen2_moe and qwen3_moe (#29440)
|
2026-07-06 20:32:12 -07:00 |
|
 Siming Dengandsiming-deng
|
df06e03662
|
[MLX] Size the attention KV pool at the compute dtype for quantized models (#30097)
Co-authored-by: siming-deng <deng_siming@apple.com>
|
2026-07-06 20:31:59 -07:00 |
|
NOOB
|
c3da0a2582
|
[MLX] Fix single-token chunked-prefill continuation misrouted as decode (#30181)
|
2026-07-06 20:31:33 -07:00 |
|
 
|
3a679459e5
|
[bench] Add agentic-trace multi-turn dataset to bench_serving (#29215)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-06 19:45:44 -07:00 |
|
 
|
e85ef54877
|
Support Cutedsl BF16 GEMM JIT kernel (#30117)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-07-06 19:16:16 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
267ff1b5f9
|
Fix LTX2 RoPE JIT kernel CI (#30278)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-07 10:15:21 +08:00 |
|
 laixinandPeng Zhang
|
6279805962
|
[DSv4] Loading Time Weight Dequant (#27867)
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
|
2026-07-07 09:54:34 +08:00 |
|
Yuwei An
|
3cbb7568bd
|
[Experimental] Full Cuda Graph Support for Prefill (#27988)
|
2026-07-06 18:13:03 -07:00 |
|
Cheng Wan
|
c861896721
|
[refactor] Resolve config declarations onto server_args at the end of __post_init__ (#30297)
|
2026-07-06 18:04:44 -07:00 |
|
 sglang-botandsglang-bot
|
cf4edda956
|
docs: sync LMSYS SGLang blog cards (#30311)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-07 01:04:21 +00:00 |
|
 Rahul VijayaraghavanandMa Mingfei
|
4b5c612257
|
Skip redundant moe_sum_reduce for single-expert routing on XPU (#22660)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-07 09:03:16 +08:00 |
|
Rockdu
|
7047afafec
|
[diffusion] fix: slice img_shapes per-sample in rollout response extractor (#29989)
|
2026-07-07 08:56:52 +08:00 |
|
Ma Mingfei
|
30fb0dd851
|
[CPU] add fused_qk_gemma_norm and refactor norm kernel implementation (#30216)
|
2026-07-07 08:52:59 +08:00 |
|
Mick
|
6c1fb8a937
|
[diffusion] fix: fix ragged-caption dynamic-batching accuracy bug in ernie-Image (#30241)
|
2026-07-07 08:41:51 +08:00 |
|