huangtingwei
|
c93a559e5a
|
Fix MoE LoRA wrapper exposing moe_runner_config (#26710)
|
2026-05-30 01:59:15 -07:00 |
|
 
|
3b61a1f935
|
[Bugfix] Optimize metadata allocation and transfer for mooncake intraNode NVLink (#26707)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-30 16:52:01 +08:00 |
|
Alison Shao
|
edfe8d34e8
|
[CI] ci-coverage-overview: schedule + manual only, include XPU/MUSA/multimodal_gen (#26619)
|
2026-05-30 16:39:18 +08:00 |
|
Jimmy Shong
|
716e670d3d
|
[bugfix]: size CuteDSL MoE allgather buffers for the worst-case forward (#26696)
|
2026-05-30 00:27:20 -07:00 |
|
Zhangheng
|
7662210406
|
[UnifiedTree]: Support eviction priority (#26549)
|
2026-05-30 15:19:43 +08:00 |
|
gjsheu
|
02aeed5387
|
[NPU] DFlash Speculative Decoding Support NPU (#23122)
|
2026-05-30 15:13:59 +08:00 |
|
AndyLi429
|
fe4b29d391
|
[Bugfix] Fix Ascend NPU CP attention for batch size > 1 (#26705)
|
2026-05-30 15:07:39 +08:00 |
|
Liangsheng Yin
|
1c79015434
|
Drop dead ScheduleBatch return_routed_experts/return_indexer_topk fields (#26760)
|
2026-05-29 23:16:09 -07:00 |
|
  
|
6f1c9fc77b
|
[RL] Fix crash when the reqs in a batch have a mix of return_routed_experts = True and False. (#26423)
Co-authored-by: root <root@slurm-h200-209-231.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-05-29 20:46:10 -07:00 |
|
Liangsheng Yin
|
1f850e67f2
|
[core] Fix crashes on the gpu_only spec_v2 path (#26738)
|
2026-05-29 19:12:03 -07:00 |
|
Liangsheng Yin
|
c2ac37dbcc
|
[Bug] ngram verify: keep batch.seq_lens_sum in sync after accept (#26753)
|
2026-05-29 18:00:11 -07:00 |
|
Erik Wijmans
|
95cd2fd29f
|
[lora] More efficient pinned memory (#20876)
|
2026-05-30 09:04:59 +09:00 |
|
Bruce Changlong Xu
|
a5e6a8887a
|
[attention] Fallback to Triton merge_state when FlashInfer hits CUDA thread limit (#23993)
|
2026-05-29 16:30:49 -07:00 |
|
 Byron HsuandByron Hsu
|
6ea69efb7f
|
[RL] Forward Kimi K2.5 weight hooks to language model (#26744)
Co-authored-by: Byron Hsu <24364830+ByronHsu@users.noreply.github.com>
|
2026-05-29 15:08:58 -07:00 |
|
Chao Shi
|
6ce49e5f4c
|
[Utils] Support configure log level at runtime (#26583)
|
2026-05-29 14:49:06 -07:00 |
|
  
|
cf66693b35
|
[Model] Add Qwen3-MoE MTP (#26468)
Co-authored-by: Byron Hsu <byronhsu@Byrons-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: root <root@slurm-h200-209-231.slurm-compute.tenant-slurm.svc.cluster.local>
|
2026-05-29 14:31:05 -07:00 |
|
 Lianmin Zhengandcctry
|
4ff1296f5e
|
Optimize get load calls (/v1/loads) using shared-memory load snapshots (#26348)
Co-authored-by: cctry <cctry@meta.com>
|
2026-05-29 13:40:26 -07:00 |
|
Qiaolin Yu
|
3cecc77ccb
|
[perf] Fuse NVFP4 gate_up_gemm + swiglu + output FP4 quant (#26626)
|
2026-05-29 13:16:24 -07:00 |
|
 Liangsheng YinandQiaolin-Yu
|
6b5f0d0ccb
|
[core] Make spec_v2 seq_lens_cpu optional via backend needs_cpu_seq_lens; Triton opts out (#26128)
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
|
2026-05-29 13:00:32 -07:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
ff8ed7a302
|
[refactor] unify cuda-graph capture/replay across attention backends (#26665)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-29 12:46:42 -07:00 |
|
 Bingxu ChenandCursor Agent
|
f113ece5cc
|
Revert "improve: combine vit calls for images from different reqs from one batch (#25910)" (#26442)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
|
2026-05-29 11:13:34 -07:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
ec075d8bc5
|
Fix DRAFT_EXTEND_V2 CG metadata: align test fixture and Triton with production seq_lens convention (#26651)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-29 02:46:45 -07:00 |
|
akhoroshev
|
4585f8eb95
|
[refactor] remove unused op_mlp (#26673)
|
2026-05-29 02:38:56 -07:00 |
|
 chenxb002andkjuuii
|
8652001b6a
|
fix: use req.req_pool_idx instead of loop variable for req_to_token i… (#26534)
Co-authored-by: kjuuii <1375341936@qq.com>
|
2026-05-29 02:34:42 -07:00 |
|
Liangsheng Yin
|
ed85bcf8c3
|
pin kernels<0.15 (#26704)
|
2026-05-29 01:46:57 -07:00 |
|
Teng Ma
|
544f3039d5
|
[PD] Fix IB device validation for JSON mappings (#26114)
|
2026-05-29 16:44:49 +08:00 |
|
Rita Brugarolas
|
9062f583db
|
[ROCm] Eliminate redundant contiguous copy in MLA attention on ROCm MXFP4 (#25463)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
|
2026-05-29 01:28:32 -07:00 |
|
 
|
3ecf2c76ad
|
[CPU] Add GPT-OSS model optimization for CPU (#16775)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-29 16:05:26 +08:00 |
|
Bingxu Chen
|
5601b7139d
|
[core] Make overlap-schedule WAR barrier CUDA-only (#26646)
|
2026-05-29 01:02:31 -07:00 |
|
Niko Ma
|
4d1163e6a9
|
[PD][MoRI] Align hybrid state transfer with per-component schema (#26539)
|
2026-05-29 00:54:46 -07:00 |
|
Chizheng Fang
|
a42a7654a2
|
Update MooncakeStore batch tests to use v1 APIs (#25880)
Signed-off-by: fangchizheng <fangchizheng@mail.ustc.edu.cn>
|
2026-05-29 00:18:05 -07:00 |
|
Arik
|
ace730db48
|
[AMD] Work around HIP TPOT regression from Event.wait() in MTP seq lens resolution (#26672)
|
2026-05-29 00:18:02 -07:00 |
|
 Aditya SharmaandXiaodong Ye
|
b2eed9e16d
|
[Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-05-29 15:09:33 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
08ec19872c
|
[HotFix][Ling 2.6] Fix HybridLinearAttn dispatcher for Ling-2.6 (#26474)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-29 14:50:15 +08:00 |
|
 王鹤男andwhn09
|
5850aa14c3
|
fix(mooncake): honour MOONCAKE_PROTOCOL so EFA hardware can select efa transport (#25083)
Co-authored-by: whn09 <whn09@users.noreply.github.com>
|
2026-05-29 14:21:15 +08:00 |
|
  
|
73c99e3361
|
Ensure multi-node MM embedding cache consistency in insert_batch (#25959)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: siyu <liusy58@linux.alibaba.com>
|
2026-05-29 13:59:46 +08:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
2dfbc3d781
|
test: strengthen CG-replay coverage with prod-fill padding, metadata invariants, and pad-ratio sweep (#26658)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-28 22:43:29 -07:00 |
|
 Colin Zandyichiche@amd.com
|
226649e3b7
|
[Fix] Fix FP8 Online Quantization (#26415)
Co-authored-by: yichiche@amd.com <jacky.cheng>
|
2026-05-28 22:00:09 -07:00 |
|
Yongji Wu
|
f16816f043
|
fix: copy seq_lens in TRTLLM MHA draft decode cuda graph capture (#26521)
|
2026-05-28 21:55:33 -07:00 |
|
yuefeng Wu
|
b47366fbf9
|
[NPU]: Optimize xgrammar token bitmask on NPU with AscendC (#24133)
|
2026-05-28 21:38:06 -07:00 |
|
 
|
40f91e6697
|
[Bugfix] [DSA] [Hisparse] Broadcast TP Rank 0 Topk Indexes to other TPs (#24654)
Co-authored-by: xz-keg <xuzou_keg@outlook.com>
Co-authored-by: xuzou <xu.zou@aminer.cn>
|
2026-05-28 21:14:46 -07:00 |
|
Dawid Majchrowski
|
3ea9607d1c
|
[diffusion] model: update to new model format (#26492)
|
2026-05-29 12:08:45 +08:00 |
|
Lianmin Zheng
|
dc4e7bc479
|
Fix TRTLLM MHA draft decode cache seqlens replay (#26655)
|
2026-05-28 20:58:16 -07:00 |
|
Xinyuan Tong
|
79c844527c
|
Upgrade xgrammar to 0.2.1 (#25676)
|
2026-05-29 11:40:07 +08:00 |
|
YC Yen-Ching Tseng
|
272066566f
|
[AMD] Pin compressed-tensors<0.16.0 for srt_hip (fixes ROCm 7.2 nightly build) (#26591)
|
2026-05-29 11:34:45 +08:00 |
|
McZyWu
|
b1173c8c14
|
[NPU] Enhance accuracy for model Step3_5 from 0 to 88% (#24582)
|
2026-05-29 11:29:30 +08:00 |
|
 LucQueenandZhengWG
|
36d0a6e08e
|
[EPD] Optimize the Mooncake backend (#22587)
Co-authored-by: ZhengWG <zwg0606@gmail.com>
|
2026-05-29 10:42:24 +08:00 |
|
Erik Wijmans
|
54b06f199c
|
[lora] Share MoE LoRA Info (#24160)
|
2026-05-29 11:01:47 +09:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
e381312664
|
Revert "Fix FA DRAFT_EXTEND_V2 cache extent" (#26628)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-28 18:45:33 -07:00 |
|
Rohit Kumar Singh
|
0f8104ef15
|
[XPU] Fix Device Assignment (#26257)
|
2026-05-29 09:38:11 +08:00 |
|