Commit Graph
8903 Commits
Author SHA1 Message Date
Cheng WanandClaude Sonnet 4.6 ff8ed7a302 [refactor] unify cuda-graph capture/replay across attention backends (#26665)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-29 12:46:42 -07:00
Bingxu ChenandCursor Agent f113ece5cc Revert "improve: combine vit calls for images from different reqs from one batch (#25910)" (#26442)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-29 11:13:34 -07:00
Cheng WanandClaude Sonnet 4.6 ec075d8bc5 Fix DRAFT_EXTEND_V2 CG metadata: align test fixture and Triton with production seq_lens convention (#26651)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 02:46:45 -07:00
akhoroshev 4585f8eb95 [refactor] remove unused op_mlp (#26673) 2026-05-29 02:38:56 -07:00
chenxb002andkjuuii 8652001b6a fix: use req.req_pool_idx instead of loop variable for req_to_token i… (#26534)
Co-authored-by: kjuuii <1375341936@qq.com>
2026-05-29 02:34:42 -07:00
Liangsheng Yin ed85bcf8c3 pin kernels<0.15 (#26704) 2026-05-29 01:46:57 -07:00
Teng Ma 544f3039d5 [PD] Fix IB device validation for JSON mappings (#26114) 2026-05-29 16:44:49 +08:00
Rita Brugarolas 9062f583db [ROCm] Eliminate redundant contiguous copy in MLA attention on ROCm MXFP4 (#25463)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
2026-05-29 01:28:32 -07:00
3ecf2c76ad [CPU] Add GPT-OSS model optimization for CPU (#16775)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
2026-05-29 16:05:26 +08:00
Bingxu Chen 5601b7139d [core] Make overlap-schedule WAR barrier CUDA-only (#26646) 2026-05-29 01:02:31 -07:00
Niko Ma 4d1163e6a9 [PD][MoRI] Align hybrid state transfer with per-component schema (#26539) 2026-05-29 00:54:46 -07:00
Chizheng Fang a42a7654a2 Update MooncakeStore batch tests to use v1 APIs (#25880)
Signed-off-by: fangchizheng <fangchizheng@mail.ustc.edu.cn>
2026-05-29 00:18:05 -07:00
Arik ace730db48 [AMD] Work around HIP TPOT regression from Event.wait() in MTP seq lens resolution (#26672) 2026-05-29 00:18:02 -07:00
Aditya SharmaandXiaodong Ye b2eed9e16d [Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-05-29 15:09:33 +08:00
Yuan Luoandluoyuan.luo 08ec19872c [HotFix][Ling 2.6] Fix HybridLinearAttn dispatcher for Ling-2.6 (#26474)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-05-29 14:50:15 +08:00
王鹤男andwhn09 5850aa14c3 fix(mooncake): honour MOONCAKE_PROTOCOL so EFA hardware can select efa transport (#25083)
Co-authored-by: whn09 <whn09@users.noreply.github.com>
2026-05-29 14:21:15 +08:00
73c99e3361 Ensure multi-node MM embedding cache consistency in insert_batch (#25959)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: siyu <liusy58@linux.alibaba.com>
2026-05-29 13:59:46 +08:00
Cheng WanandClaude Sonnet 4.6 2dfbc3d781 test: strengthen CG-replay coverage with prod-fill padding, metadata invariants, and pad-ratio sweep (#26658)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 22:43:29 -07:00
Colin Zandyichiche@amd.com 226649e3b7 [Fix] Fix FP8 Online Quantization (#26415)
Co-authored-by: yichiche@amd.com <jacky.cheng>
2026-05-28 22:00:09 -07:00
Yongji Wu f16816f043 fix: copy seq_lens in TRTLLM MHA draft decode cuda graph capture (#26521) 2026-05-28 21:55:33 -07:00
yuefeng Wu b47366fbf9 [NPU]: Optimize xgrammar token bitmask on NPU with AscendC (#24133) 2026-05-28 21:38:06 -07:00
40f91e6697 [Bugfix] [DSA] [Hisparse] Broadcast TP Rank 0 Topk Indexes to other TPs (#24654)
Co-authored-by: xz-keg <xuzou_keg@outlook.com>
Co-authored-by: xuzou <xu.zou@aminer.cn>
2026-05-28 21:14:46 -07:00
Dawid Majchrowski 3ea9607d1c [diffusion] model: update to new model format (#26492) 2026-05-29 12:08:45 +08:00
Lianmin Zheng dc4e7bc479 Fix TRTLLM MHA draft decode cache seqlens replay (#26655) 2026-05-28 20:58:16 -07:00
Xinyuan Tong 79c844527c Upgrade xgrammar to 0.2.1 (#25676) 2026-05-29 11:40:07 +08:00
YC Yen-Ching Tseng 272066566f [AMD] Pin compressed-tensors<0.16.0 for srt_hip (fixes ROCm 7.2 nightly build) (#26591) 2026-05-29 11:34:45 +08:00
McZyWu b1173c8c14 [NPU] Enhance accuracy for model Step3_5 from 0 to 88% (#24582) 2026-05-29 11:29:30 +08:00
LucQueenandZhengWG 36d0a6e08e [EPD] Optimize the Mooncake backend (#22587)
Co-authored-by: ZhengWG <zwg0606@gmail.com>
2026-05-29 10:42:24 +08:00
Erik Wijmans 54b06f199c [lora] Share MoE LoRA Info (#24160) 2026-05-29 11:01:47 +09:00
Cheng WanandClaude Sonnet 4.6 e381312664 Revert "Fix FA DRAFT_EXTEND_V2 cache extent" (#26628)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:45:33 -07:00
Rohit Kumar Singh 0f8104ef15 [XPU] Fix Device Assignment (#26257) 2026-05-29 09:38:11 +08:00
McZyWu 1c2857b064 bugfix: --decrypted-draft-config-file not applied (#25960) 2026-05-29 09:15:05 +08:00
Kurkur fffdfb6fcb [Fix][NPU] Preserve existing packed_modules_mapping when merging model-level fused module mappings (#25755) 2026-05-29 09:11:36 +08:00
Cheng WanandClaude Opus 4.7 f66f56c6bd Add attention-backend unit-test suite under test/registered/attention/unittest (#26517)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 17:30:31 -07:00
3bdea78ad1 model: support Step-3.7-Flash (#26565)
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: luotingdan <luotingdan@stepfun.com>
2026-05-29 08:00:54 +08:00
Yujun Dong 3e255fd493 fix: Graceful fallback to CustomAllReduce when full_nvlink is not True (#25650) 2026-05-28 16:28:33 -07:00
Cheng WanandClaude Opus 4.7 4f92e63c99 Let unittest._ShouldStop propagate through retry() so subTest+failfast works (#26616)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 16:26:17 -07:00
Liangsheng Yin ec78fa6518 test/registered: cleanup pure model e2e tests (moves, splits, dedup, kit) (#26610) 2026-05-28 15:41:46 -07:00
Jimmy Shong f838adb7d4 bench_serving: add Zipfian shared-prefix sampling to generated-shared-prefix (#26378) 2026-05-28 14:39:46 -07:00
97d129f8c6 # feat(bench): add SPEED-Bench dataset support to bench_serving (#24149)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-05-28 14:37:00 -07:00
Khoa PhamandCursor 93445e6359 [spec decoding] support kimi-k2.6-eagle3.1-mla draft (#26506)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-28 14:26:39 -07:00
shuwennandmaodoudou168 68706e615a [SPEC] fix: use effective max draft tokens for adaptive spec initiali… (#26354)
Co-authored-by: maodoudou168 <maodoudou168@users.noreply.github.com>
2026-05-28 13:33:11 -07:00
Cameron Quilici 3ca8cb470f [BugFix] preserve cached token details in multi-tokenizer output (#26590) 2026-05-28 13:31:55 -07:00
Mohammad Miadh Angkad 690d4cdd94 Revert "[CI] FA3: ascending cuda-graph capture to avoid varlen workspace IMA (#26532) (#26550)" (#26600) 2026-05-28 13:17:37 -07:00
Vladislav Nosivskoy 34ea682a07 [UnifiedTree] gate load back pre-evict on full-attn availability only (#26302)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
2026-05-29 00:30:03 +08:00
Yaochen Hanandronnie_zheng e33bbbb467 [5/N] Quantization Refactor: GPTQ schemes and kernel split (#26402)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-28 17:15:16 +03:00
mispa-ms d616b8edad [diffusion][jit_kernel] perf: varlen FA fast path for USPAttention masked branch (#26318) 2026-05-28 21:26:09 +08:00
be32df33b9 [MUSA] Fix startup with patched torchada (#26437)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
2026-05-28 20:57:55 +08:00
Jiajun Li f4eac50389 Fix GemmaRMSNorm gemma_weight buffer storage for Qwen3.5 (#26430) 2026-05-28 18:42:46 +08:00
Liangsheng Yin 8e0ed75f2d Remove dead fields and always-False plumbing across SB / FB / LogitsMetadata (#26551) 2026-05-28 03:15:04 -07:00