Commit Graph
8877 Commits
Author SHA1 Message Date
McZyWu b1173c8c14 [NPU] Enhance accuracy for model Step3_5 from 0 to 88% (#24582) 2026-05-29 11:29:30 +08:00
LucQueenandZhengWG 36d0a6e08e [EPD] Optimize the Mooncake backend (#22587)
Co-authored-by: ZhengWG <zwg0606@gmail.com>
2026-05-29 10:42:24 +08:00
Erik Wijmans 54b06f199c [lora] Share MoE LoRA Info (#24160) 2026-05-29 11:01:47 +09:00
Cheng WanandClaude Sonnet 4.6 e381312664 Revert "Fix FA DRAFT_EXTEND_V2 cache extent" (#26628)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:45:33 -07:00
Rohit Kumar Singh 0f8104ef15 [XPU] Fix Device Assignment (#26257) 2026-05-29 09:38:11 +08:00
McZyWu 1c2857b064 bugfix: --decrypted-draft-config-file not applied (#25960) 2026-05-29 09:15:05 +08:00
Kurkur fffdfb6fcb [Fix][NPU] Preserve existing packed_modules_mapping when merging model-level fused module mappings (#25755) 2026-05-29 09:11:36 +08:00
Cheng WanandClaude Opus 4.7 f66f56c6bd Add attention-backend unit-test suite under test/registered/attention/unittest (#26517)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 17:30:31 -07:00
3bdea78ad1 model: support Step-3.7-Flash (#26565)
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: luotingdan <luotingdan@stepfun.com>
2026-05-29 08:00:54 +08:00
Yujun Dong 3e255fd493 fix: Graceful fallback to CustomAllReduce when full_nvlink is not True (#25650) 2026-05-28 16:28:33 -07:00
Cheng WanandClaude Opus 4.7 4f92e63c99 Let unittest._ShouldStop propagate through retry() so subTest+failfast works (#26616)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 16:26:17 -07:00
Liangsheng Yin ec78fa6518 test/registered: cleanup pure model e2e tests (moves, splits, dedup, kit) (#26610) 2026-05-28 15:41:46 -07:00
Jimmy Shong f838adb7d4 bench_serving: add Zipfian shared-prefix sampling to generated-shared-prefix (#26378) 2026-05-28 14:39:46 -07:00
97d129f8c6 # feat(bench): add SPEED-Bench dataset support to bench_serving (#24149)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-05-28 14:37:00 -07:00
Khoa PhamandCursor 93445e6359 [spec decoding] support kimi-k2.6-eagle3.1-mla draft (#26506)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-28 14:26:39 -07:00
shuwennandmaodoudou168 68706e615a [SPEC] fix: use effective max draft tokens for adaptive spec initiali… (#26354)
Co-authored-by: maodoudou168 <maodoudou168@users.noreply.github.com>
2026-05-28 13:33:11 -07:00
Cameron Quilici 3ca8cb470f [BugFix] preserve cached token details in multi-tokenizer output (#26590) 2026-05-28 13:31:55 -07:00
Mohammad Miadh Angkad 690d4cdd94 Revert "[CI] FA3: ascending cuda-graph capture to avoid varlen workspace IMA (#26532) (#26550)" (#26600) 2026-05-28 13:17:37 -07:00
Vladislav Nosivskoy 34ea682a07 [UnifiedTree] gate load back pre-evict on full-attn availability only (#26302)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
2026-05-29 00:30:03 +08:00
Yaochen Hanandronnie_zheng e33bbbb467 [5/N] Quantization Refactor: GPTQ schemes and kernel split (#26402)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-28 17:15:16 +03:00
mispa-ms d616b8edad [diffusion][jit_kernel] perf: varlen FA fast path for USPAttention masked branch (#26318) 2026-05-28 21:26:09 +08:00
be32df33b9 [MUSA] Fix startup with patched torchada (#26437)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
2026-05-28 20:57:55 +08:00
Jiajun Li f4eac50389 Fix GemmaRMSNorm gemma_weight buffer storage for Qwen3.5 (#26430) 2026-05-28 18:42:46 +08:00
Liangsheng Yin 8e0ed75f2d Remove dead fields and always-False plumbing across SB / FB / LogitsMetadata (#26551) 2026-05-28 03:15:04 -07:00
syy-hw c397a21167 [Ascend NPU] Enable GLM-4.6V series models inference (#26146) 2026-05-28 17:27:12 +08:00
Xingyu Liu 578d27e56a [bugfix] Honor cast_x_before_out_mul in RMSNorm.forward_cuda residual path (#25920)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
2026-05-28 01:22:46 -07:00
Cheng WanandClaude Opus 4.7 12e28bdf0c Fix FlashInfer SWA EXTEND-with-prefix correctness in merge_state path (#26513)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 01:16:58 -07:00
Cheng WanandClaude Opus 4.7 00cd6fb3d9 Add sliding-window mask support to TorchNativeAttnBackend (#26516)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 01:06:57 -07:00
ZeyuanChen2000 a245cae3d1 [NPU] fix model ERNIE-4.5-21B-A3B-PT bias need 1D error (#26038) 2026-05-28 16:05:39 +08:00
Cheng WanandClaude Opus 4.7 8ca09a30f1 Allow Optional key/value in unified_attention_with_output split-op (MLA absorb fix) (#26515)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 01:04:07 -07:00
Cheng WanandClaude Opus 4.7 b429a30428 Expose Flex attention causal/decode masks as static methods (#26514)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 01:03:32 -07:00
Cheng WanandClaude Opus 4.7 e5f5d84780 Fix FA DRAFT_EXTEND_V2 cache extent (#26512)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 00:56:23 -07:00
Xingyu Liu 770c51b127 [Bug] Forward fixed_split_size in SWA / cross-attention paths of FlashInfer backend (#26412)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
2026-05-28 00:52:26 -07:00
Liangsheng Yin 8dca6291c7 [CI] FA3: ascending cuda-graph capture to avoid varlen workspace IMA (#26532) (#26550) 2026-05-28 00:50:35 -07:00
DarkSharpnessandClaude Opus 4.7 8f21b3e2ef [Refactor] JIT kernel benchmark (#25274)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 00:49:55 -07:00
hanwlax 794fdd39ef fix: adapt dots_vlm for transformers v5 (#25829) 2026-05-28 15:26:13 +08:00
Brayden Zhongandb8zhong 50e0b3b77f Support Flashinfer Cute-DSL MLA attention (#24737)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-28 00:21:32 -07:00
Brayden Zhongandb8zhong e31ea50df8 Remove DeepGEMM for indexer GEMM in piecewise NSA path (#26494)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-28 00:15:23 -07:00
Liangsheng Yin 686ef50672 Group ScheduleBatch and ForwardBatch fields by data-flow role (#26022) 2026-05-28 00:11:20 -07:00
Brayden Zhongandb8zhong b4808d44da Use Cute-DSL MXFP8 quantize kernels (#25486)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-28 00:01:31 -07:00
Brayden Zhong 97dd6aad60 Add a little env var for disabling Flashinfer autotune cache (#26193) 2026-05-27 23:59:59 -07:00
Xinyuan Tong bed20249f1 fix(tool_call): reland schema type normalization (#26433) 2026-05-28 14:31:18 +08:00
Mike QiuandMike_Qiu 8ff66b707f feat: convert mm_hashes to str in encode_server for Mooncake key compat (#26487)
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
2026-05-28 14:16:29 +08:00
Xiaoyu Zhang e60f799b40 Enable Kimi-K2.5 piecewise CUDA graph (#26382) 2026-05-27 22:51:33 -07:00
Jay Chun b437a0d066 Fix PD decode radix cache double-counting cached_tokens (#25973) 2026-05-28 11:50:03 +08:00
Chunyuan WU 714fdd9723 Fix MiniMax-M2.7 on CPU (#25061) 2026-05-28 10:53:13 +08:00
Shaoting 14c1bb2721 [Feat][LMCache] Support LMCache mp mode (#24089)
Signed-off-by: Shaoting-Feng <stfeng@uw.edu>
2026-05-28 10:15:09 +08:00
jvzibro 421bda6d85 [Bug Fix] Remove H20 device check for FlashInfer AllReduce Fusion (#26470) 2026-05-28 00:45:55 +00:00
YAMY eae03ce3b2 refactor(dsv4): route MHC prenorm through DeepGEMM wrapper (#26238) 2026-05-27 17:45:45 -07:00
Netanel Haber 0abe6a85a5 Support NemotronHPuzzleForCausalLM (#24429)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
2026-05-27 16:12:44 -07:00