Commit Graph
13001 Commits
Author SHA1 Message Date
Cheng WanandClaude Sonnet 4.6 d765dfd043 refactor(attn): init hisparse_coordinator before attn_backend; replace lazy property with init-time capture (#26012)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-21 16:03:42 -07:00
Cheng WanandClaude Sonnet 4.6 c5251a98a9 feat(model_runner): remove pool/backend refs from ForwardBatch via ForwardContext (#25983)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-21 14:01:49 -07:00
Liangsheng Yin 44ec2ee18d [core] Unify output_tokens_buf in FutureMap (#25922) 2026-05-21 13:56:44 -07:00
Liangsheng Yin c9a0e55eb5 [Spec] Polish FutureMap after #25879: rename callback, async guard, cleanup (#25962) 2026-05-21 13:56:22 -07:00
zijiexiaandClaude Opus 4.7 17dadebd4e [Docs] DeepSeek-V4: switch H200 FP4 Pro to flashinfer_mxfp4, Flash Balanced too (#25923)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 13:51:40 -07:00
Jimmy Shong 1a85586738 [Fix]: Restrict Kimi-K2.5 shared-experts fusion to Quark MXFP4 checkpoints (#25974) 2026-05-21 13:07:45 -07:00
Yuhao Yang 81d686d9fa Default MegaMoE to W4A8 for Max-Throughput recipe (#26004) 2026-05-21 11:54:16 -07:00
Cheng WanandClaude Opus 4.7 b765faee30 [MoE Refactor] deprecate forward_npu and NpuFuseEPMoE (#25678)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 10:25:17 -07:00
Ethan (Yusheng) SuandCursor a24c374f84 [lora] Remove synchronous .any().item() guard in LoRA MoE prefill path (#25531)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 23:58:57 +08:00
Mick ca9dc17be4 [diffusion] chore: adjust layer wise-offload strategy (#25930) 2026-05-21 23:48:58 +08:00
Kangyan-ZhouandClaude Opus 4.7 049bb83134 [CI] bot-cherry-pick: surface created PR number/URL in job summary (#26001)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-21 22:58:21 +08:00
liuxianglong17 32f996b75a Avoiding the problem of printing a large number of compatibility warn… (#25956) 2026-05-21 22:10:23 +08:00
Kangyan-ZhouandLiangsheng Yin caa9f08294 [CI] Force-reinstall nvidia-cutlass-dsl-libs-cu13 last to avoid wheel-mix TypeError (#25958)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-05-21 22:01:42 +08:00
amote-i ac83d8a339 docs: delete deprecated args from npu supported features (#25995) 2026-05-21 20:20:29 +08:00
Kangyan-Zhou 32352f7edf [CI] Fix bot-cherry-pick: use state == "MERGED" instead of invalid merged field (#25987) 2026-05-21 20:10:12 +08:00
Shangming Cai fbebdd5105 [CI] Enable nixl disaggregation test for decode radix cache (#25990) 2026-05-21 19:29:24 +08:00
loading66 2e0d2d4c18 [NPU][DOCS]Add best practice and benchmark result parameter description (#25875) 2026-05-21 19:08:10 +08:00
Kangyan-Zhou 64f21b1589 [CI] Improve bot-cherry-pick: accept PR number, require merged, explicit title (#25981) 2026-05-21 18:32:57 +08:00
8562d5ae94 [AMD] Relaxing Timeout for AMD stage-a (#25978)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Bingxu Chen <Bingxu.Chen@amd.com>
Co-authored-by: bingxche <bingxche@amd.com>
2026-05-21 17:32:51 +08:00
jianzhao-xu f66881f03c [NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#25384) 2026-05-21 17:32:25 +08:00
Bingxu ChenandCursor e72e3145a0 [AMD] Upgrade AITER (#25896)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 02:10:43 -07:00
Liangsheng Yin baeac179f7 [Spec] Route seq_lens through FutureMap; drop verify_done.wait (#25879) 2026-05-21 01:51:40 -07:00
DarkSharpnessandClaude 19f55c0e6d [Refactor] major JIT kernel clean up for dsv4 (#25884)
Co-authored-by: Claude <noreply@anthropic.com>
2026-05-21 01:14:31 -07:00
Liangsheng Yin 1b3d8da827 cap API quota for runner-utilization / amd-ci-job-monitor (#25965) 2026-05-21 01:13:37 -07:00
hxie c3f9bc9818 Fix nixl mla key and backup skipping (#24376) 2026-05-21 00:48:28 -07:00
看海的人 b9ae8353d2 [NPU] Support model DeepSeek-OCR and DeepSeek-OCR-2 (#25257) 2026-05-21 15:21:20 +08:00
Liwansi 190488e9a8 [NPU] Support chunk prefill for Qwen3.5/Qwen3.6 models (#25839) 2026-05-21 14:44:25 +08:00
Kangyan-Zhou 4ea8282cb7 [Revert] nvidia-cutlass-dsl[cu13] 4.5.1 -> 4.5.0 (#25938) 2026-05-21 14:36:56 +08:00
Xinyuan Tong 40faf44f7a [auto-detect] match Ring-2.6/Ling XML kv tool-call format via vocab signature (#25366) 2026-05-20 23:34:52 -07:00
Bingxu ChenandCursor Agent 45cadc215f [AMD][CI] Clean up AMD nightly + pr-test workflows (#25266)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-20 23:30:26 -07:00
ybyang b5b9c809e1 fix(model-gateway): rustfmt nightly in conversations/handlers.rs (#25947) 2026-05-20 23:29:51 -07:00
sushil Dubey c4f14650b9 fix act fun for xpu (#23809) 2026-05-21 14:02:21 +08:00
Lianmin ZhengandJaewon 8fa56a0ab1 Fix FlashInfer A2A token cap sizing (#25907)
Co-authored-by: Jaewon <52840625+jaewonlee-fb@users.noreply.github.com>
2026-05-20 23:01:28 -07:00
xutizhou e8608bdcb5 Fix EPLB redundant experts with shared expert fusion and Waterfill (#25367) 2026-05-20 22:58:08 -07:00
Charles Chen 847cbada9c Support Gemma4 MoE NVFP4 (#25054) 2026-05-20 22:45:15 -07:00
Cheng WanandClaude Sonnet 4.6 888a8794ef [Fix] DSV4 cached_loc invalidated when SWA mapping is rebuilt (#25889)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-20 22:38:12 -07:00
YAMY 3a6de13cd8 perf(dsv4): add MHC token-count prewarm (#25810) 2026-05-20 22:22:41 -07:00
Mick 1ac3e33622 [diffusion] optimize: enable inference mode in pipeline executor (#25891) 2026-05-21 13:20:24 +08:00
Brilliant Hanabi e56db8bd24 fix: use base GPU ID CUDA device for multimodal processor (#21191) 2026-05-21 13:18:37 +08:00
84ea47eb22 [CPU] Fix issues when running llama3.2-11B vision model with image tasks (#8666)
Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
2026-05-21 13:09:18 +08:00
Cheng WanandClaude Sonnet 4.6 79b937aefb [Refactor] Encapsulate SWA loc translation inside SWAKVPool with per-batch cache invalidation (#25824)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-20 21:26:32 -07:00
YC Yen-Ching Tseng 90efa9c83f [AMD] Fix AMD stage-a-test-small-1-gpu (#25932) 2026-05-20 20:51:49 -07:00
Randall LinandCursor 791a2f057f Add overridable hooks for custom chat serving implementations (#25807)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 11:21:25 +08:00
Zheng Wengang 74c6294ba9 [BugFix][EPD]Fix Qwen3VLMoe encoder-only AttributeError (#25759) 2026-05-21 11:07:21 +08:00
Mohammad Miadh Angkad a449ee4822 [Deps] Use cu13 extra for nvidia cutlass dsl (#25576) 2026-05-21 10:31:27 +08:00
jiayisunxandMa Mingfei 34479c19bd [XPU] upgrade triton-xpu version to 3.7.1 (#25730)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-21 10:29:20 +08:00
Kangyan-ZhouandClaude Opus 4.7 4868b92d47 [CI] Fix bot-cherry-pick auth: GITHUB_TOKEN for push, dedicated PAT for PR (#25926)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 10:28:42 +08:00
Cheng WanandClaude Sonnet 4.6 a528eb7564 fix: rustfmt service_discovery.rs warn! line length (#25927)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-20 19:25:11 -07:00
silencejade e603beab55 [NPU] Add Qwen3.5-397B-A17B best practice doc (#25594) 2026-05-21 10:02:12 +08:00
Hanming Lu ddf3817924 Revert "[AMD]fix: use CUDA event for targeted draft-to-verify sync in… (#25917) 2026-05-20 18:49:01 -07:00