Commit Graph
18707 Commits
Author SHA1 Message Date
Shangming Cai 378e66d248 [PD] Remove outdated backend whitelist for decode radix cache (#28238)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-06-15 22:36:05 +08:00
Trevor Morris 20f4272109 fix: Fix DSR1 perf regression due to unnecessarily falling back to triton gemm (#28073) 2026-06-15 09:45:09 -04:00
YAMY d5899b95c4 fix(qwen3.5): keep CUDA dual-stream overlap (regressed by #25885) (#27868) 2026-06-15 09:44:21 -04:00
littleyellowbicycle 09e9c4fde3 【bugfix】The NPU's forward_dsa_prepare_npu also needs special handling for is_nextn (#28118) 2026-06-15 21:31:39 +08:00
Kurkur e985422b2b [Fix][MTP][MM] Fix EAGLE v2 chunked-prefill next-token chain crash on multimodal models due to placeholder tokens (#27863) 2026-06-15 21:30:45 +08:00
qinsir5522 81166f382d Fix inaccuracies and add NPU constraints in ascend_npu_profiling.mdx. (#28283) 2026-06-15 21:12:18 +08:00
jianzhao-xu f8d1d397b6 [NPU] fix ascend_docs (#28279) 2026-06-15 21:11:55 +08:00
McZyWu 8bdb007e58 [NPU] Docs op performance optimize (#28277) 2026-06-15 20:28:18 +08:00
Mick 818808d152 [diffusion] optimize: optimize causal conv3d vae padding (#28204) 2026-06-15 20:18:08 +08:00
loading66 f768344b1a [DOCS][NPU]Supplementary Notes (#28295) 2026-06-15 20:06:05 +08:00
longxin9715 7bd1a9d163 Update documentation for Ascend NPU Guide (#28284) 2026-06-15 20:02:50 +08:00
amote-i edd5eff519 [NPU] [DOC] fix issues in ascend_npu_support_new_models (#28296) 2026-06-15 19:58:33 +08:00
iridiumine 3df6e2f968 [NPU] Add MiMo-V2-Flash manual testcases (#28223) 2026-06-15 19:57:49 +08:00
eb349efb14 [EPD][BugFix] Fix encode_with_global_cache_mooncake (#28031)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: Michael Qiu <qiudayu.qdy@antgroup.com>
2026-06-15 19:47:52 +08:00
c4ec39a785 [AMD] refactor sparse MLA decode kernel for Deepseek V4 triton backend (#28265)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
2026-06-15 03:26:59 -07:00
Thomas Wang da12f36629 [AMD] Refactor unified_kv attention metadata to data class and fuse c4/128 out_loc (#28275) 2026-06-15 02:59:01 -07:00
Lianmin Zhengandgemini-code-assist[bot] 3b419f66da [JIT] Track angle-bracket includes in source hash (#28273)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-15 02:58:14 -07:00
Zhonghua Deng 2a33724c9b [perf] Reuse a pooled HTTP session for multimodal URL downloads (#28056)
Signed-off-by: Abatom <abzhonghua@gmail.com>
2026-06-15 17:47:47 +08:00
McZyWu bf186cf8fc bugfix revise interface get cpu copy for npu mem pool to align with gpu (#27802) 2026-06-15 17:24:10 +08:00
monkeyLoveding 1180b70440 triton-ascend update (#27624) 2026-06-15 17:20:19 +08:00
Bingxu Chen 9864059e2b [AMD] Update AITER commit (#28249) 2026-06-15 02:04:34 -07:00
Bingxu ChenandCursor 19c78552dc [AMD] Restrict CI image fallback to versioned tags (#28263)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-15 16:08:48 +08:00
Xinyuan Tongandzijiexia 33f99831f8 docs(minimax-m3): refresh B200 benchmarks (tp8, piecewise) + add GPQA (#28207)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-15 00:15:38 -07:00
Jia-Wei Jiang a88ba6cc0b [Fix] Reduce power of two to constant time (#28228)
Signed-off-by: JiangJiaWei1103 <waynechuang97@gmail.com>
2026-06-14 23:25:22 -07:00
Duyi-Wang 63df86f5e7 [AMD] Skip eplb bookkeeping and topk remap when EPLB is not in use on mori-ep / HIP (#22985) (#28188) 2026-06-14 23:00:23 -07:00
Mick 578e936d8d [diffusion] feat: persist torch.compile inductor/triton cache across restarts (#28205) 2026-06-15 13:34:19 +08:00
07b9108348 [Diffusion] FLUX: fuse FeedForward GELU into up-proj GEMM (cublasLt epilogue) (#28166)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 13:00:26 +08:00
Prajjandprajjwal1 441b75ee69 [quantization] NVFP4 MoE: split fused w13 gate/up global scales (#27588)
Co-authored-by: prajjwal1 <prajjwal1@protonmail.com>
2026-06-14 21:18:36 -07:00
Ryan Zzzandzhujunyu ce9fad7196 [Bugfix][DeepSeek-V4] Fix Spec V2 Draft Input ID Dtype for DP Collectives (#28043)
Co-authored-by: zhujunyu <zhujunyu.666@bytedance.com>
2026-06-14 21:11:53 -07:00
weireweire bf38a0b03d Fix disaggregated decode load token accounting (#25736) 2026-06-15 11:37:41 +08:00
0417951a86 [Bug Fix] Validate tokenizer-dependent features with skip_tokenizer_init (#27882)
Co-authored-by: Randall <randall@iterationlab.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-14 20:07:52 -07:00
Zyann 37505eca27 feat: report multimodal (image/audio/video) token counts in usage.prompt_tokens_details (#27122) 2026-06-15 11:04:10 +08:00
Kangyan-ZhouandClaude Opus 4.8 c0dfe4c8ec [router] Reconcile workers that registered without resolving model_ids (#27980)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 18:47:50 -07:00
Mohammad Miadh Angkad 69b02ea68a [Distributed] Guard torch symm mem all-reduce sizes (#24548) 2026-06-14 18:28:57 -07:00
Xinyuan Tong 1a66059c4e [Spec] Restore index_share_for_mtp_iteration in EAGLE V2 draft worker (#28192) 2026-06-14 18:01:11 -07:00
Michael c127ba6483 [AMD] ci: fix scheduled AMD runs startup failure when calling extra-a suite (#28214) 2026-06-14 17:43:47 -07:00
Xinyuan Tong 000fc975c7 ci(docker): support layered overlay images in release-docker-dev (#28206) 2026-06-14 16:34:02 -07:00
Lianmin Zheng f18d38d040 Revert "[AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28213) 2026-06-14 13:34:06 -07:00
Mick 3cb29f6747 [diffusion] feat: use regional torch.compile (compile_repeated_blocks) for DiT of diffusers backend (#28193) 2026-06-15 00:34:10 +08:00
Mick ec36dde580 [diffusion] feat: add --warmup-mode enum server arg (#28184) 2026-06-14 23:09:04 +08:00
Mick 582bd23f71 [diffusion] feat: enable spatial-shard vae decode across GPUs (#28071) 2026-06-14 20:19:44 +08:00
d72314808f [JIT Kernel] Multi-GPU test/bench framework for custom all-reduce + TP QKNorm (#26706)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: ziyi.xu <ziyi.xu@radixark.ai>
2026-06-14 17:20:36 +08:00
Mick 5331de0f8c [diffusion] chore: resolve model_index.json Hub-first with local-cache fallback (#28177) 2026-06-14 16:48:38 +08:00
Humphrey 8c334e2224 fix(io_struct): index extra_key per sub-request in batched GenerateReqInput (#26971) 2026-06-14 00:50:38 -07:00
Mick 1456eb612d [diffusion] CI: tighten perf baselines (#28123) 2026-06-14 15:50:21 +08:00
Jimmy Shongandgithub-actions[bot] 54acffc864 Eval accuracy gpqa aime25 mixins (#27102)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-14 00:49:12 -07:00
Qiaolin Yu f293ddf3ce [perf] reduce overhead of fill_ids list reconstruction and decref (#27965) 2026-06-14 00:41:11 -07:00
Yongji Wu f2d7d67603 numa: bind within allowed CPUs when affinity is already constrained (#26983) 2026-06-14 00:38:40 -07:00
b796338271 Fix prefill delayer wait histograms always observing 0 (#25975)
Co-authored-by: kingjameschan <170807154+kingjameschan@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Humphrey <181440142+humphreysun98@users.noreply.github.com>
2026-06-14 00:35:29 -07:00
David Wang 8c5320b37e dflash add sliding window attention draft layer support (#27469) 2026-06-14 00:32:02 -07:00