YAMY
|
f870bf1ed0
|
[dsv4] Prewarm MHC prenorm kernel at startup (#27986)
|
2026-06-15 13:26:26 -07:00 |
|
 Lianmin ZhengandIan O'Connell
|
7e629a2f8c
|
Allow overriding tokenizer path in benchmark harness (#28280)
Co-authored-by: Ian O'Connell <ianoc@meta.com>
|
2026-06-15 13:07:50 -07:00 |
|
 
|
33719cfb31
|
[PD] Optimize SWA allocation (#28085)
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-15 11:01:55 -07:00 |
|
Shangming Cai
|
378e66d248
|
[PD] Remove outdated backend whitelist for decode radix cache (#28238)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-06-15 22:36:05 +08:00 |
|
Trevor Morris
|
20f4272109
|
fix: Fix DSR1 perf regression due to unnecessarily falling back to triton gemm (#28073)
|
2026-06-15 09:45:09 -04:00 |
|
YAMY
|
d5899b95c4
|
fix(qwen3.5): keep CUDA dual-stream overlap (regressed by #25885) (#27868)
|
2026-06-15 09:44:21 -04:00 |
|
littleyellowbicycle
|
09e9c4fde3
|
【bugfix】The NPU's forward_dsa_prepare_npu also needs special handling for is_nextn (#28118)
|
2026-06-15 21:31:39 +08:00 |
|
Kurkur
|
e985422b2b
|
[Fix][MTP][MM] Fix EAGLE v2 chunked-prefill next-token chain crash on multimodal models due to placeholder tokens (#27863)
|
2026-06-15 21:30:45 +08:00 |
|
Mick
|
818808d152
|
[diffusion] optimize: optimize causal conv3d vae padding (#28204)
|
2026-06-15 20:18:08 +08:00 |
|
iridiumine
|
3df6e2f968
|
[NPU] Add MiMo-V2-Flash manual testcases (#28223)
|
2026-06-15 19:57:49 +08:00 |
|
 
|
eb349efb14
|
[EPD][BugFix] Fix encode_with_global_cache_mooncake (#28031)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: Michael Qiu <qiudayu.qdy@antgroup.com>
|
2026-06-15 19:47:52 +08:00 |
|
 
|
c4ec39a785
|
[AMD] refactor sparse MLA decode kernel for Deepseek V4 triton backend (#28265)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
|
2026-06-15 03:26:59 -07:00 |
|
Thomas Wang
|
da12f36629
|
[AMD] Refactor unified_kv attention metadata to data class and fuse c4/128 out_loc (#28275)
|
2026-06-15 02:59:01 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Lianmin Zhengandgemini-code-assist[bot]
|
3b419f66da
|
[JIT] Track angle-bracket includes in source hash (#28273)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-06-15 02:58:14 -07:00 |
|
Zhonghua Deng
|
2a33724c9b
|
[perf] Reuse a pooled HTTP session for multimodal URL downloads (#28056)
Signed-off-by: Abatom <abzhonghua@gmail.com>
|
2026-06-15 17:47:47 +08:00 |
|
McZyWu
|
bf186cf8fc
|
bugfix revise interface get cpu copy for npu mem pool to align with gpu (#27802)
|
2026-06-15 17:24:10 +08:00 |
|
Duyi-Wang
|
63df86f5e7
|
[AMD] Skip eplb bookkeeping and topk remap when EPLB is not in use on mori-ep / HIP (#22985) (#28188)
|
2026-06-14 23:00:23 -07:00 |
|
Mick
|
578e936d8d
|
[diffusion] feat: persist torch.compile inductor/triton cache across restarts (#28205)
|
2026-06-15 13:34:19 +08:00 |
|
 
|
07b9108348
|
[Diffusion] FLUX: fuse FeedForward GELU into up-proj GEMM (cublasLt epilogue) (#28166)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-15 13:00:26 +08:00 |
|
 Prajjandprajjwal1
|
441b75ee69
|
[quantization] NVFP4 MoE: split fused w13 gate/up global scales (#27588)
Co-authored-by: prajjwal1 <prajjwal1@protonmail.com>
|
2026-06-14 21:18:36 -07:00 |
|
 Ryan Zzzandzhujunyu
|
ce9fad7196
|
[Bugfix][DeepSeek-V4] Fix Spec V2 Draft Input ID Dtype for DP Collectives (#28043)
Co-authored-by: zhujunyu <zhujunyu.666@bytedance.com>
|
2026-06-14 21:11:53 -07:00 |
|
weireweire
|
bf38a0b03d
|
Fix disaggregated decode load token accounting (#25736)
|
2026-06-15 11:37:41 +08:00 |
|
 
|
0417951a86
|
[Bug Fix] Validate tokenizer-dependent features with skip_tokenizer_init (#27882)
Co-authored-by: Randall <randall@iterationlab.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-14 20:07:52 -07:00 |
|
Zyann
|
37505eca27
|
feat: report multimodal (image/audio/video) token counts in usage.prompt_tokens_details (#27122)
|
2026-06-15 11:04:10 +08:00 |
|
Mohammad Miadh Angkad
|
69b02ea68a
|
[Distributed] Guard torch symm mem all-reduce sizes (#24548)
|
2026-06-14 18:28:57 -07:00 |
|
Xinyuan Tong
|
1a66059c4e
|
[Spec] Restore index_share_for_mtp_iteration in EAGLE V2 draft worker (#28192)
|
2026-06-14 18:01:11 -07:00 |
|
Lianmin Zheng
|
f18d38d040
|
Revert "[AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28213)
|
2026-06-14 13:34:06 -07:00 |
|
Mick
|
3cb29f6747
|
[diffusion] feat: use regional torch.compile (compile_repeated_blocks) for DiT of diffusers backend (#28193)
|
2026-06-15 00:34:10 +08:00 |
|
Mick
|
ec36dde580
|
[diffusion] feat: add --warmup-mode enum server arg (#28184)
|
2026-06-14 23:09:04 +08:00 |
|
Mick
|
582bd23f71
|
[diffusion] feat: enable spatial-shard vae decode across GPUs (#28071)
|
2026-06-14 20:19:44 +08:00 |
|
 
|
d72314808f
|
[JIT Kernel] Multi-GPU test/bench framework for custom all-reduce + TP QKNorm (#26706)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: ziyi.xu <ziyi.xu@radixark.ai>
|
2026-06-14 17:20:36 +08:00 |
|
Mick
|
5331de0f8c
|
[diffusion] chore: resolve model_index.json Hub-first with local-cache fallback (#28177)
|
2026-06-14 16:48:38 +08:00 |
|
Humphrey
|
8c334e2224
|
fix(io_struct): index extra_key per sub-request in batched GenerateReqInput (#26971)
|
2026-06-14 00:50:38 -07:00 |
|
Mick
|
1456eb612d
|
[diffusion] CI: tighten perf baselines (#28123)
|
2026-06-14 15:50:21 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Jimmy Shongandgithub-actions[bot]
|
54acffc864
|
Eval accuracy gpqa aime25 mixins (#27102)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-06-14 00:49:12 -07:00 |
|
Qiaolin Yu
|
f293ddf3ce
|
[perf] reduce overhead of fill_ids list reconstruction and decref (#27965)
|
2026-06-14 00:41:11 -07:00 |
|
Yongji Wu
|
f2d7d67603
|
numa: bind within allowed CPUs when affinity is already constrained (#26983)
|
2026-06-14 00:38:40 -07:00 |
|
  
|
b796338271
|
Fix prefill delayer wait histograms always observing 0 (#25975)
Co-authored-by: kingjameschan <170807154+kingjameschan@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Humphrey <181440142+humphreysun98@users.noreply.github.com>
|
2026-06-14 00:35:29 -07:00 |
|
David Wang
|
8c5320b37e
|
dflash add sliding window attention draft layer support (#27469)
|
2026-06-14 00:32:02 -07:00 |
|
Liangsheng Yin
|
bb48405c31
|
Unify NVTX annotation helpers and split the enable gate per subsystem (#28165)
|
2026-06-14 00:04:59 -07:00 |
|
ybyang
|
50993554d8
|
fix(health): make health-check rid unique across tokenizer workers (#28143)
|
2026-06-14 00:01:52 -07:00 |
|
 
|
f79a6b5c33
|
Support GLM-4.7 function calling via structural tags (#28149)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-06-14 14:57:07 +08:00 |
|
Mick
|
37ed10bd24
|
[diffusion] UX: reduce attention backend log noise (#28169)
|
2026-06-14 14:55:13 +08:00 |
|
Ting SUN
|
b250bea994
|
fix(sampling): reject non-finite temperature in SamplingParams.verify (#28153)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
2026-06-13 23:41:36 -07:00 |
|
Yuzhen Zhou
|
171037c3e7
|
Fix Qwen3.5 deterministic batch-invariant logprobs (#27869)
|
2026-06-13 23:23:06 -07:00 |
|
Yanbin Jiang
|
1747b88c5e
|
[LoRA] Support DSA indexer LoRA targets for GLM-5.1 / DeepSeek-V3.2-family models (#28110)
|
2026-06-13 23:02:33 -07:00 |
|
Mick
|
31ac743484
|
[diffusion] chore: improve server warmup coverage (#28127)
|
2026-06-14 13:35:42 +08:00 |
|
Mohammad Miadh Angkad
|
91c63aeb4d
|
Fix stale CUDA graph benchmark and docs refs (#28041)
|
2026-06-13 21:51:42 -07:00 |
|
Jared Wen
|
5da3b37a9d
|
[CI] add Precision Regression Test on Nightly Run CI (#26902)
|
2026-06-14 12:43:57 +08:00 |
|
JoyFuture
|
a3fd5c24be
|
feat: add NVTX markers for the scheduler main loop (#27901)
|
2026-06-13 17:16:53 -07:00 |
|