Mohammad Miadh Angkad
|
1b9dfa14e6
|
Fix FlashInfer MNNVL workspace size check (#32318)
|
2026-07-29 02:12:18 -07:00 |
|
Yihao Wang
|
227dadd79a
|
[diffusion] feat: support resident layers for DiT (#31538)
|
2026-07-29 16:52:55 +08:00 |
|
 Brayden ZhongandBrayden Zhong
|
7dcebca255
|
Fix nightly CI: NVFP4 cuda-graph crash, NVILA batching, CuTe paged-KV zero-size, Kimi-VL OOM (#32118)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-29 01:39:55 -07:00 |
|
 
|
f05c92fb6d
|
✨ [llm][npu][quant] Add W8A8 MXFP8 quantization for Qwen3 MoE on Ascend NPU (#30768)
Co-authored-by: Артем Савкин <58187114+OrangeRedeng@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-29 10:39:36 +03:00 |
|
Mick
|
da5528db30
|
fix(vlm): materialize Qwen3-VL features on the vision device (#31596)
|
2026-07-29 15:25:04 +08:00 |
|
 weireweireandweireweire
|
bd47ec97ff
|
[EAGLE] Handle NaNs in fused top-k=1 (#32396)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-07-29 00:07:42 -07:00 |
|
 Jyothirmai KottuandMick
|
7c248dde7f
|
[diffusion] fix: don't self-kill diffusion worker when PID 1 is the real parent (#31361)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-07-29 14:35:56 +08:00 |
|
siweil
|
9bdbb180b1
|
[Disagg][NIXL] Fix heterogeneous attn-TP KV transfer for replicated GQA heads (NIXL_ERR_NOT_FOUND) (#31968)
|
2026-07-29 14:13:23 +08:00 |
|
 
|
ef6c07008b
|
Support DCP for Kimi Linear model (#32612)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
|
2026-07-28 22:59:58 -07:00 |
|
Liangsheng Yin
|
c4fc241fd3
|
[Perf] Free KV pages by segment in the paged allocator without a device sync (#32701)
|
2026-07-28 22:24:16 -07:00 |
|
Brayden Zhong
|
f01a0c7f97
|
Fixing MXFP8 online quantization pipeline (#31510)
|
2026-07-28 21:26:13 -07:00 |
|
YC Yen-Ching Tseng
|
68673fe6c5
|
[AMD] add Gemma3RMSNorm.forward_hip to unbreak ROCm (#32613)
|
2026-07-28 21:19:17 -07:00 |
|
Xiaojun(Robin) Zhang
|
1af0167493
|
[EPD][VLM] Fix Kimi-VL 2D encoder grids (#32104)
Signed-off-by: Xiaojun Zhang <zhangxiaojunhust@gmail.com>
|
2026-07-29 11:46:17 +08:00 |
|
Junlin Wu
|
d6fcfe02d6
|
🐛 [llm][npu][quant] Fix ModelSlim MXFP4 packed weight loading (#32013)
|
2026-07-29 11:34:41 +08:00 |
|
Yihao Wang
|
9ea964a535
|
[diffusion] fix: per-shard FP8 scale shape for single-GPU fused linears (#32157)
|
2026-07-28 20:28:01 -07:00 |
|
Liangsheng Yin
|
14bd315d6e
|
[Refactor] Remove dead allocator backup_state / restore_state (#32709)
|
2026-07-28 19:49:47 -07:00 |
|
  
|
339bef7fad
|
[MLX] Fix overlap-loop request bookkeeping and graceful shutdown (#32447)
Co-authored-by: xiaolin2004 <uwowmhdjwpwpwdhwkw@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-07-28 19:04:53 -07:00 |
|
 
|
580b1acbe6
|
[MLX] Move fused swiglu tests to test/registered so CI collects them (#32448)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-07-28 19:04:24 -07:00 |
|
 Leon GaoandAndrew Gu
|
ee678910f7
|
[Kernel] Skip KV writes to reserved padding slots (#32477)
Co-authored-by: Andrew Gu <andrew@thinkingmachines.ai>
|
2026-07-29 09:58:18 +08:00 |
|
gjsheu
|
d86492fea0
|
[NPU] adapt dflash v2 on npu (#31739)
|
2026-07-29 09:39:19 +08:00 |
|
Lianmin Zheng
|
16a52bff23
|
[Refactor] Move sampling tokenizer validation helper (#32694)
|
2026-07-28 16:48:03 -07:00 |
|
Lianmin Zheng
|
9ca4023b13
|
[Core] Clean up array-like msgspec structs (#32688)
|
2026-07-28 16:25:22 -07:00 |
|
Mick
|
70ea37e7e0
|
vlm: reject moss vision metadata mismatches (#31957)
|
2026-07-29 07:09:54 +08:00 |
|
Xiaoyu Zhang
|
c9947b087b
|
Enable multimodal prefill BCG for VL and audio models (#30872)
|
2026-07-29 06:47:40 +08:00 |
|
Xiaoyu Zhang
|
7778dd23ea
|
[diffusion] refactor: remove stale kernels and dead code (#32651)
|
2026-07-29 06:23:23 +08:00 |
|
Void
|
7f438a6031
|
feat: SM120 (Blackwell Desktop) support for GLM-5.1 inference (#26928)
|
2026-07-28 14:52:34 -07:00 |
|
Ethan (Yusheng) Su
|
0a49226d19
|
[LoRA] 1/n Per-rank tensor serialization for load_lora_adapter_from_tensors under dp_size > 1 (#32580)
|
2026-07-28 14:28:56 -07:00 |
|
YAMY
|
dd67452b4f
|
[Cleanup] Move mamba-max-states-per-path validation into _handle_mamba_backend (#32502)
|
2026-07-28 14:21:21 -07:00 |
|
 paulzhang-tmandClaude Fable 5
|
4e5a05148a
|
[FullCG] Support chunked cached-prefix prefill (#30825)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-28 11:20:25 -07:00 |
|
 IvanShan177andClaude Opus 4.8
|
1eee8fbdcc
|
[PD] Drain NIXL completion notifications before enforcing the WaitingForInput timeout (#32267)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-29 00:20:03 +08:00 |
|
 Xinyuan TongandFAN YUCHEN
|
ee236086db
|
Fix invalid escape warnings in tool parsers (#28370)
Co-authored-by: FAN YUCHEN <2994114386@qq.com>
|
2026-07-28 23:31:47 +08:00 |
|
Mick
|
84cdfde5b2
|
[diffusion] fix: fix diffusion output stability on mps (#30017)
|
2026-07-28 21:55:03 +08:00 |
|
 
|
32c30c0f96
|
Return 400 instead of 500 for unfetchable or unparseable multimodal inputs (#31417)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-07-28 05:23:42 -07:00 |
|
Mick
|
75017c3fa0
|
[diffusion] fix: keep fused qk-norm-rope out of dynamo tracing (#31849)
|
2026-07-28 19:51:00 +08:00 |
|
Mick
|
a24906a091
|
[diffusion] feat: add dynamic cuDNN SDPA attention backend (#30090)
|
2026-07-28 19:50:24 +08:00 |
|
Talantan1102
|
5558dbad00
|
[NPU] Optimize DeepSeek-V4 performance (#31931)
|
2026-07-28 19:45:46 +08:00 |
|
 James LiuandClaude Opus 4.8
|
51397af885
|
Pack aux hidden states into a preallocated buffer (#28956)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-28 04:37:18 -07:00 |
|
Xiaoyu Zhang
|
9cffc2ba52
|
[Kernel] Remove unused implementations and stale registry entries (#32636)
|
2026-07-28 18:12:10 +08:00 |
|
 hjzhangandhjzhang
|
dde03d7c4a
|
[JIT] Restore the previous division behavior in per-token group quantization (#32616)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
|
2026-07-28 17:56:36 +08:00 |
|
Zhangheng
|
60d6914f17
|
[UnifiedTree]: move /mem_cache/unifed_cache_component dir to /mem_cache/unified_cache (#32484)
|
2026-07-28 16:28:21 +08:00 |
|
 Xinyuan Tongandliyucheng09
|
fc8b328f5c
|
[Model] Support standalone text-only Qwen3.5 checkpoints (#32401)
Co-authored-by: liyucheng09 <liyucheng09@gmail.com>
|
2026-07-28 14:08:39 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
d9cf7b0a8b
|
[MTP] Cut spec-v2 host-seam overhead in hybrid-linear MTP decode (#32219)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-07-28 11:38:33 +08:00 |
|
longxin9715
|
356c11d5d9
|
[Fix] --mm-process-config crash when video config contains (#30260)
|
2026-07-28 09:24:40 +08:00 |
|
 Lianmin Zhengandcctry
|
c19a333944
|
[mm] Handle per-item embeddings in cache misses (#32498)
Co-authored-by: cctry <cctry@meta.com>
|
2026-07-27 17:16:17 -07:00 |
|
 Caio RochaandCheng Wan
|
5a46e16f01
|
[Fix] Enable graph capture and MSCCL++ for attention TP groups (#31629)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-07-27 15:56:41 -07:00 |
|
Mohammad Miadh Angkad
|
3005af0941
|
Fix compressed-tensors NVFP4 MoE W13 layout (#32430)
|
2026-07-27 14:56:42 -07:00 |
|
Kangrui Du
|
8a311d1c88
|
[diffusion] fix: preserve tensor stride when offloading rollout weights to pinned host memory (#32420)
|
2026-07-27 12:05:13 -07:00 |
|
James Liu
|
1da062f018
|
[Inkling] Add minimal DFLASH support (#31840)
|
2026-07-27 12:01:13 -07:00 |
|
 
|
8d6549bc40
|
[Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-27 21:31:04 +08:00 |
|
inkcherry
|
5656de2d9a
|
[PD] pool decode bootstrap HTTP sessions (#31543)
|
2026-07-27 21:07:23 +08:00 |
|