 Xinyuan Tongandhnyls2002
|
85cdf1178d
|
[CI] Prune redundant CPU test overhead (#34309)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-08-13 19:51:17 -07:00 |
|
 Yuhao YangandClaude
|
6ad3f2d8fd
|
docs: link dots3.note checkpoints, add H100 cells (#34797)
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-08-14 02:42:27 +00:00 |
|
 ashwini rathiandarathi-hlab
|
456c5551cc
|
[XPU][test] Add cache_salt=None to _make_req in test_lmcache_radix_cache.py (#34728)
Co-authored-by: arathi-hlab <arathi.hlab@users.noreply.github.com>
|
2026-08-14 10:27:15 +08:00 |
|
 SII-yangdianandSII-yangdian
|
f2b2b567aa
|
perf(jit_kernel/deepseek_v4): optimize paged_mqa_metadata (#25855)
Co-authored-by: SII-yangdian <yangdian@sii.edu.cn>
|
2026-08-13 19:00:22 -07:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
704e512836
|
[GDN] Honor configured linear-attn verify backend in the kernel dispatcher (#34592)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-14 08:59:52 +08:00 |
|
Michael
|
34219ed9a7
|
[AMD] CI: pin antlr4-python3-runtime back after lmms-eval (unblocks ROCm 7.2 stage-b evals) (#34768)
|
2026-08-14 08:41:16 +08:00 |
|
Ziang Li
|
9d34c2809f
|
[FlashInfer v0.6.16] Support FlashInfer CuTe DSL NVFP4 MoE quantization (#28354)
|
2026-08-13 17:33:46 -07:00 |
|
ormandj
|
c4271c3fe1
|
[DSpark] Fix EP1 decode performance regression (#34759)
|
2026-08-13 17:00:54 -07:00 |
|
 weireweireandweireweire
|
54c44fef73
|
[Perf] Skip trivial DSV4 nonpaged indexer logits (#33857)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-08-13 16:21:30 -07:00 |
|
Liangsheng Yin
|
151a314829
|
[Fix] Make the DSpark draft num_token_non_padded host-to-device copy non-blocking (#34782)
|
2026-08-13 16:10:54 -07:00 |
|
zijiexia
|
abdef3c38e
|
[Qwen] Update Docker image tag for MI300X to v0.5.17-rocm700-mi30x-20… (#34770)
|
2026-08-13 15:59:42 -07:00 |
|
 MichaelandMichael
|
9720922671
|
[AMD][CI] Restore gfx942 Grok-1 INT4 and Grok-2 schedules (#34761)
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
|
2026-08-13 15:57:54 -07:00 |
|
 Khoa PhamandClaude Opus 5
|
f3beb2c529
|
[CI] Disable the prefill CUDA graph on the P worker of test_kimi_linear_pd_dcp4 (#34779)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-13 15:26:21 -07:00 |
|
 Khoa PhamandClaude Opus 5
|
81fe452810
|
[DCP] Share one pack kernel between both a2a backends (#34651)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-13 15:06:02 -07:00 |
|
Lianmin Zheng
|
8bbca87780
|
[Core] Organize environment variable registry (#34730)
|
2026-08-13 14:25:37 -07:00 |
|
 Khoa PhamandCursor
|
652a2709d1
|
[Docs] Add decode context parallelism to advanced features (#34654)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-13 13:51:57 -07:00 |
|
jthomson04
|
903439044a
|
perf(kv-events): coalesce cache events (#31479)
|
2026-08-13 13:36:53 -07:00 |
|
Liangsheng Yin
|
8554d9a5bc
|
[Fix] Carry the backend on Kimi-K3 deferred preprocessing configs (#34766)
|
2026-08-13 13:30:33 -07:00 |
|
Lukas Humbel
|
8ad04a9bee
|
docs(nixl): document OBJ throughput target (#30405)
|
2026-08-13 13:04:13 -07:00 |
|
 kangwangamdandBingxu Chen
|
29b067245b
|
[AMD] CI: drop the spaces from SGL_EVAL_SPEC (fixes ROCm 7.2 stage-a sgl-eval install) (#34689)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-08-13 12:48:58 -07:00 |
|
Mohammad Miadh Angkad
|
96db53ec70
|
[CI] Fix test_resolution_is_reproducible after cuda_ipc became opt-in (#34746)
|
2026-08-13 12:47:03 -07:00 |
|
 
|
6b94d39f13
|
[Model Loading] Overlap checkpoint staging with CUDA graph capture during startup (#32017)
Co-authored-by: Wenhui Zhu <wzhu59@asu.edu>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-13 12:26:25 -07:00 |
|
Shangming Cai
|
0772e79ee7
|
[CI][PD] Pin nccl rendezvous port per side to fix flaky disaggregation tests (#34755)
|
2026-08-14 01:44:20 +08:00 |
|
jambow0320
|
a82f8e1777
|
[PD] Add the missing Prefill bootstrap timeout for NIXL (#34692)
|
2026-08-14 01:08:09 +08:00 |
|
  
|
69a31ce342
|
fix: make Cache-DiT actually cache on MiniMax-H3 (#33827)
Signed-off-by: YZLi <yuanli@nvidia.com>
Signed-off-by: yunch <yunch@nvidia.com>
Co-authored-by: YZLi <yuanli@nvidia.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-14 00:39:58 +08:00 |
|
Xiaoyu Zhang
|
c255fbc4fe
|
[Diffusion] Add @triple-mu as a code owner (#34748)
|
2026-08-14 00:37:49 +08:00 |
|
triple-mu
|
ea7a6e0e99
|
fix(diffusion): unshard FSDP root group for custom encoder entry points (#34575)
|
2026-08-13 23:24:00 +08:00 |
|
Xiaoyu Zhang
|
ebca0bbde4
|
[Diffusion][ERNIE] Fuse QKNorm with full-width RoPE (#34620)
|
2026-08-13 23:23:21 +08:00 |
|
Xiaoyu Zhang
|
82f7afb881
|
[Diffusion] Make auto residency decisions component-scoped (#34615)
|
2026-08-13 23:20:55 +08:00 |
|
 
|
07821e9d56
|
[AMD][CI] Stop scheduling Grok-1 and Grok-2 on MI30x (#34643)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
|
2026-08-13 22:22:42 +08:00 |
|
 
|
c1142677a8
|
[do not merge] add new cookbooks (#34658)
Co-authored-by: Jianfei Wang <jianfei.wangg@outlook.com>
Co-authored-by: Jianfei Wang <905787410@qq.com>
|
2026-08-13 13:35:15 +00:00 |
|
Xiaoyu Zhang
|
74c0322342
|
[Diffusion][FLUX.2] Fuse eager AdaLN and packed SwiGLU (#34616)
|
2026-08-13 19:55:53 +08:00 |
|
Xiaoyu Zhang
|
3c1791a7df
|
[Diffusion][HunyuanVideo] Fuse eager QKV packing and high-quality QKNorm (#34617)
|
2026-08-13 19:54:00 +08:00 |
|
 triple-muandXiaoyu Zhang
|
993e24df75
|
[diffusion] Add --dit-layerwise-residency-policy for strided DiT residency (#34534)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-13 19:49:56 +08:00 |
|
 Liu ZhenlongandXiaoyu Zhang
|
b764194e81
|
[diffusion] model: support LongCat-Image (#23274)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-13 19:48:26 +08:00 |
|
Estrella-xx
|
eea2e5d6e5
|
Optimize delayed sample and mrope position computation (#32637)
|
2026-08-13 19:15:47 +08:00 |
|
Baizhou Zhang
|
9c0d4cba3f
|
[CI] Split Kimi K2.5 performance batches by config (#34669)
|
2026-08-13 03:20:50 -07:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
a23670ddbf
|
[diffusion] Wan2.2-TI2V: fuse per-token adaLN table add into contiguous slices + hoist rope cache (denoise -13.1% H100 / -12.6% H200, bit-exact; eager beats compile) (#34584)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-13 17:26:46 +08:00 |
|
Rain Jiang
|
ba23846ccf
|
move the PD bootstrap registry under api_server::disaggregation (#33895)
|
2026-08-13 02:13:43 -07:00 |
|
  
|
34206c0017
|
[AMD] Run V4 MTP target-verify through the decode kernel (#34597)
Co-authored-by: RolaoDenthu <xinyis10@illinois.edu>
Co-authored-by: 1am9trash <1am9trash>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
|
2026-08-13 02:00:45 -07:00 |
|
Rain Jiang
|
fd1e04d952
|
refactor error responses into shared utils::response helpers (#33894)
|
2026-08-13 01:30:29 -07:00 |
|
Rainchar9119
|
dbebc1deb4
|
[Perf] Occupancy tuning for DSA indexer fp8-quant Q kernel (#32755)
Signed-off-by: Rainchar9119 <1134601163@qq.com>
|
2026-08-13 01:15:27 -07:00 |
|
  
|
fad376d3ee
|
[CPU][QUANT] add amx cpu support for auto-round (#29593)
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Signed-off-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: Weiwei Zhang <WeiweiZhang1@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-08-13 15:50:59 +08:00 |
|
 Xinyi SongandThomas Wang
|
c034120cb8
|
[AMD] Enable draft-extend CUDA graph and reduce bubble for MTP (#29202)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-08-13 00:50:19 -07:00 |
|
 AMD-yanfeiwangandZhangheng
|
a34f81251f
|
fix(hicache/umbp): support DeepSeek-V4 hybrid HostPoolGroup (multi-po… (#30762)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-08-13 15:40:07 +08:00 |
|
jacky.cheng
|
b7f87a2513
|
[AMD][Perf] Fuse GatedDeltaNet QKVZBA split/reshape/cat into a single Triton kernel for Qwen3.5-architecture MoE on HIP (#34421)
|
2026-08-13 00:18:35 -07:00 |
|
Ke Bao
|
5a5c3d309b
|
Drop the mmlu case from the unified radix cache kit (#34667)
|
2026-08-13 15:16:05 +08:00 |
|
 MickandClaude Opus 5
|
969921b32d
|
[diffusion] chore: track minimax-h3 in the nightly diffusion benchmark (#34655)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-13 15:12:26 +08:00 |
|
Xinguo Zhu
|
3f6ef01322
|
Optimize MiniMax-M2.7 on CPU (#31956)
|
2026-08-13 15:04:16 +08:00 |
|
Gabriel Wu
|
889c2f31aa
|
feat(attention): add architecture-owned SM12x FA4 kernels (#32991)
|
2026-08-12 23:58:42 -07:00 |
|