 Xiaoyu ZhangandClaude Opus 4.8
|
441910f926
|
[diffusion] LTX-2 quality=high fused RMSNorm+modulate + FFN GELU epilogue (H200 ltx23-one-stage denoise 45.85->43.24 s, ~matches torch.compile) (#34172)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-08-10 09:46:17 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
56ef810cad
|
[diffusion] BCG: auto-capture the default warmup resolution instead of hard-requiring --warmup-resolutions (H200 SANA denoise 0.73->0.457 s with a single flag) (#34174)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-08-10 08:37:32 +08:00 |
|
Xiaoyu Zhang
|
f6cbdc1dd1
|
docs(diffusion): refresh skills for latest runtime (#34143)
|
2026-08-09 10:30:46 +08:00 |
|
Xiaoyu Zhang
|
38c007dfe5
|
[diffusion] FLUX.1: route the adaLN LN+modulate sites through the bit-exact fused LayerNorm+modulate kernel (H200 1024^2 lossless denoise -1.2%, e2e wall -2.9%) (#34126)
|
2026-08-09 09:52:37 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
6424fec326
|
[diffusion] Bit-exact data-movement elimination for the Wan causal VAE decoder (H200 LongLive2 704x1280x61f: decode 2.80->2.32 s lossless / 2.12->1.67 s quality=high, e2e -10.7%) (#34125)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-09 09:50:56 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
33ed5d4413
|
[diffusion] perf_logger: SYNC_STAGE_PROFILING must drain the GPU queue for stage records too (fixes 2-3x inflated DecodingStage readings) (#34124)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-09 09:49:52 +08:00 |
|
Xiaoyu Zhang
|
dc9624deb2
|
[diffusion] Clean up kernels and shared fast paths (#34085)
|
2026-08-09 00:37:00 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
891445676c
|
[diffusion] Sana: bit-exact fused aten LayerNorm+modulate under BCG (H200 denoise -4.8%) (#34015)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 16:05:25 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
5dffa06fe1
|
[diffusion] GLM-Image bit-exact fused aten LayerNorm+modulate / qk-LN (H200 30-step denoise -8.1%) (#34008)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 13:26:11 +08:00 |
|
Xiaoyu Zhang
|
148f15b0af
|
[diffusion] FLUX.1 fused adaLN modulate (bit-exact) + RoPE cache hoist, LN-affine folding behind quality=high (H200 e2e -3.5% lossless / -6.9% high) (#34004)
|
2026-08-08 13:07:42 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
6c7498113f
|
[diffusion] Enable breakable CUDA graph for SANA (H200 1024px e2e -26%, bit-exact) (#33989)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 23:54:39 +08:00 |
|
Xiaoyu Zhang
|
d4be483efb
|
[diffusion] Enable breakable CUDA graph for LTX-2 (H200 two-stage e2e 10.75 s -> 6.90 s, 1.56x) (#33885)
|
2026-08-07 22:29:14 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
572434e2f6
|
[diffusion] Z-Image bit-exact fused qk-norm (H200 Turbo 1024px e2e -6.4%) (#33886)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 17:03:28 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
591cfb0881
|
[diffusion] FLUX.2 bit-exact residual-gate fast path (H200 klein-4B 50-step denoise -1.2%) (#33823)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 22:54:40 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
dd98c9572a
|
[diffusion] Generalize the FLUX.2 VAE decoder fast path to AutoencoderKL (Z-Image / FLUX.1) behind quality=high (#33818)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 22:53:15 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
3654740347
|
[diffusion] ERNIE-Image bit-exact fused RMSNorm+scale/shift (H200 1024^2 e2e 15.63 -> 15.00 s, denoise -3.3%) (#33854)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 19:58:44 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
295784723a
|
[diffusion] Ideogram 4: fuse RMSNorm modulate/gate chains via the Z-Image Triton suite behind quality=high (H200 e2e -2.9%/-3.4%) (#33822)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 19:57:34 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
eff6a11350
|
[diffusion] FLUX.1 bit-exact residual-gate fast path + tanh-GELU epilogue behind quality=high (H200 e2e -1.1% lossless / -4.3% high) (#33819)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 19:56:38 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
b6876fc652
|
[diffusion] ERNIE-Image bit-exact residual-gate fast path (H200 1024^2 e2e 16.17 -> 15.75 s) (#33734)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 13:43:09 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4c0a8940fa
|
[Kernel] Unify BaseFusedOp and MultiPlatformOp dispatch (#33205)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 08:52:09 +08:00 |
|
 Xiaoyu ZhangandMohammad Miadh Angkad
|
ba12a16a62
|
[diffusion] Prefer cuDNN SDPA over FA4 for dense attention on sm_100 (B200) (#33655)
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-08-06 08:49:10 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
3425c93666
|
[diffusion] Wan VAE RMSNorm+SiLU fusion behind quality=high (H200 FastWan2.2 e2e 9.611 -> 9.125 s) (#33546)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-05 21:33:35 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
a5888c956f
|
[diffusion] Pack Ulysses Q/K/V input all-to-all into one collective + reusable a2a staging buffers (#33667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-05 19:15:46 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
95d0e57e83
|
[diffusion] Fuse DiT FFN tanh-GELU into up-proj GEMM (cublasLt epilogue) behind quality=high (Qwen-Image 1024^2 denoise 12.36 -> 12.05 s on H200) (#33536)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-04 23:49:43 +08:00 |
|
Xiaoyu Zhang
|
0d0c7d853f
|
[diffusion] FLUX.2 VAE decoder fast path behind quality=high (H200: 1024^2 97.6->29.2 ms, 2048^2 437.2->168.5 ms) (#33451)
|
2026-08-04 23:48:19 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
c6f2a9c1d4
|
[diffusion] Restrict request-level quality to two validated tiers: lossless (default) and high (#33453)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-04 11:43:16 +08:00 |
|
Xiaoyu Zhang
|
f829fb3ff7
|
[diffusion] Fix component accuracy topology reuse (#33317)
|
2026-08-04 08:35:03 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4ef1660cd8
|
[Docs] MiniMax-H3: add measured H200 Ulysses4 vs TP2+Ulysses2 topology data (#33398)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-03 12:54:29 -07:00 |
|
Xiaoyu Zhang
|
b64fd800d4
|
docs(diffusion): update skills for MiniMax-H3 (#33282)
|
2026-08-03 12:44:01 +08:00 |
|
  
|
fb207b72b0
|
feat(kernels): port standalone Kimi K3 kernels (#32890)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zhangxiaolei <zhangxiaolei.666@bytedance.com>
|
2026-08-01 13:26:47 +08:00 |
|
Xiaoyu Zhang
|
c5bd3d7dce
|
[diffusion][benchmark] Add reproducible request-manifest offline benchmark (#32917)
|
2026-07-30 22:11:18 +08:00 |
|
Xiaoyu Zhang
|
7784ac8f91
|
[diffusion][docs] Fix Cosmos3 model sizes (#32916)
|
2026-07-30 22:10:23 +08:00 |
|
Xiaoyu Zhang
|
2e9c82b359
|
[Kernel] Remove unreachable AOT headers (#32842)
|
2026-07-30 22:08:57 +08:00 |
|
Xiaoyu Zhang
|
1d9c292547
|
[Kernel] Add inventory guards and clean benchmark layout (#32788)
|
2026-07-30 09:03:24 +08:00 |
|
Xiaoyu Zhang
|
0ebbe43dbb
|
fix(diffusion): size VSA top-k from padded blocks (#32695)
|
2026-07-29 21:58:41 +08:00 |
|
Xiaoyu Zhang
|
4f5b50c576
|
perf(diffusion): decode Wan VAE in BF16 (#32697)
|
2026-07-29 21:57:50 +08:00 |
|
Xiaoyu Zhang
|
917e900d4d
|
feat(diffusion): add regional torch compile (#32696)
|
2026-07-29 21:57:05 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
8742a1a0f8
|
Add a benchmark script for the HPC-Ops bf16xfp32 router GEMM (#32642)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-29 18:17:08 +08:00 |
|
Xiaoyu Zhang
|
c32c4ef79c
|
[Kernel] Move sgl-kernel under sglang.kernels.aot (#32648)
|
2026-07-29 17:25:00 +08:00 |
|
Xiaoyu Zhang
|
c9947b087b
|
Enable multimodal prefill BCG for VL and audio models (#30872)
|
2026-07-29 06:47:40 +08:00 |
|
Xiaoyu Zhang
|
7778dd23ea
|
[diffusion] refactor: remove stale kernels and dead code (#32651)
|
2026-07-29 06:23:23 +08:00 |
|
Xiaoyu Zhang
|
9cffc2ba52
|
[Kernel] Remove unused implementations and stale registry entries (#32636)
|
2026-07-28 18:12:10 +08:00 |
|
 
|
8d6549bc40
|
[Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-27 21:31:04 +08:00 |
|
 
|
3d91a569ce
|
[MoE Backend] Add HPC-Ops FP8 MoE runner backend (#30541)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-24 19:46:11 +08:00 |
|
 
|
841fa293b5
|
[Fix] Reject online weight updates while the HPC-Ops router GEMM split cache is active (#31943)
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-24 19:30:13 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
d4a0dfbc31
|
[Fix] Two root causes of the H100 deepep TBO CI break: scale-tensor use-after-free + missing non-finite quant sanitization (#32188)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-24 07:29:28 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
62aa85d9aa
|
[Kernel] Sweep missed dedicated kernels into kernels.ops (moe/quant siblings + dspark) (RFC #29630) (#32160)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 17:07:16 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
11b0e5c5ad
|
[Kernel] Classification cleanup: unify _jit_ naming, drop empty/model groups, add elementwise (RFC #29630) (#32148)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 13:47:02 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
2d1a7be8c4
|
[Kernel] Reclassify kernel tests by ops group + move helpers out of the package (RFC #29630) (#32128)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 12:18:27 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
99f636a86f
|
[Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 08:35:09 +08:00 |
|
 
|
0a6d1930c3
|
[Attention Backend] Add HPC-Ops attention backend (#30540)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-22 22:06:22 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
74338e94f1
|
[Kernel] Phase 4 batch-3: migrate tangled JIT subsystems + new groups into kernels.ops (RFC #29630) (#32045)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-22 21:15:03 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
977ea336cd
|
[Kernel] Phase 4 batch-2: migrate JIT operator groups into kernels.ops (no shims) (RFC #29630) (#32015)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-22 17:49:51 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
246b3c3eaf
|
[Kernel] Phase 3+4: move JIT infra + operator groups into sglang.kernels (RFC #29630) (#31666)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-22 11:28:27 +08:00 |
|
Xiaoyu Zhang
|
075bd97952
|
[Benchmark] Remove obsolete auto-benchmark remnants (#31941)
|
2026-07-21 20:44:52 +08:00 |
|
 
|
e4eea7ce2f
|
Optimize LongCat-Flash router GEMM with the HPC-Ops bf16xfp32 kernel (#30247)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-21 20:05:17 +08:00 |
|
Xiaoyu Zhang
|
829e9ce9d5
|
Lower AutoRound quantization MMLU threshold (#31748)
|
2026-07-20 13:40:48 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
216b750c8f
|
[Kernel] Sweep decoupled scattered kernels into sglang.kernels.ops (RFC #29630) (#31582)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-18 19:07:07 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
619609aa5a
|
[Kernel] Simplify sglang.kernels tests to idiomatic pytest style (RFC #29630) (#31546)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-17 15:06:26 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
1ac1ffea0c
|
[Kernel] Fill non-CUDA coverage: HIP (aiter/rocm-triton) + Ascend NPU backends (RFC #29630) (#31307)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-17 14:05:35 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
8432eafd3d
|
[Kernel] Decouple KernelBackend from device + device-based CapabilityRequirement (RFC #29630) (#31292)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-17 10:35:34 +08:00 |
|
Xiaoyu Zhang
|
e73f323464
|
[JIT] Reduce MoE fused gate CI test sweep (#31400)
|
2026-07-16 14:42:45 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
22453ca63c
|
docker: build HPC-Ops into the GPU image (#31390)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-16 11:07:01 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4aadf94146
|
[Kernel] Relocate vendored fla and mamba kernel trees to sglang.kernels (RFC #29630, Phase 2.5, 7/7) (#30795)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 12:52:15 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
c00131ebaa
|
[Kernel] Migrate linear-attention, MiniMax-sparse and diffusion kernels to sglang.kernels (RFC #29630, Phase 2.5, 6/7) (#30793)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 11:21:36 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ba5be86d42
|
[Kernel] Migrate DSA + DSV4 attention kernels to sglang.kernels (RFC #29630, Phase 2.5, 5/7) (#30792)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 11:11:22 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ee000f6734
|
[CI] Fix SGLANG_JIT_KERNEL_RUN_FULL_TESTS never activating the nightly full jit-kernel sweep (#31042)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 17:32:29 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
1a35440c4a
|
[Kernel] Migrate generic attention kernels to sglang.kernels (RFC #29630, Phase 2.5, 4/7) (#30789)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 16:53:46 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
e9ef06c560
|
[Kernel] Migrate top-level srt/layers stray kernels to sglang.kernels (RFC #29630, Phase 2.5, 3/7) (#30787)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 09:20:59 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ee464fedc6
|
[Kernel] Migrate scattered MoE kernels to sglang.kernels (RFC #29630, Phase 2.5, 2/7) (#30786)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 09:03:21 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
4c997310f5
|
[Kernel] Hotfix: update sgl-kernel imports of relocated fp8_kernel (RFC #29630 #30784) (#31089)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-14 08:41:23 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
874fc07d9b
|
[Kernel] Migrate scattered quantization kernels to sglang.kernels (RFC #29630, Phase 2.5, 1/7) (#30784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-13 16:17:01 +08:00 |
|
Xiaoyu Zhang
|
65abb23842
|
Add diffusion BCG prompt conditioning guard (#30782)
|
2026-07-11 13:14:31 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
6ed9843b57
|
[Kernel] Introduce sglang.kernels namespace and migrate scattered triton_ops kernels (RFC #29630, Phase 2) (#30044)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-10 21:41:08 +08:00 |
|
Xiaoyu Zhang
|
e9493a015c
|
Fix diffusion BCG lifetime and add Z-Image-Turbo CI (#30584)
|
2026-07-10 21:19:59 +08:00 |
|
Xiaoyu Zhang
|
b8ca06fdad
|
Fix zero expert routed ids for MoE backends (#30387)
|
2026-07-08 21:23:25 +08:00 |
|
  
|
33c3dfd7e0
|
[diffusion] Enable breakable CUDA graph (BCG) for diffusion DiTs (#27436)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: BBuf <bbuf@sglang.local>
|
2026-07-08 14:45:48 +08:00 |
|
 Xiaoyu ZhangandZijie Xia
|
ead1e490b5
|
[Doc] Add LongCat 2.0 FP8 cookbook (#30320)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-07-07 11:48:13 -07:00 |
|
 
|
e339c83f82
|
[Model] Support LongCat 2.0 FP8 (#30275)
Co-authored-by: sunjiaqi11 <sunjiaqi11@meituan.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-07-07 19:51:12 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
1da7d3a50b
|
[MoE] Retire the AOT moe_fused_gate / kimi_k2_moe_fused_gate gate kernels (#26771) (#29997)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-07 13:53:17 +08:00 |
|
Xiaoyu Zhang
|
931b00f1b0
|
[diffusion] Clean up duplicate helper definitions (#30159)
|
2026-07-05 22:05:11 +08:00 |
|
Xiaoyu Zhang
|
6dd0cefb2a
|
[CI] Revert ModelOpt NVFP4 threshold relax (#29844)
|
2026-07-04 21:17:32 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
a2d7eb303e
|
[MoE] Consolidate ungrouped + grouped gate/topk onto one Triton router (#26771) — faster than AOT on B200/H100/H200, at parity with flashinfer (#29771)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-03 11:18:59 +08:00 |
|
Xiaoyu Zhang
|
b276a9acee
|
chore: cleanup garbage code (#29770)
|
2026-07-02 16:14:01 +08:00 |
|
  
|
df0dfbaa45
|
[Kernel] Strengthen kernel shape coverage (#29636)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-07-01 15:44:20 +08:00 |
|
Xiaoyu Zhang
|
8205aa3603
|
chore: clean diffusion dead code (#29789)
|
2026-07-01 15:42:16 +08:00 |
|
Xiaoyu Zhang
|
47ae1241d3
|
[CI] Relax ModelOpt NVFP4 diffusion consistency thresholds (#29767)
|
2026-07-01 14:45:08 +08:00 |
|
Xiaoyu Zhang
|
fcb9f229b3
|
[KDA-Pilot] Add LTX2 QKNorm split-RoPE CUDA fast path (#29708)
|
2026-07-01 14:42:07 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
3add35e26d
|
[Diffusion] Reuse shared AlignedVector and tidy jit_kernel/diffusion (#29664)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-30 11:38:22 +08:00 |
|
Xiaoyu Zhang
|
c36f166364
|
[skill] Remove outdated llm-serving-auto-benchmark skill (#29487)
|
2026-06-27 14:19:11 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
495f13fa12
|
[KDA-Pilot] Add diffusion residual-gate CUDA fast path for LTX2 (#29361)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-27 12:59:41 +08:00 |
|
Xiaoyu Zhang
|
8524678889
|
Fix MiniMax MSA fallback when fmha plan is unavailable (#29250)
|
2026-06-26 23:14:31 +08:00 |
|
Xiaoyu Zhang
|
18b0e5757e
|
[Diffusion] Fuse LTX2 Ada values (#29390)
|
2026-06-26 23:13:15 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
5996b54bd3
|
[KDA-Pilot] Add diffusion causal Conv3D cat-pad CUDA fast path for Cosmos3 (#29281)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-26 15:06:49 +08:00 |
|
Xiaoyu Zhang
|
4d06d4c97f
|
Sync Gemma4 hardware table with Blackwell recipes (#29266)
|
2026-06-25 22:57:33 +08:00 |
|
 
|
52c32035eb
|
[diffusion] Add Qwen-Image ModelOpt NVFP4 support (#28928)
Co-authored-by: jingyu-ml <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-25 22:56:09 +08:00 |
|
Xiaoyu Zhang
|
ddda4f9028
|
[Bugfix] Fix Ministral3 init argument forwarding (#29111)
|
2026-06-25 18:26:45 +08:00 |
|
Xiaoyu Zhang
|
7c9804ef21
|
Add MiMo V2.5 Blackwell vision FA4 recipe (#29253)
|
2026-06-25 13:47:32 +08:00 |
|
Xiaoyu Zhang
|
efbe67d237
|
Tune Gemma4 26B-A4B B200 memory recipe (#29252)
|
2026-06-25 11:31:25 +08:00 |
|
Xiaoyu Zhang
|
26e1d4d847
|
[KDA-Pilot] Add B200 diffusion norm-scale-shift CUDA fast path for Qwen-Image (#27392)
|
2026-06-24 14:38:36 +08:00 |
|