Trevor Morris
|
20f4272109
|
fix: Fix DSR1 perf regression due to unnecessarily falling back to triton gemm (#28073)
|
2026-06-15 09:45:09 -04:00 |
|
 Prajjandprajjwal1
|
441b75ee69
|
[quantization] NVFP4 MoE: split fused w13 gate/up global scales (#27588)
Co-authored-by: prajjwal1 <prajjwal1@protonmail.com>
|
2026-06-14 21:18:36 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
0d9a2a9de3
|
[MoE Refactor] Migrate SM90 Cutlass W4A16 to MoeRunner (#26489)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-30 02:02:56 -07:00 |
|
Liangsheng Yin
|
b7d62bd724
|
[CI] Rename basic CI stage-a/b/c -> base-a/b/c for symmetry with extra CI (#25420)
|
2026-05-15 18:26:55 -07:00 |
|
  
|
0c19540550
|
[Fix] Fix gpt oss triton kernels and upgrade flashinfer back to 0.6.11.post1 (#25335)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: mmangkad <mmangkad@users.noreply.github.com>
|
2026-05-15 01:04:56 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
1913cb4dbb
|
Skip CI tests added in #24816 (broken on main) (#25329)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-14 18:12:13 -07:00 |
|
Liangsheng Yin
|
22d3f3996c
|
ci: decouple stage and runner for cuda registry (#25197)
|
2026-05-13 17:28:21 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
28758d37dd
|
Add FlashInfer SM90 cutlass MXFP4 MoE backend (W4A16) for GPT-OSS + DeepSeek-V4 (#24816)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-13 14:53:18 -07:00 |
|
 Xu Zouandxz-keg
|
ca7a8cc61d
|
[Bugfix] Fix a bug causing NVFP4 to be tested on all gpus like SM90 devices. (#24604)
Co-authored-by: xz-keg <xuzou_keg@outlook.com>
|
2026-05-08 11:51:30 -07:00 |
|
Sam (Kesen Li)
|
73e93bebd6
|
[1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954)
|
2026-04-29 01:45:48 -07:00 |
|