Xiaoyu Zhang
|
0e4d1b49d3
|
[Codex] Remove stale DeepSeek V4 JIT kernels (#25764)
|
2026-05-19 20:04:32 +08:00 |
|
Kevin Li
|
fbfddfd5c7
|
fix (jit kernel): elementwise activation C++ error (#25695)
|
2026-05-19 15:23:52 +08:00 |
|
 Xiaoyu ZhangandCodex
|
2424303dfb
|
[codex] Optimize hidden-size 512 RMSNorm dispatch (#24710)
Co-authored-by: Codex <codex@example.com>
|
2026-05-19 09:26:10 +08:00 |
|
+3        
|
866793c502
|
Amd/deepseek v4 rebase main 0509 (#24933)
Co-authored-by: root <root@smci355-ccs-aus-m12-33.cs-aus.dcgpu>
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-18 09:15:07 -07:00 |
|
 Qingfu WenandR0CKSTAR
|
3bf7e346fc
|
[MUSA][Diffusion] Improve wan model inference speed using torch.compile (#25256)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-05-17 22:10:24 +08:00 |
|
Xiaoyu Zhang
|
93bacc25ed
|
[codex] Optimize LTX2 split rotary kernel (#24732)
|
2026-05-16 20:58:38 +08:00 |
|
Liangsheng Yin
|
b7d62bd724
|
[CI] Rename basic CI stage-a/b/c -> base-a/b/c for symmetry with extra CI (#25420)
|
2026-05-15 18:26:55 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
ee93795476
|
perf(mla): hybrid Triton fused cat+FP8-quantize for MLA chunked-prefill K/V (#25333)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-15 10:51:00 -07:00 |
|
Cheng Wan
|
dca9ba6321
|
perf(mla): TMA bulk-store set_mla_kv_buffer (up to 12× over baseline) (#25311)
|
2026-05-14 18:23:41 -07:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
bc265c5f82
|
[AMD] Add amd jit resolve token ids bench ci (#25210)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
|
2026-05-14 00:13:12 -07:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
7b128e143a
|
[AMD] Add amd jit clamp position bench ci (#25209)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
|
2026-05-14 00:03:56 -07:00 |
|
 
|
34c0029f0a
|
[diffusion] [AMD] feat: support online MXFP4 and fp8 quantization (#21431)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-05-14 08:52:01 +08:00 |
|
shiyu7
|
37f18438c5
|
[rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper (#24986)
|
2026-05-13 16:33:46 -07:00 |
|
 
|
e2290b155a
|
Port KV Compression V2 from deepseek_v4_dev (#24890)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: DarkSharpness <2040703891@qq.com>
|
2026-05-13 22:40:38 +08:00 |
|
YC Yen-Ching Tseng
|
cf92ccbf18
|
[AMD] Run jit kernel PR test through run_suite.py register mechanism (#24987)
|
2026-05-13 02:57:08 -07:00 |
|
 
|
e9dea79755
|
(3/n - prefill optimize)[LoRA][MoE] Optimize virtual experts: remove CPU-GPU sync & multi-block CUDA JIT histogram (#24262)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-11 16:36:57 -07:00 |
|
Mick
|
73b8eda103
|
[diffusion] fix: fix FA3 varlen out argument handling (#24688)
|
2026-05-08 19:01:49 +08:00 |
|
+6        
|
35870d55ac
|
Deepseek V4 (#23882)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Qiaolin Yu <90088090+qiaolin-yu@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <11704492+yushengsu-thu@users.noreply.github.com>
Co-authored-by: Mingyi <27337995+wisclmy0611@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+againstentropy@users.noreply.github.com>
|
2026-05-07 18:32:21 -07:00 |
|
Liangsheng Yin
|
eaf074d50e
|
propagate pytest exit code from test __main__ entries (#24487)
|
2026-05-06 18:46:52 -07:00 |
|
Xiaoyu Zhang
|
d86f2916cc
|
Fix diffusion fallback guards and validation (#23335)
|
2026-05-07 00:05:43 +08:00 |
|
Mick
|
177babcc38
|
[diffusion] optimize: fuse LTX2 split rotary embedding (#24411)
|
2026-05-05 16:07:40 +08:00 |
|
Liangsheng Yin
|
84f3b44916
|
[tiny] misc cleanups across configs, attention, jit_kernel (#24350)
|
2026-05-04 03:17:14 -07:00 |
|
Xiaoyu Zhang
|
b712dd48fe
|
[codex] diffusion: enable group norm silu fuse by default (#23148)
|
2026-05-02 20:55:51 +08:00 |
|
Xiaoyu Zhang
|
1360848ee1
|
Optimize large GroupNorm SiLU apply (#23938)
|
2026-05-02 20:54:46 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Qi Yuhangandgemini-code-assist[bot]
|
3f7c95d6cc
|
[JIT Kernel][1/2]Migrate MXFP8 Group GEMM & Quant into JIT (#23833)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-29 22:50:09 +08:00 |
|
Khoa Pham
|
ddcacaf1bd
|
Fix failing test_nvidia_nemotron_3_nano by fixing test_grouped_topk (#23874)
|
2026-04-28 15:03:58 -07:00 |
|
Qingfu Wen
|
dc1eac4903
|
[MUSA][Diffusion] Fix fa3 API on MT MUSA (#23646)
|
2026-04-28 13:01:35 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
c7878dbb6d
|
[MoE] Deprecate act_and_mul_triton; fold filter_expert into JIT silu/gelu_and_mul (#23707)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-26 01:41:35 -07:00 |
|
Mick
|
03849496ad
|
jit_kernel: tolerate FA3 kernels without out arg (#23717)
|
2026-04-25 23:42:33 +08:00 |
|
 Артем Савкинandronnie_zheng
|
bd523dd60d
|
[NPU] [Bugfix] [Diffusion] Fixed gray images at the generation output (#23266)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-25 10:20:38 +03:00 |
|
  
|
82254bd9c5
|
[JIT Kernel] Reland JIT activation (#22094)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-24 23:00:28 -07:00 |
|
 Jia GuoandClaude Opus 4.6
|
587fd15bd2
|
perf: eliminate attention DtoD copy by passing pre-allocated output to FA (#21985)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-24 12:05:16 -07:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
 Jimmy ShongandSGLang CI
|
68a8ed9b11
|
[Fix/Kernel] Add JIT rmsnorm_hf kernel to fix transformers backend MMLU accuracy regression (#22931)
Co-authored-by: SGLang CI <ci@sglang.ai>
|
2026-04-23 12:00:31 +08:00 |
|
Yanbin Jiang
|
917d2aa1dc
|
[LoRA] Fix EP + per-expert MoE LoRA illegal memory access (#23178)
|
2026-04-22 14:22:32 -07:00 |
|
jianan-gu
|
2cf3ac515b
|
[Diffusion][CPU] Init CPU platform support for SGLang Diffusion (#20816)
|
2026-04-21 14:25:54 +08:00 |
|
Yuhao Yang
|
5595f6e988
|
Fix trtllm mla chunked-prefill zero-length bug (#22291) (#22688)
|
2026-04-20 22:10:13 -07:00 |
|
Liangsheng Yin
|
6cc2eee50d
|
[misc] CI hygiene: enforce __main__ entry, drop silent-skipped tests, fix rerun-test protoc (#23305)
|
2026-04-20 21:16:24 -07:00 |
|
   
|
6ecd6f84db
|
[CI] Add per-job uv venv isolation and upgrade CI version to Cuda 13 (#23119)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-19 05:32:36 -07:00 |
|
Xiaoyu Zhang
|
cd6ad80c00
|
diffusion: add HunyuanVideo GroupNorm+SiLU fast path (#22814)
|
2026-04-18 23:38:49 +08:00 |
|
Xiaoyu Zhang
|
615d6c93b2
|
[codex] Add flashinfer TRTLLM backend for diffusion NVFP4 (#22717)
|
2026-04-18 09:06:28 +08:00 |
|
Xiaoyu Zhang
|
f97c608caa
|
[diffusion] quant: add FLUX.1-dev modelopt nvfp4 support (#22672)
|
2026-04-14 15:00:59 +08:00 |
|
Lianmin Zheng
|
f81b6e8f51
|
[Misc] Add @cache_once to is_arch_support_pdl in jit_kernel (#22724)
|
2026-04-13 14:42:49 -07:00 |
|
 DarkSharpnessandMingyang Jiang
|
314d6ecf08
|
[Feature][JIT Kernel] Fused TP QK norm For Minimax (#20673)
Co-authored-by: Mingyang Jiang <13463932+jmydurant@users.noreply.github.com>
|
2026-04-13 20:29:47 +08:00 |
|
Zhangheng
|
5549d910c6
|
[hisparse]: Adding ci for hisparse kvcache-swap-in jit-kernel (#22155)
|
2026-04-13 12:50:29 +08:00 |
|
Zhangheng
|
305b42935a
|
[HiSparse]: Add benchmark for hisparse kernel (#22187)
|
2026-04-13 12:49:18 +08:00 |
|
Mick
|
bf022e177c
|
Revert "[Diffusion] Add FLUX.1-dev ModelOpt NVFP4 support (#22574)" (#22649)
|
2026-04-13 11:17:32 +08:00 |
|
Xiaoyu Zhang
|
03a1a7b81c
|
[Diffusion] Add FLUX.1-dev ModelOpt NVFP4 support (#22574)
|
2026-04-13 07:57:41 +08:00 |
|
 sushil DubeyandMa Mingfei
|
e26c73c4e9
|
[diffusion] platform: support Intel XPU (#17920)
Signed-off-by: sushil.dubey <sushil.dubey@intel.com>
Signed-off-by: Sushil Dubey <sushil.dubey@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-11 15:09:02 +08:00 |
|
Khoa Pham
|
aeeff58cd4
|
[Spec][Ngram] Clean up unused stateless batchMatch (#22487)
|
2026-04-10 21:52:56 -07:00 |
|