       
|
02236fa38c
|
Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
|
2026-07-19 22:57:37 -07:00 |
|
 
|
132ade55cd
|
[Kernel] Rewrite JIT custom all-reduce (v2) with a decoupled kernel/storage design (#31049)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: root <root@GPUC5A6.maas>
|
2026-07-17 18:37:22 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandRahul Vijayaraghavan
|
37f94cb7a0
|
[Intel GPU] DeepSeek V4 13/N: use sgl-kernel implementation of kernels in V2 Compressor to run on XPU (#28439)
Signed-off-by: P V R K Jyothendra Varma <polisettyvarma@gmail.com>
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Rahul Vijayaraghavan <rahul.vijayaraghavan@intel.com>
|
2026-07-17 09:16:04 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ee464fedc6
|
[Kernel] Migrate scattered MoE kernels to sglang.kernels (RFC #29630, Phase 2.5, 2/7) (#30786)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 09:03:21 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
874fc07d9b
|
[Kernel] Migrate scattered quantization kernels to sglang.kernels (RFC #29630, Phase 2.5, 1/7) (#30784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-13 16:17:01 +08:00 |
|
Alison Shao
|
f3c3eea608
|
ci: make multi-GPU jit test hangs attributable from the CI log (#29925)
|
2026-07-07 19:56:32 -07:00 |
|
zhangxiaolei
|
56f22cd520
|
Optimize C128 state pool allocation using request state pool (#28612)
|
2026-06-30 19:11:29 -07:00 |
|
Xinyuan Tong
|
7c23d2255a
|
[minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712)
|
2026-06-22 13:10:43 -07:00 |
|
 
|
d72314808f
|
[JIT Kernel] Multi-GPU test/bench framework for custom all-reduce + TP QKNorm (#26706)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: ziyi.xu <ziyi.xu@radixark.ai>
|
2026-06-14 17:20:36 +08:00 |
|
Cheng Wan
|
97a0031799
|
[lint] Enable Ruff UP037 to drop redundant quoted annotations (#27984)
|
2026-06-11 17:38:05 -07:00 |
|
Mohammad Miadh Angkad
|
7e245afefe
|
[CI] Fix registered sigmoid gate mul test location (#27909)
|
2026-06-11 18:52:37 +08:00 |
|
jacky.cheng
|
22c7285a26
|
[AMD] Fuse sigmoid + mul attention output gate into single Triton kernel (#27630)
|
2026-06-11 02:05:05 -07:00 |
|
Mohammad Miadh Angkad
|
bdf3ef6421
|
[CI] Fix registered QK Gemma RMSNorm test location (#27839)
|
2026-06-10 15:21:59 -07:00 |
|
jacky.cheng
|
0da18f8d91
|
[AMD][Perf] Fuse QK RMSNorm + gate extraction Triton kernel for Qwen3.5 on HIP (#27656)
|
2026-06-10 14:29:33 -07:00 |
|
Liangsheng Yin
|
186f1e300a
|
[CI] Move JIT kernel tests + benchmarks to test/registered/jit; add in-package guard (#27644)
|
2026-06-09 12:37:39 -07:00 |
|
YC Yen-Ching Tseng
|
a26587dd4e
|
[AMD][diffusion] Add FlyDSL fused normalization kernels for ROCm diffusion models optimization (#22786)
|
2026-06-08 02:42:39 -07:00 |
|
 
|
e513c13e2e
|
Optimize ngram decode token table update (#24756)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-06 14:13:45 +08:00 |
|
 Brayden ZhongandBrayden Zhong
|
38ae22e08c
|
Nemotron perf changes (#26733)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-05 22:31:46 -07:00 |
|
huangtingwei
|
00fefef16b
|
[PD & HiSparse] Add DeepSeek V4 support for HiSparse direct Prefill-to-Decode DRAM (#24880)
|
2026-06-05 15:39:48 +08:00 |
|
Shaun Kotek
|
b8d7351a74
|
Feat/add w4a16 moe support to nemotron (#25655)
|
2026-06-02 22:42:26 -07:00 |
|
Alison Shao
|
365cc2ade5
|
jit_kernel tests: bump multiprocess_test timeout 90s -> 240s (cold JIT cache) (#26994)
|
2026-06-02 17:33:52 -04:00 |
|
 
|
84e1108312
|
Optimize ngram decode id computation (#24757)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-02 17:37:34 +08:00 |
|
 Jinyan ChenandJinyan Chen
|
301bcf0872
|
Add FP4 Indexer for DeepSeek V4 (#26209)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
|
2026-06-02 00:14:38 -07:00 |
|
fzyzcjy
|
6be4b32d8d
|
Add token-id verification to the KV-canary (#26818)
|
2026-05-31 09:58:51 +08:00 |
|
fzyzcjy
|
0ca610a6df
|
Add real-data KV verification to the KV-canary (#26817)
|
2026-05-31 09:58:32 +08:00 |
|
fzyzcjy
|
30a22cc360
|
Add a periodic full-radix-tree KV-canary sweep (#26812)
|
2026-05-31 09:56:42 +08:00 |
|
fzyzcjy
|
736ad1f32a
|
Add the KV-canary plan JIT kernels (#26807)
|
2026-05-31 09:53:56 +08:00 |
|
fzyzcjy
|
70e983a2da
|
Add the KV-canary write JIT kernel and reference implementation (#26806)
|
2026-05-31 09:53:34 +08:00 |
|
fzyzcjy
|
16950954c6
|
Add the KV-canary verify JIT kernel and reference implementation (#26805)
|
2026-05-31 09:52:52 +08:00 |
|
 
|
e279b0bf72
|
Optimize large add_constant tensors (#24755)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-05-30 22:25:19 +08:00 |
|
mispa-ms
|
d616b8edad
|
[diffusion][jit_kernel] perf: varlen FA fast path for USPAttention masked branch (#26318)
|
2026-05-28 21:26:09 +08:00 |
|
Xingyu Liu
|
578d27e56a
|
[bugfix] Honor cast_x_before_out_mul in RMSNorm.forward_cuda residual path (#25920)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
|
2026-05-28 01:22:46 -07:00 |
|
Xiaoyu Zhang
|
533ef41112
|
[Diffusion] Default NVFP4 backend to FlashInfer TRTLLM (#25523)
|
2026-05-25 18:14:06 +08:00 |
|
Xiaoyu Zhang
|
ccbbae00ea
|
[codex] Reland Wan2.2 ModelOpt CI checkpoints (#25857)
|
2026-05-20 22:15:25 +08:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
8131641bc6
|
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-20 00:18:04 -07:00 |
|
 Xiaoyu ZhangandCodex
|
2424303dfb
|
[codex] Optimize hidden-size 512 RMSNorm dispatch (#24710)
Co-authored-by: Codex <codex@example.com>
|
2026-05-19 09:26:10 +08:00 |
|
Liangsheng Yin
|
b7d62bd724
|
[CI] Rename basic CI stage-a/b/c -> base-a/b/c for symmetry with extra CI (#25420)
|
2026-05-15 18:26:55 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
ee93795476
|
perf(mla): hybrid Triton fused cat+FP8-quantize for MLA chunked-prefill K/V (#25333)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-15 10:51:00 -07:00 |
|
Cheng Wan
|
dca9ba6321
|
perf(mla): TMA bulk-store set_mla_kv_buffer (up to 12× over baseline) (#25311)
|
2026-05-14 18:23:41 -07:00 |
|
 
|
e2290b155a
|
Port KV Compression V2 from deepseek_v4_dev (#24890)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: DarkSharpness <2040703891@qq.com>
|
2026-05-13 22:40:38 +08:00 |
|
YC Yen-Ching Tseng
|
cf92ccbf18
|
[AMD] Run jit kernel PR test through run_suite.py register mechanism (#24987)
|
2026-05-13 02:57:08 -07:00 |
|
Liangsheng Yin
|
eaf074d50e
|
propagate pytest exit code from test __main__ entries (#24487)
|
2026-05-06 18:46:52 -07:00 |
|
Xiaoyu Zhang
|
d86f2916cc
|
Fix diffusion fallback guards and validation (#23335)
|
2026-05-07 00:05:43 +08:00 |
|
Xiaoyu Zhang
|
b712dd48fe
|
[codex] diffusion: enable group norm silu fuse by default (#23148)
|
2026-05-02 20:55:51 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Qi Yuhangandgemini-code-assist[bot]
|
3f7c95d6cc
|
[JIT Kernel][1/2]Migrate MXFP8 Group GEMM & Quant into JIT (#23833)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-29 22:50:09 +08:00 |
|
Khoa Pham
|
ddcacaf1bd
|
Fix failing test_nvidia_nemotron_3_nano by fixing test_grouped_topk (#23874)
|
2026-04-28 15:03:58 -07:00 |
|
Qingfu Wen
|
dc1eac4903
|
[MUSA][Diffusion] Fix fa3 API on MT MUSA (#23646)
|
2026-04-28 13:01:35 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
c7878dbb6d
|
[MoE] Deprecate act_and_mul_triton; fold filter_expert into JIT silu/gelu_and_mul (#23707)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-26 01:41:35 -07:00 |
|
  
|
82254bd9c5
|
[JIT Kernel] Reland JIT activation (#22094)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-24 23:00:28 -07:00 |
|
 Jimmy ShongandSGLang CI
|
68a8ed9b11
|
[Fix/Kernel] Add JIT rmsnorm_hf kernel to fix transformers backend MMLU accuracy regression (#22931)
Co-authored-by: SGLang CI <ci@sglang.ai>
|
2026-04-23 12:00:31 +08:00 |
|