Baizhou Zhang
|
2e70e4f4f6
|
[CI] Little renaming of gb200 CI workflow (#22608)
|
2026-04-11 17:52:42 -07:00 |
|
Baizhou Zhang
|
d14d368191
|
[Kernel] Set sgl_per_token_group_quant_8bit_v2 as default choice (#22467)
|
2026-04-11 01:59:57 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
3c46ff2ac5
|
fix: restore CPU flash_attn test to use sgl_kernel directly (#22573)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:39:20 -07:00 |
|
Baizhou Zhang
|
60acdc31f2
|
[Fix] Fix several bugs on DSA models (#22430)
|
2026-04-09 12:46:23 -07:00 |
|
Baizhou Zhang
|
606aa11ea8
|
[DSA] Enable all reduce fusion for DSA models (#22390)
|
2026-04-09 12:42:44 -07:00 |
|
Baizhou Zhang
|
5e1a9834f9
|
Update ci permission (#22387)
|
2026-04-08 15:42:06 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
4e5b8cb041
|
Fix get_version_tag.py to handle dot-separated post versions (#22385)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-08 15:18:22 -07:00 |
|
Baizhou Zhang
|
213af1d4f7
|
Add CI tests for GLM-5 (#22285)
|
2026-04-08 01:05:36 -07:00 |
|
Baizhou Zhang
|
20ee59bcfc
|
[Misc] Remove unused cu13 docker release workflow (#22167)
|
2026-04-05 16:01:02 -07:00 |
|
Baizhou Zhang
|
c5fa364b80
|
[Hotfix] Fix router gemm on sm103 (#22134)
|
2026-04-05 09:33:14 -07:00 |
|
Baizhou Zhang
|
106baedbfb
|
[Doc] Update GLM-5 instructions in sglang documentation (#21716)
|
2026-04-05 03:13:07 -07:00 |
|
Baizhou Zhang
|
088203454b
|
[Fix] Fix nightly tests (#22140)
|
2026-04-05 02:26:42 -07:00 |
|
Baizhou Zhang
|
723ed6c3c4
|
[CI]Temporary ban auto benchmark tool test (#22138)
|
2026-04-04 23:18:19 -07:00 |
|
Baizhou Zhang
|
bf984ae65d
|
Revert "[Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang" (#22098)
|
2026-04-04 02:17:19 -07:00 |
|
Baizhou Zhang
|
ac1e437f6a
|
Revert "[Feature] JIT activation and update skills (by codex)" (#22078)
|
2026-04-03 15:04:15 -07:00 |
|
Baizhou Zhang
|
97adf8a290
|
[misc] Add hint for kernel release trigger (#22036)
|
2026-04-03 03:31:44 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
98ac40192b
|
[Workflow] Fix kernel release build failures for aarch64 and wheel renaming (#22018)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-03 03:23:03 -07:00 |
|
Baizhou Zhang
|
75de479680
|
[Misc] Update CI permission (#22014)
|
2026-04-02 23:37:05 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
5c082c307a
|
[Workflow] Fix kernel release jobs skipped on push events (#22011)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-02 23:03:13 -07:00 |
|
Baizhou Zhang
|
0a709cfe02
|
[Workflow] Avoid triggering nightly tests in kernel bump workflow (#22010)
|
2026-04-02 22:40:33 -07:00 |
|
Baizhou Zhang
|
efa7b2d5d3
|
Revert "[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine)" (#22002)
|
2026-04-02 20:42:13 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
29d8e959d7
|
[CI] Remove stale Ascend suite entries from test/srt/run_suite.py (#21978)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-02 16:47:19 -07:00 |
|
Baizhou Zhang
|
c7d03a6215
|
Revert "Rollback flashmla to older version [1/2]" (#21922)
|
2026-04-02 00:27:02 -07:00 |
|
Baizhou Zhang
|
fbc1f92453
|
[DSA] Set trtllm kernels as nsa default for Blackwell (#21914)
|
2026-04-02 00:22:27 -07:00 |
|
Baizhou Zhang
|
5e12c4e08e
|
[DSA] Support trtllm sparse mla kernel for prefill batches (#21783)
|
2026-04-01 13:55:05 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
f60f2ccc10
|
[Fix] Fall back to triton MOE for GPT-OSS on Blackwell with driver >= 595 (#21780)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-31 15:52:10 -07:00 |
|
Baizhou Zhang
|
d52757fe97
|
[CI]Remove msgm-en and mmlu tests which cause timeout (#21733)
|
2026-03-31 01:10:05 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
62a63eeff7
|
[Fix] Fix weight_loader property assignment for qwen3-next FP8 models (#21662)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-30 01:35:59 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Baizhou Zhangandgemini-code-assist[bot]
|
5b19c9a05d
|
[Doc] Update tips for developer new-comers (#21659)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-29 22:40:36 -07:00 |
|
Baizhou Zhang
|
edd4d54023
|
[Clean] Remove deprecated environs (#21536)
|
2026-03-28 00:35:44 -07:00 |
|
Baizhou Zhang
|
6ef4318ec0
|
[CI] Move v32 cp test to deepep running suite (#21585)
|
2026-03-27 22:49:06 -07:00 |
|
Baizhou Zhang
|
ec29bbb286
|
Split workflow for releasing runtime docker (#21563)
|
2026-03-27 15:05:52 -07:00 |
|
Baizhou Zhang
|
4e905febd2
|
[CI] Relax several thresholds in flaky CIs (#21562)
|
2026-03-27 13:16:49 -07:00 |
|
Baizhou Zhang
|
0138129d3c
|
[CI] Fix nemotron nvfp4 test estimated time (#21516)
|
2026-03-26 21:53:09 -07:00 |
|
Baizhou Zhang
|
a93065679b
|
Revert "bugfix for weight loading for qwen3-next" (#21496)
|
2026-03-26 16:17:18 -07:00 |
|
Baizhou Zhang
|
dbe871efdd
|
Rollback flashmla to older version [1/2] (#21430)
|
2026-03-25 17:49:54 -07:00 |
|
Baizhou Zhang
|
e90cba715c
|
Revert "[Bugfix] Disable ci for .md files" (#21420)
|
2026-03-25 12:00:20 -07:00 |
|
Baizhou Zhang
|
f5c225eeba
|
[CI] Fix TestQwen35WithHiCache (#21371)
|
2026-03-25 00:04:59 -07:00 |
|
Baizhou Zhang
|
2b75fed0dd
|
Workaround of DSA performance drop on B200 + DP (#21337)
|
2026-03-24 22:21:07 -07:00 |
|
Baizhou Zhang
|
1046dbe038
|
[Fix] Fix trtllm fp4 moe kernel not found error (#21343)
|
2026-03-24 16:38:05 -07:00 |
|
Baizhou Zhang
|
d173cfecd6
|
Update access of cut branch workflow and delete deprecated release workflow (#21231)
|
2026-03-23 15:18:10 -07:00 |
|
Baizhou Zhang
|
ed316a26ef
|
Fix CP in-seq-split method for DeepSeek V32 and update related tests (#21192)
|
2026-03-23 12:34:10 -07:00 |
|
Baizhou Zhang
|
9ca68a5904
|
Remove flaky test for test_lora_update.py (#21088)
|
2026-03-21 00:52:40 -07:00 |
|
Baizhou Zhang
|
67cad3e69e
|
Revert "Support CuteDSL mm_fp4 backend" (#21077)
|
2026-03-20 22:47:47 -07:00 |
|
Baizhou Zhang
|
5f3393c04c
|
Fix deepseek-v32-fp4 b200 ci (#21072)
|
2026-03-20 22:28:40 -07:00 |
|
Baizhou Zhang
|
c82d20d48e
|
Fix DeepSeek V32 FP4 test (#20984)
|
2026-03-20 01:04:32 -07:00 |
|
Baizhou Zhang
|
42f4b7276c
|
Revert "feat(mm)(grpc): compute M-RoPE positions for preprocessed VL inputs" (#20956)
|
2026-03-19 18:03:04 -07:00 |
|
Baizhou Zhang
|
826eb21bca
|
Fix sglang-kernel dependency on CI runners (#20715)
|
2026-03-16 15:23:01 -07:00 |
|
Baizhou Zhang
|
e96a3752a0
|
Update timeout for some ci tests (#20664)
|
2026-03-15 23:58:30 -07:00 |
|
Baizhou Zhang
|
39008955ff
|
Revert "[AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars." (#20602)
|
2026-03-14 12:12:42 -07:00 |
|
Baizhou Zhang
|
f8668d9e78
|
[Fix] Add fallback for flashinfer allreduce fusion (#20384)
|
2026-03-13 01:24:55 -07:00 |
|
Baizhou Zhang
|
7771fcd4b3
|
Update CodeOwners (#20329)
|
2026-03-11 19:11:09 -07:00 |
|
Baizhou Zhang
|
a6ae89fe3c
|
Revert "chore: bump sgl-kernel version to 0.3.21.post1" (#20229)
|
2026-03-09 20:32:19 -07:00 |
|
Baizhou Zhang
|
be63f982b7
|
[V32/GLM5] Control the threshold of applying dense attention with an environ (#20062)
|
2026-03-09 14:36:10 -07:00 |
|
Baizhou Zhang
|
61d530e8ac
|
[CI] Fix lint (#20209)
|
2026-03-09 14:09:59 -07:00 |
|
Baizhou Zhang
|
d28f35240a
|
[V32/GLM5] Change default setting of V32 nvfp4 on TP4 (#20086)
|
2026-03-07 15:13:25 -08:00 |
|
Baizhou Zhang
|
04e364d538
|
[V32] Enhance deepseek v32 related tests (#19985)
|
2026-03-05 20:12:49 -08:00 |
|
Baizhou Zhang
|
51e5dc845a
|
Revert "[Kernel Slimming] Migrate NVFP4 kernels to JIT" (#20005)
|
2026-03-05 19:40:00 -08:00 |
|
Baizhou Zhang
|
10c65df48a
|
[Bug] Fix lora tp bug on H200 (#19769)
|
2026-03-04 20:11:02 -08:00 |
|
Baizhou Zhang
|
f07d668ba1
|
Remove flashinfer version argument from cu13 docker release workflow (#19862)
|
2026-03-04 00:26:06 -08:00 |
|
Baizhou Zhang
|
52dcade4aa
|
Fix flashinfer bump workflow (#19855)
|
2026-03-04 00:09:58 -08:00 |
|
Baizhou Zhang
|
78ddf05afd
|
[Fix] Install tomli in flashinfer bumping workflow (#19841)
|
2026-03-03 23:12:22 -08:00 |
|
 Baizhou ZhangandClaude
|
c287d9b645
|
chore: add flashinfer version bump workflow (#19837)
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-03-03 22:53:34 -08:00 |
|
Baizhou Zhang
|
776709efe8
|
[3/n] deepseek_v2.py Refactor: Migrate MLA forward method in deepseek_v2.py (#19122)
|
2026-02-27 13:37:29 -08:00 |
|
Baizhou Zhang
|
2472e47d73
|
Revert "Refactor graph input buffers (#18991)" (#19173)
|
2026-02-23 13:09:54 +08:00 |
|
Baizhou Zhang
|
fa80b9beba
|
[CI] Skip some subtests for tool call parser (#19172)
|
2026-02-23 12:20:12 +08:00 |
|
Baizhou Zhang
|
43f83525c0
|
Revert "[AMD] support two batch overlapping for mori ep #17953" (#19161)
|
2026-02-23 01:19:23 +08:00 |
|
Baizhou Zhang
|
1f7a813051
|
[CI]Extend timeout for test_text_models_perf.py (#19155)
|
2026-02-22 20:26:56 +08:00 |
|
Baizhou Zhang
|
c36a10aabb
|
Tiny update pull-requests permission of release-branch-cut.yml (#19121)
|
2026-02-21 20:14:31 +08:00 |
|
Baizhou Zhang
|
9a32f8ccb9
|
[CI] Move test_load_lora_from_tensor test to H100 (#18797)
|
2026-02-13 21:28:00 +08:00 |
|
Baizhou Zhang
|
947927bdb5
|
[V3.2] Change default CP token split method to --round-robin-split (#18613)
|
2026-02-11 20:14:35 +08:00 |
|
Baizhou Zhang
|
2d38b8aca0
|
Revert "[sgl-kernel] upgrade deepgemm" (#18562)
|
2026-02-11 01:17:40 +08:00 |
|
Baizhou Zhang
|
615a02dcd4
|
Revert "optimize get_topk_ragged by fusing get k and k_scale triton kernel" (#18471)
|
2026-02-09 16:37:19 +08:00 |
|
Baizhou Zhang
|
eb4cf1dfc4
|
[CI] Skip some flaky subtests for test_multi_lora_backend.py (#18408)
|
2026-02-07 19:06:53 +08:00 |
|
Baizhou Zhang
|
9fbec79906
|
Revert "[Build] Enable full kernel in aarch64 wheel" (#18385)
|
2026-02-07 09:19:07 +08:00 |
|
Baizhou Zhang
|
f2e0048d06
|
Add CI permission for Shunkangz, dongjiyingdjy, samuellees (#18377)
|
2026-02-07 01:19:02 +08:00 |
|
Baizhou Zhang
|
d279520ba5
|
[DeepGemm] Add a flag for fast warmup (#18111)
|
2026-02-04 14:12:13 +08:00 |
|
Baizhou Zhang
|
c7d53fa26a
|
Set torch url index in pyproject.toml (#16802)
|
2026-02-01 13:23:52 +08:00 |
|
Baizhou Zhang
|
1d942e4eef
|
[DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657)
|
2026-01-27 22:10:57 +08:00 |
|
Baizhou Zhang
|
832c756549
|
[Doc] Tiny update description on torch compile (#17819)
|
2026-01-27 18:59:04 +08:00 |
|
Baizhou Zhang
|
0dfe46dafb
|
[Docker] Install cudnn==9.16 for cuda 13 image to avoid check error (#17668)
|
2026-01-24 11:27:03 +08:00 |
|
Baizhou Zhang
|
283a2daeaa
|
[hotfix] Reenable all reduce fusion on sm100 (#17591)
|
2026-01-22 23:36:38 +08:00 |
|
Baizhou Zhang
|
8dae6ec03c
|
Add xyjixyjixyji to CI_Permission (#17559)
|
2026-01-21 23:27:25 -08:00 |
|
Baizhou Zhang
|
e2d33531f3
|
[Kernel] Little refactor of flashinfer allreduce norm fusion (#17474)
|
2026-01-22 13:31:57 +08:00 |
|
 Baizhou Zhangandiforgetmyname
|
fafa171529
|
[hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
|
2026-01-22 12:29:55 +08:00 |
|
Baizhou Zhang
|
3373545b9f
|
[HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518)
|
2026-01-22 12:17:02 +08:00 |
|
Baizhou Zhang
|
8251a74d5f
|
[Tiny] Backward compatibility for fp4 gemm flags (#17466)
|
2026-01-21 14:34:40 +08:00 |
|
Baizhou Zhang
|
a54d75bf2e
|
[Fix] Set fa3 as default MHA backend on Hopper (#17425)
|
2026-01-21 13:54:09 +08:00 |
|
Baizhou Zhang
|
c3f9c30f99
|
[Minor] Change lora_target_modules to "all" in CI tests (#17386)
|
2026-01-21 11:46:36 +08:00 |
|
Baizhou Zhang
|
6ea491e439
|
Overlap shared experts with deepep dispatch for single batch overlap on Blackwell (#17289)
|
2026-01-21 02:56:55 +08:00 |
|
Baizhou Zhang
|
55c616427d
|
Add flag that enables NCCL mlp sync batch for overlap scheduler (#17288)
|
2026-01-20 23:06:55 +08:00 |
|
Baizhou Zhang
|
ea879c7739
|
[Minor] Correct sglang version when installing from source (#17315)
|
2026-01-18 19:36:16 -08:00 |
|
Baizhou Zhang
|
8b9e9357fe
|
[2/n] deepseek_v2.py Refactor: Migrate MHA forward method in deepseek_v2.py (#16817)
|
2026-01-17 09:36:25 +08:00 |
|
Baizhou Zhang
|
a04675892e
|
Update flashinfer to 0.6.1 (#15551)
|
2026-01-17 00:48:30 +08:00 |
|
Baizhou Zhang
|
8b99af9af8
|
[Doc] Tiny update Cuda 13 environment instructions (#17174)
|
2026-01-16 06:12:26 +08:00 |
|
Baizhou Zhang
|
f9fc50acd6
|
[Tiny] Rename test_sparse_flash_attn.py to fix CI (#16895)
|
2026-01-11 18:18:29 +08:00 |
|
Baizhou Zhang
|
8b5d426340
|
[CI]Move fa4 e2e test to 4-gpu-b200 runner (#16889)
|
2026-01-11 15:53:38 +08:00 |
|
Baizhou Zhang
|
9fd2358cc2
|
Update Cutedsl version and pin cuda-python version (#16838)
|
2026-01-10 17:08:43 +08:00 |
|
Baizhou Zhang
|
7f393d9512
|
[Docker] Add nightly dev docker for Cuda 13 (#16862)
|
2026-01-10 14:56:53 +08:00 |
|
Baizhou Zhang
|
94fc26aad8
|
[Doc]Update note for Cuda 13 container usage (#16805)
|
2026-01-10 14:03:19 +08:00 |
|