Baizhou Zhang
|
9ca68a5904
|
Remove flaky test for test_lora_update.py (#21088)
|
2026-03-21 00:52:40 -07:00 |
|
Baizhou Zhang
|
67cad3e69e
|
Revert "Support CuteDSL mm_fp4 backend" (#21077)
|
2026-03-20 22:47:47 -07:00 |
|
Baizhou Zhang
|
5f3393c04c
|
Fix deepseek-v32-fp4 b200 ci (#21072)
|
2026-03-20 22:28:40 -07:00 |
|
Baizhou Zhang
|
c82d20d48e
|
Fix DeepSeek V32 FP4 test (#20984)
|
2026-03-20 01:04:32 -07:00 |
|
Baizhou Zhang
|
42f4b7276c
|
Revert "feat(mm)(grpc): compute M-RoPE positions for preprocessed VL inputs" (#20956)
|
2026-03-19 18:03:04 -07:00 |
|
Baizhou Zhang
|
826eb21bca
|
Fix sglang-kernel dependency on CI runners (#20715)
|
2026-03-16 15:23:01 -07:00 |
|
Baizhou Zhang
|
e96a3752a0
|
Update timeout for some ci tests (#20664)
|
2026-03-15 23:58:30 -07:00 |
|
Baizhou Zhang
|
39008955ff
|
Revert "[AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars." (#20602)
|
2026-03-14 12:12:42 -07:00 |
|
Baizhou Zhang
|
f8668d9e78
|
[Fix] Add fallback for flashinfer allreduce fusion (#20384)
|
2026-03-13 01:24:55 -07:00 |
|
Baizhou Zhang
|
7771fcd4b3
|
Update CodeOwners (#20329)
|
2026-03-11 19:11:09 -07:00 |
|
Baizhou Zhang
|
a6ae89fe3c
|
Revert "chore: bump sgl-kernel version to 0.3.21.post1" (#20229)
|
2026-03-09 20:32:19 -07:00 |
|
Baizhou Zhang
|
be63f982b7
|
[V32/GLM5] Control the threshold of applying dense attention with an environ (#20062)
|
2026-03-09 14:36:10 -07:00 |
|
Baizhou Zhang
|
61d530e8ac
|
[CI] Fix lint (#20209)
|
2026-03-09 14:09:59 -07:00 |
|
Baizhou Zhang
|
d28f35240a
|
[V32/GLM5] Change default setting of V32 nvfp4 on TP4 (#20086)
|
2026-03-07 15:13:25 -08:00 |
|
Baizhou Zhang
|
04e364d538
|
[V32] Enhance deepseek v32 related tests (#19985)
|
2026-03-05 20:12:49 -08:00 |
|
Baizhou Zhang
|
51e5dc845a
|
Revert "[Kernel Slimming] Migrate NVFP4 kernels to JIT" (#20005)
|
2026-03-05 19:40:00 -08:00 |
|
Baizhou Zhang
|
10c65df48a
|
[Bug] Fix lora tp bug on H200 (#19769)
|
2026-03-04 20:11:02 -08:00 |
|
Baizhou Zhang
|
f07d668ba1
|
Remove flashinfer version argument from cu13 docker release workflow (#19862)
|
2026-03-04 00:26:06 -08:00 |
|
Baizhou Zhang
|
52dcade4aa
|
Fix flashinfer bump workflow (#19855)
|
2026-03-04 00:09:58 -08:00 |
|
Baizhou Zhang
|
78ddf05afd
|
[Fix] Install tomli in flashinfer bumping workflow (#19841)
|
2026-03-03 23:12:22 -08:00 |
|
 Baizhou ZhangandClaude
|
c287d9b645
|
chore: add flashinfer version bump workflow (#19837)
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-03-03 22:53:34 -08:00 |
|
Baizhou Zhang
|
776709efe8
|
[3/n] deepseek_v2.py Refactor: Migrate MLA forward method in deepseek_v2.py (#19122)
|
2026-02-27 13:37:29 -08:00 |
|
Baizhou Zhang
|
2472e47d73
|
Revert "Refactor graph input buffers (#18991)" (#19173)
|
2026-02-23 13:09:54 +08:00 |
|
Baizhou Zhang
|
fa80b9beba
|
[CI] Skip some subtests for tool call parser (#19172)
|
2026-02-23 12:20:12 +08:00 |
|
Baizhou Zhang
|
43f83525c0
|
Revert "[AMD] support two batch overlapping for mori ep #17953" (#19161)
|
2026-02-23 01:19:23 +08:00 |
|
Baizhou Zhang
|
1f7a813051
|
[CI]Extend timeout for test_text_models_perf.py (#19155)
|
2026-02-22 20:26:56 +08:00 |
|
Baizhou Zhang
|
c36a10aabb
|
Tiny update pull-requests permission of release-branch-cut.yml (#19121)
|
2026-02-21 20:14:31 +08:00 |
|
Baizhou Zhang
|
9a32f8ccb9
|
[CI] Move test_load_lora_from_tensor test to H100 (#18797)
|
2026-02-13 21:28:00 +08:00 |
|
Baizhou Zhang
|
947927bdb5
|
[V3.2] Change default CP token split method to --round-robin-split (#18613)
|
2026-02-11 20:14:35 +08:00 |
|
Baizhou Zhang
|
2d38b8aca0
|
Revert "[sgl-kernel] upgrade deepgemm" (#18562)
|
2026-02-11 01:17:40 +08:00 |
|
Baizhou Zhang
|
615a02dcd4
|
Revert "optimize get_topk_ragged by fusing get k and k_scale triton kernel" (#18471)
|
2026-02-09 16:37:19 +08:00 |
|
Baizhou Zhang
|
eb4cf1dfc4
|
[CI] Skip some flaky subtests for test_multi_lora_backend.py (#18408)
|
2026-02-07 19:06:53 +08:00 |
|
Baizhou Zhang
|
9fbec79906
|
Revert "[Build] Enable full kernel in aarch64 wheel" (#18385)
|
2026-02-07 09:19:07 +08:00 |
|
Baizhou Zhang
|
f2e0048d06
|
Add CI permission for Shunkangz, dongjiyingdjy, samuellees (#18377)
|
2026-02-07 01:19:02 +08:00 |
|
Baizhou Zhang
|
d279520ba5
|
[DeepGemm] Add a flag for fast warmup (#18111)
|
2026-02-04 14:12:13 +08:00 |
|
Baizhou Zhang
|
c7d53fa26a
|
Set torch url index in pyproject.toml (#16802)
|
2026-02-01 13:23:52 +08:00 |
|
Baizhou Zhang
|
1d942e4eef
|
[DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657)
|
2026-01-27 22:10:57 +08:00 |
|
Baizhou Zhang
|
832c756549
|
[Doc] Tiny update description on torch compile (#17819)
|
2026-01-27 18:59:04 +08:00 |
|
Baizhou Zhang
|
0dfe46dafb
|
[Docker] Install cudnn==9.16 for cuda 13 image to avoid check error (#17668)
|
2026-01-24 11:27:03 +08:00 |
|
Baizhou Zhang
|
283a2daeaa
|
[hotfix] Reenable all reduce fusion on sm100 (#17591)
|
2026-01-22 23:36:38 +08:00 |
|
Baizhou Zhang
|
8dae6ec03c
|
Add xyjixyjixyji to CI_Permission (#17559)
|
2026-01-21 23:27:25 -08:00 |
|
Baizhou Zhang
|
e2d33531f3
|
[Kernel] Little refactor of flashinfer allreduce norm fusion (#17474)
|
2026-01-22 13:31:57 +08:00 |
|
 Baizhou Zhangandiforgetmyname
|
fafa171529
|
[hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
|
2026-01-22 12:29:55 +08:00 |
|
Baizhou Zhang
|
3373545b9f
|
[HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518)
|
2026-01-22 12:17:02 +08:00 |
|
Baizhou Zhang
|
8251a74d5f
|
[Tiny] Backward compatibility for fp4 gemm flags (#17466)
|
2026-01-21 14:34:40 +08:00 |
|
Baizhou Zhang
|
a54d75bf2e
|
[Fix] Set fa3 as default MHA backend on Hopper (#17425)
|
2026-01-21 13:54:09 +08:00 |
|
Baizhou Zhang
|
c3f9c30f99
|
[Minor] Change lora_target_modules to "all" in CI tests (#17386)
|
2026-01-21 11:46:36 +08:00 |
|
Baizhou Zhang
|
6ea491e439
|
Overlap shared experts with deepep dispatch for single batch overlap on Blackwell (#17289)
|
2026-01-21 02:56:55 +08:00 |
|
Baizhou Zhang
|
55c616427d
|
Add flag that enables NCCL mlp sync batch for overlap scheduler (#17288)
|
2026-01-20 23:06:55 +08:00 |
|
Baizhou Zhang
|
ea879c7739
|
[Minor] Correct sglang version when installing from source (#17315)
|
2026-01-18 19:36:16 -08:00 |
|
Baizhou Zhang
|
8b9e9357fe
|
[2/n] deepseek_v2.py Refactor: Migrate MHA forward method in deepseek_v2.py (#16817)
|
2026-01-17 09:36:25 +08:00 |
|
Baizhou Zhang
|
a04675892e
|
Update flashinfer to 0.6.1 (#15551)
|
2026-01-17 00:48:30 +08:00 |
|
Baizhou Zhang
|
8b99af9af8
|
[Doc] Tiny update Cuda 13 environment instructions (#17174)
|
2026-01-16 06:12:26 +08:00 |
|
Baizhou Zhang
|
f9fc50acd6
|
[Tiny] Rename test_sparse_flash_attn.py to fix CI (#16895)
|
2026-01-11 18:18:29 +08:00 |
|
Baizhou Zhang
|
8b5d426340
|
[CI]Move fa4 e2e test to 4-gpu-b200 runner (#16889)
|
2026-01-11 15:53:38 +08:00 |
|
Baizhou Zhang
|
9fd2358cc2
|
Update Cutedsl version and pin cuda-python version (#16838)
|
2026-01-10 17:08:43 +08:00 |
|
Baizhou Zhang
|
7f393d9512
|
[Docker] Add nightly dev docker for Cuda 13 (#16862)
|
2026-01-10 14:56:53 +08:00 |
|
Baizhou Zhang
|
94fc26aad8
|
[Doc]Update note for Cuda 13 container usage (#16805)
|
2026-01-10 14:03:19 +08:00 |
|
Baizhou Zhang
|
38dc5839dd
|
[1/n]deepseek_v2.py Refactor: attention backend handlers and forward method definition (#16306)
|
2026-01-08 09:22:31 +08:00 |
|
Baizhou Zhang
|
153c69f63d
|
[CI] Enable dpsk v31 test on nightly H200 (#16660)
|
2026-01-07 23:21:19 +08:00 |
|
Baizhou Zhang
|
7d757d6f17
|
Clean Some Environment Variables for DeepSeek V32 (#15938)
|
2026-01-07 14:00:16 +08:00 |
|
Baizhou Zhang
|
6ffe1fc02f
|
[Fix]Pin mooncake version to 0.3.7.post2 in grace blackwell (#16502)
|
2026-01-06 14:11:40 +08:00 |
|
Baizhou Zhang
|
bb23a8fe77
|
[Tiny]Remove progress bar for fp8 ue8m0 quant when unneeded (#16177)
|
2026-01-03 10:54:12 +08:00 |
|
Baizhou Zhang
|
f07e76b229
|
Multiple refactors of DeepSeek V32 and context parallel (#16305)
|
2026-01-03 02:21:22 +08:00 |
|
Baizhou Zhang
|
70a769bc56
|
Fix NPU docker release workflow (#16253)
|
2026-01-01 10:53:45 +08:00 |
|
Baizhou Zhang
|
e47afa0237
|
[DP]Fix sync bubble in adjust_num_token_non_padded_for_attn_tp (#16178)
|
2025-12-31 17:58:18 +08:00 |
|
Baizhou Zhang
|
f35b5da521
|
[CI] Append test variant name to markdown report header in nightly test (#16166)
|
2025-12-31 00:09:24 +08:00 |
|
 Baizhou ZhangandKangyan Zhou
|
98225be6e5
|
[CI] Fixing release with cut branch workflow (#16153)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
|
2025-12-30 17:18:02 +08:00 |
|
Baizhou Zhang
|
208e6a9dac
|
[Doc]Update MTP moe backends for EP document (#16013)
|
2025-12-28 19:36:23 +08:00 |
|
Baizhou Zhang
|
656f4d69a1
|
Refactor fp8 nextn layer for DeepSeek nvfp4 checkpoint (#15353)
|
2025-12-28 11:57:09 +08:00 |
|
Baizhou Zhang
|
468931b572
|
[Tiny]Move deepseek fp4 cutlass moe test to per-commit test (#15565)
|
2025-12-21 18:08:07 -08:00 |
|
Baizhou Zhang
|
2b0ddf89f5
|
[Tiny]Add warning for deepgemm on Blackwell (#15352)
|
2025-12-18 13:45:26 -08:00 |
|
Baizhou Zhang
|
8451e22758
|
[DeepSeek-V32]Update nightly performance benchmark (#15308)
|
2025-12-17 11:25:31 -08:00 |
|
Baizhou Zhang
|
d92c1f8cbd
|
Fix lora doc (#15282)
|
2025-12-16 15:53:38 -08:00 |
|
Baizhou Zhang
|
28a19e494b
|
Fix lint (#15281)
|
2025-12-16 13:36:01 -08:00 |
|
Baizhou Zhang
|
0261c4aff7
|
[misc] Upgrade cutedsl to 4.3.1 (#14857)
|
2025-12-16 12:11:56 -08:00 |
|
Baizhou Zhang
|
c843419562
|
Remove duplicate bs=1 in nightly benchmark (#15162)
|
2025-12-15 22:09:22 -08:00 |
|
Baizhou Zhang
|
ab3ffd1c8e
|
Add nightly accuracy test for DeepSeek V3.2 (#14935)
|
2025-12-13 12:11:16 -08:00 |
|
Baizhou Zhang
|
8698867479
|
[CI]Add gb200 runner back (#15024)
|
2025-12-12 20:19:34 -08:00 |
|
Baizhou Zhang
|
7dcad45cad
|
[CI] Temp disable gb200 test (#14865)
|
2025-12-10 19:05:16 -08:00 |
|
Baizhou Zhang
|
e5201bda34
|
[CI] Unblock gb200 cutedsl test (#14469)
|
2025-12-08 17:58:25 -08:00 |
|
Baizhou Zhang
|
6799847ebf
|
[CI]Unblock and split spec v2+dp test (#14551)
|
2025-12-07 17:39:25 -08:00 |
|
Baizhou Zhang
|
673c11ba73
|
[Minor] Temporarily skipping deepep large mtp test (#14586)
|
2025-12-07 13:59:16 -08:00 |
|
Baizhou Zhang
|
9dfa01a435
|
[Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538)
|
2025-12-06 12:29:16 -08:00 |
|
Baizhou Zhang
|
bc388471d2
|
[1/n] Fix hanging during DeepGemm Warmup (#14493)
|
2025-12-06 10:44:02 -08:00 |
|
Baizhou Zhang
|
42fcf5438f
|
Revert "tiny remove deprecated endpoint call" (#14533)
|
2025-12-05 23:48:54 -08:00 |
|
Baizhou Zhang
|
80a575e4e8
|
Add YAMY1234 to CI Permission (#14475)
|
2025-12-04 21:25:49 -08:00 |
|
 Baizhou ZhangandXinyuan Tong
|
7e78825d5a
|
[Tiny]Small fixes in deepseek v32 doc (#14372)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-03 11:35:40 -08:00 |
|
Baizhou Zhang
|
4bcc5879af
|
[Doc] Fix DeepSeek V32 Doc (#14336)
|
2025-12-02 21:06:55 -08:00 |
|
Baizhou Zhang
|
922054079c
|
[Doc] Update DeepSeek-V3.2 document (#14321)
|
2025-12-02 18:19:39 -08:00 |
|
Baizhou Zhang
|
03888b9de5
|
[Minor] Upgrade cutedsl version in Dockerfile (#13968)
|
2025-12-01 17:15:26 -08:00 |
|
Baizhou Zhang
|
eb5008846a
|
[CI] Fix test_deepep_large.py (#14247)
|
2025-12-01 15:18:48 -08:00 |
|
Baizhou Zhang
|
f1115cf58d
|
Revert "[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs" (#14171)
|
2025-11-30 12:49:46 -08:00 |
|
Baizhou Zhang
|
7b03cc6482
|
[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs (#14065)
|
2025-11-30 10:11:42 -08:00 |
|
Baizhou Zhang
|
051ad83347
|
[chore] Arrange NV packages in Dockerfile (#13749)
|
2025-11-27 18:08:27 -05:00 |
|
Baizhou Zhang
|
7ab548ef64
|
[2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960)
|
2025-11-27 09:00:26 -05:00 |
|
Baizhou Zhang
|
8a9b8b8457
|
Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015)
|
2025-11-26 10:45:23 -08:00 |
|
Baizhou Zhang
|
873382a910
|
[Tiny]Upgrade README for sgl-kernel (#13945)
|
2025-11-25 16:46:30 -08:00 |
|
Baizhou Zhang
|
808b6dfdea
|
[Minor] Fix lint (#13938)
|
2025-11-25 10:57:23 -08:00 |
|
Baizhou Zhang
|
04b52fa8d6
|
[chore]Upgrade flashinfer to 0.5.3 (#13751)
|
2025-11-23 23:38:36 -08:00 |
|