Commit Graph
100 Commits
Author SHA1 Message Date
Baizhou Zhang 2e70e4f4f6 [CI] Little renaming of gb200 CI workflow (#22608) 2026-04-11 17:52:42 -07:00
Baizhou Zhang d14d368191 [Kernel] Set sgl_per_token_group_quant_8bit_v2 as default choice (#22467) 2026-04-11 01:59:57 -07:00
Baizhou ZhangandClaude Opus 4.6 3c46ff2ac5 fix: restore CPU flash_attn test to use sgl_kernel directly (#22573)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-10 21:39:20 -07:00
Baizhou Zhang 60acdc31f2 [Fix] Fix several bugs on DSA models (#22430) 2026-04-09 12:46:23 -07:00
Baizhou Zhang 606aa11ea8 [DSA] Enable all reduce fusion for DSA models (#22390) 2026-04-09 12:42:44 -07:00
Baizhou Zhang 5e1a9834f9 Update ci permission (#22387) 2026-04-08 15:42:06 -07:00
Baizhou ZhangandClaude Opus 4.6 4e5b8cb041 Fix get_version_tag.py to handle dot-separated post versions (#22385)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 15:18:22 -07:00
Baizhou Zhang 213af1d4f7 Add CI tests for GLM-5 (#22285) 2026-04-08 01:05:36 -07:00
Baizhou Zhang 20ee59bcfc [Misc] Remove unused cu13 docker release workflow (#22167) 2026-04-05 16:01:02 -07:00
Baizhou Zhang c5fa364b80 [Hotfix] Fix router gemm on sm103 (#22134) 2026-04-05 09:33:14 -07:00
Baizhou Zhang 106baedbfb [Doc] Update GLM-5 instructions in sglang documentation (#21716) 2026-04-05 03:13:07 -07:00
Baizhou Zhang 088203454b [Fix] Fix nightly tests (#22140) 2026-04-05 02:26:42 -07:00
Baizhou Zhang 723ed6c3c4 [CI]Temporary ban auto benchmark tool test (#22138) 2026-04-04 23:18:19 -07:00
Baizhou Zhang bf984ae65d Revert "[Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang" (#22098) 2026-04-04 02:17:19 -07:00
Baizhou Zhang ac1e437f6a Revert "[Feature] JIT activation and update skills (by codex)" (#22078) 2026-04-03 15:04:15 -07:00
Baizhou Zhang 97adf8a290 [misc] Add hint for kernel release trigger (#22036) 2026-04-03 03:31:44 -07:00
Baizhou ZhangandClaude Opus 4.6 98ac40192b [Workflow] Fix kernel release build failures for aarch64 and wheel renaming (#22018)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 03:23:03 -07:00
Baizhou Zhang 75de479680 [Misc] Update CI permission (#22014) 2026-04-02 23:37:05 -07:00
Baizhou ZhangandClaude Opus 4.6 5c082c307a [Workflow] Fix kernel release jobs skipped on push events (#22011)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 23:03:13 -07:00
Baizhou Zhang 0a709cfe02 [Workflow] Avoid triggering nightly tests in kernel bump workflow (#22010) 2026-04-02 22:40:33 -07:00
Baizhou Zhang efa7b2d5d3 Revert "[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine)" (#22002) 2026-04-02 20:42:13 -07:00
Baizhou ZhangandClaude Opus 4.6 29d8e959d7 [CI] Remove stale Ascend suite entries from test/srt/run_suite.py (#21978)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 16:47:19 -07:00
Baizhou Zhang c7d03a6215 Revert "Rollback flashmla to older version [1/2]" (#21922) 2026-04-02 00:27:02 -07:00
Baizhou Zhang fbc1f92453 [DSA] Set trtllm kernels as nsa default for Blackwell (#21914) 2026-04-02 00:22:27 -07:00
Baizhou Zhang 5e12c4e08e [DSA] Support trtllm sparse mla kernel for prefill batches (#21783) 2026-04-01 13:55:05 -07:00
Baizhou ZhangandClaude Opus 4.6 f60f2ccc10 [Fix] Fall back to triton MOE for GPT-OSS on Blackwell with driver >= 595 (#21780)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 15:52:10 -07:00
Baizhou Zhang d52757fe97 [CI]Remove msgm-en and mmlu tests which cause timeout (#21733) 2026-03-31 01:10:05 -07:00
Baizhou ZhangandClaude Opus 4.6 62a63eeff7 [Fix] Fix weight_loader property assignment for qwen3-next FP8 models (#21662)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-30 01:35:59 -07:00
Baizhou Zhangandgemini-code-assist[bot] 5b19c9a05d [Doc] Update tips for developer new-comers (#21659)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-29 22:40:36 -07:00
Baizhou Zhang edd4d54023 [Clean] Remove deprecated environs (#21536) 2026-03-28 00:35:44 -07:00
Baizhou Zhang 6ef4318ec0 [CI] Move v32 cp test to deepep running suite (#21585) 2026-03-27 22:49:06 -07:00
Baizhou Zhang ec29bbb286 Split workflow for releasing runtime docker (#21563) 2026-03-27 15:05:52 -07:00
Baizhou Zhang 4e905febd2 [CI] Relax several thresholds in flaky CIs (#21562) 2026-03-27 13:16:49 -07:00
Baizhou Zhang 0138129d3c [CI] Fix nemotron nvfp4 test estimated time (#21516) 2026-03-26 21:53:09 -07:00
Baizhou Zhang a93065679b Revert "bugfix for weight loading for qwen3-next" (#21496) 2026-03-26 16:17:18 -07:00
Baizhou Zhang dbe871efdd Rollback flashmla to older version [1/2] (#21430) 2026-03-25 17:49:54 -07:00
Baizhou Zhang e90cba715c Revert "[Bugfix] Disable ci for .md files" (#21420) 2026-03-25 12:00:20 -07:00
Baizhou Zhang f5c225eeba [CI] Fix TestQwen35WithHiCache (#21371) 2026-03-25 00:04:59 -07:00
Baizhou Zhang 2b75fed0dd Workaround of DSA performance drop on B200 + DP (#21337) 2026-03-24 22:21:07 -07:00
Baizhou Zhang 1046dbe038 [Fix] Fix trtllm fp4 moe kernel not found error (#21343) 2026-03-24 16:38:05 -07:00
Baizhou Zhang d173cfecd6 Update access of cut branch workflow and delete deprecated release workflow (#21231) 2026-03-23 15:18:10 -07:00
Baizhou Zhang ed316a26ef Fix CP in-seq-split method for DeepSeek V32 and update related tests (#21192) 2026-03-23 12:34:10 -07:00
Baizhou Zhang 9ca68a5904 Remove flaky test for test_lora_update.py (#21088) 2026-03-21 00:52:40 -07:00
Baizhou Zhang 67cad3e69e Revert "Support CuteDSL mm_fp4 backend" (#21077) 2026-03-20 22:47:47 -07:00
Baizhou Zhang 5f3393c04c Fix deepseek-v32-fp4 b200 ci (#21072) 2026-03-20 22:28:40 -07:00
Baizhou Zhang c82d20d48e Fix DeepSeek V32 FP4 test (#20984) 2026-03-20 01:04:32 -07:00
Baizhou Zhang 42f4b7276c Revert "feat(mm)(grpc): compute M-RoPE positions for preprocessed VL inputs" (#20956) 2026-03-19 18:03:04 -07:00
Baizhou Zhang 826eb21bca Fix sglang-kernel dependency on CI runners (#20715) 2026-03-16 15:23:01 -07:00
Baizhou Zhang e96a3752a0 Update timeout for some ci tests (#20664) 2026-03-15 23:58:30 -07:00
Baizhou Zhang 39008955ff Revert "[AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars." (#20602) 2026-03-14 12:12:42 -07:00
Baizhou Zhang f8668d9e78 [Fix] Add fallback for flashinfer allreduce fusion (#20384) 2026-03-13 01:24:55 -07:00
Baizhou Zhang 7771fcd4b3 Update CodeOwners (#20329) 2026-03-11 19:11:09 -07:00
Baizhou Zhang a6ae89fe3c Revert "chore: bump sgl-kernel version to 0.3.21.post1" (#20229) 2026-03-09 20:32:19 -07:00
Baizhou Zhang be63f982b7 [V32/GLM5] Control the threshold of applying dense attention with an environ (#20062) 2026-03-09 14:36:10 -07:00
Baizhou Zhang 61d530e8ac [CI] Fix lint (#20209) 2026-03-09 14:09:59 -07:00
Baizhou Zhang d28f35240a [V32/GLM5] Change default setting of V32 nvfp4 on TP4 (#20086) 2026-03-07 15:13:25 -08:00
Baizhou Zhang 04e364d538 [V32] Enhance deepseek v32 related tests (#19985) 2026-03-05 20:12:49 -08:00
Baizhou Zhang 51e5dc845a Revert "[Kernel Slimming] Migrate NVFP4 kernels to JIT" (#20005) 2026-03-05 19:40:00 -08:00
Baizhou Zhang 10c65df48a [Bug] Fix lora tp bug on H200 (#19769) 2026-03-04 20:11:02 -08:00
Baizhou Zhang f07d668ba1 Remove flashinfer version argument from cu13 docker release workflow (#19862) 2026-03-04 00:26:06 -08:00
Baizhou Zhang 52dcade4aa Fix flashinfer bump workflow (#19855) 2026-03-04 00:09:58 -08:00
Baizhou Zhang 78ddf05afd [Fix] Install tomli in flashinfer bumping workflow (#19841) 2026-03-03 23:12:22 -08:00
Baizhou ZhangandClaude c287d9b645 chore: add flashinfer version bump workflow (#19837)
Co-authored-by: Claude <noreply@anthropic.com>
2026-03-03 22:53:34 -08:00
Baizhou Zhang 776709efe8 [3/n] deepseek_v2.py Refactor: Migrate MLA forward method in deepseek_v2.py (#19122) 2026-02-27 13:37:29 -08:00
Baizhou Zhang 2472e47d73 Revert "Refactor graph input buffers (#18991)" (#19173) 2026-02-23 13:09:54 +08:00
Baizhou Zhang fa80b9beba [CI] Skip some subtests for tool call parser (#19172) 2026-02-23 12:20:12 +08:00
Baizhou Zhang 43f83525c0 Revert "[AMD] support two batch overlapping for mori ep #17953" (#19161) 2026-02-23 01:19:23 +08:00
Baizhou Zhang 1f7a813051 [CI]Extend timeout for test_text_models_perf.py (#19155) 2026-02-22 20:26:56 +08:00
Baizhou Zhang c36a10aabb Tiny update pull-requests permission of release-branch-cut.yml (#19121) 2026-02-21 20:14:31 +08:00
Baizhou Zhang 9a32f8ccb9 [CI] Move test_load_lora_from_tensor test to H100 (#18797) 2026-02-13 21:28:00 +08:00
Baizhou Zhang 947927bdb5 [V3.2] Change default CP token split method to --round-robin-split (#18613) 2026-02-11 20:14:35 +08:00
Baizhou Zhang 2d38b8aca0 Revert "[sgl-kernel] upgrade deepgemm" (#18562) 2026-02-11 01:17:40 +08:00
Baizhou Zhang 615a02dcd4 Revert "optimize get_topk_ragged by fusing get k and k_scale triton kernel" (#18471) 2026-02-09 16:37:19 +08:00
Baizhou Zhang eb4cf1dfc4 [CI] Skip some flaky subtests for test_multi_lora_backend.py (#18408) 2026-02-07 19:06:53 +08:00
Baizhou Zhang 9fbec79906 Revert "[Build] Enable full kernel in aarch64 wheel" (#18385) 2026-02-07 09:19:07 +08:00
Baizhou Zhang f2e0048d06 Add CI permission for Shunkangz, dongjiyingdjy, samuellees (#18377) 2026-02-07 01:19:02 +08:00
Baizhou Zhang d279520ba5 [DeepGemm] Add a flag for fast warmup (#18111) 2026-02-04 14:12:13 +08:00
Baizhou Zhang c7d53fa26a Set torch url index in pyproject.toml (#16802) 2026-02-01 13:23:52 +08:00
Baizhou Zhang 1d942e4eef [DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657) 2026-01-27 22:10:57 +08:00
Baizhou Zhang 832c756549 [Doc] Tiny update description on torch compile (#17819) 2026-01-27 18:59:04 +08:00
Baizhou Zhang 0dfe46dafb [Docker] Install cudnn==9.16 for cuda 13 image to avoid check error (#17668) 2026-01-24 11:27:03 +08:00
Baizhou Zhang 283a2daeaa [hotfix] Reenable all reduce fusion on sm100 (#17591) 2026-01-22 23:36:38 +08:00
Baizhou Zhang 8dae6ec03c Add xyjixyjixyji to CI_Permission (#17559) 2026-01-21 23:27:25 -08:00
Baizhou Zhang e2d33531f3 [Kernel] Little refactor of flashinfer allreduce norm fusion (#17474) 2026-01-22 13:31:57 +08:00
Baizhou Zhangandiforgetmyname fafa171529 [hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
2026-01-22 12:29:55 +08:00
Baizhou Zhang 3373545b9f [HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518) 2026-01-22 12:17:02 +08:00
Baizhou Zhang 8251a74d5f [Tiny] Backward compatibility for fp4 gemm flags (#17466) 2026-01-21 14:34:40 +08:00
Baizhou Zhang a54d75bf2e [Fix] Set fa3 as default MHA backend on Hopper (#17425) 2026-01-21 13:54:09 +08:00
Baizhou Zhang c3f9c30f99 [Minor] Change lora_target_modules to "all" in CI tests (#17386) 2026-01-21 11:46:36 +08:00
Baizhou Zhang 6ea491e439 Overlap shared experts with deepep dispatch for single batch overlap on Blackwell (#17289) 2026-01-21 02:56:55 +08:00
Baizhou Zhang 55c616427d Add flag that enables NCCL mlp sync batch for overlap scheduler (#17288) 2026-01-20 23:06:55 +08:00
Baizhou Zhang ea879c7739 [Minor] Correct sglang version when installing from source (#17315) 2026-01-18 19:36:16 -08:00
Baizhou Zhang 8b9e9357fe [2/n] deepseek_v2.py Refactor: Migrate MHA forward method in deepseek_v2.py (#16817) 2026-01-17 09:36:25 +08:00
Baizhou Zhang a04675892e Update flashinfer to 0.6.1 (#15551) 2026-01-17 00:48:30 +08:00
Baizhou Zhang 8b99af9af8 [Doc] Tiny update Cuda 13 environment instructions (#17174) 2026-01-16 06:12:26 +08:00
Baizhou Zhang f9fc50acd6 [Tiny] Rename test_sparse_flash_attn.py to fix CI (#16895) 2026-01-11 18:18:29 +08:00
Baizhou Zhang 8b5d426340 [CI]Move fa4 e2e test to 4-gpu-b200 runner (#16889) 2026-01-11 15:53:38 +08:00
Baizhou Zhang 9fd2358cc2 Update Cutedsl version and pin cuda-python version (#16838) 2026-01-10 17:08:43 +08:00
Baizhou Zhang 7f393d9512 [Docker] Add nightly dev docker for Cuda 13 (#16862) 2026-01-10 14:56:53 +08:00
Baizhou Zhang 94fc26aad8 [Doc]Update note for Cuda 13 container usage (#16805) 2026-01-10 14:03:19 +08:00