Baizhou Zhang
|
c843419562
|
Remove duplicate bs=1 in nightly benchmark (#15162)
|
2025-12-15 22:09:22 -08:00 |
|
Baizhou Zhang
|
ab3ffd1c8e
|
Add nightly accuracy test for DeepSeek V3.2 (#14935)
|
2025-12-13 12:11:16 -08:00 |
|
Baizhou Zhang
|
8698867479
|
[CI]Add gb200 runner back (#15024)
|
2025-12-12 20:19:34 -08:00 |
|
Baizhou Zhang
|
7dcad45cad
|
[CI] Temp disable gb200 test (#14865)
|
2025-12-10 19:05:16 -08:00 |
|
Baizhou Zhang
|
e5201bda34
|
[CI] Unblock gb200 cutedsl test (#14469)
|
2025-12-08 17:58:25 -08:00 |
|
Baizhou Zhang
|
6799847ebf
|
[CI]Unblock and split spec v2+dp test (#14551)
|
2025-12-07 17:39:25 -08:00 |
|
Baizhou Zhang
|
673c11ba73
|
[Minor] Temporarily skipping deepep large mtp test (#14586)
|
2025-12-07 13:59:16 -08:00 |
|
Baizhou Zhang
|
9dfa01a435
|
[Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538)
|
2025-12-06 12:29:16 -08:00 |
|
Baizhou Zhang
|
bc388471d2
|
[1/n] Fix hanging during DeepGemm Warmup (#14493)
|
2025-12-06 10:44:02 -08:00 |
|
Baizhou Zhang
|
42fcf5438f
|
Revert "tiny remove deprecated endpoint call" (#14533)
|
2025-12-05 23:48:54 -08:00 |
|
Baizhou Zhang
|
80a575e4e8
|
Add YAMY1234 to CI Permission (#14475)
|
2025-12-04 21:25:49 -08:00 |
|
 Baizhou ZhangandXinyuan Tong
|
7e78825d5a
|
[Tiny]Small fixes in deepseek v32 doc (#14372)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-03 11:35:40 -08:00 |
|
Baizhou Zhang
|
4bcc5879af
|
[Doc] Fix DeepSeek V32 Doc (#14336)
|
2025-12-02 21:06:55 -08:00 |
|
Baizhou Zhang
|
922054079c
|
[Doc] Update DeepSeek-V3.2 document (#14321)
|
2025-12-02 18:19:39 -08:00 |
|
Baizhou Zhang
|
03888b9de5
|
[Minor] Upgrade cutedsl version in Dockerfile (#13968)
|
2025-12-01 17:15:26 -08:00 |
|
Baizhou Zhang
|
eb5008846a
|
[CI] Fix test_deepep_large.py (#14247)
|
2025-12-01 15:18:48 -08:00 |
|
Baizhou Zhang
|
f1115cf58d
|
Revert "[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs" (#14171)
|
2025-11-30 12:49:46 -08:00 |
|
Baizhou Zhang
|
7b03cc6482
|
[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs (#14065)
|
2025-11-30 10:11:42 -08:00 |
|
Baizhou Zhang
|
051ad83347
|
[chore] Arrange NV packages in Dockerfile (#13749)
|
2025-11-27 18:08:27 -05:00 |
|
Baizhou Zhang
|
7ab548ef64
|
[2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960)
|
2025-11-27 09:00:26 -05:00 |
|
Baizhou Zhang
|
8a9b8b8457
|
Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015)
|
2025-11-26 10:45:23 -08:00 |
|
Baizhou Zhang
|
873382a910
|
[Tiny]Upgrade README for sgl-kernel (#13945)
|
2025-11-25 16:46:30 -08:00 |
|
Baizhou Zhang
|
808b6dfdea
|
[Minor] Fix lint (#13938)
|
2025-11-25 10:57:23 -08:00 |
|
Baizhou Zhang
|
04b52fa8d6
|
[chore]Upgrade flashinfer to 0.5.3 (#13751)
|
2025-11-23 23:38:36 -08:00 |
|
 Baizhou Zhangandfy1214
|
4683e244fe
|
[1/2] Refactor DeepGeem requant for FP8 Linear on Blackwell (#13601)
Co-authored-by: fy1214
|
2025-11-23 16:07:56 -08:00 |
|
Baizhou Zhang
|
c9bd1aca32
|
[CI] Tiny refactoring sgl-kernel tests (#13813)
|
2025-11-23 12:45:17 -08:00 |
|
Baizhou Zhang
|
8bfce9b08d
|
[Tiny] Renaming environ for NVFP4 dispatch (#13756)
|
2025-11-22 00:05:20 -08:00 |
|
Baizhou Zhang
|
9f59194f29
|
[Fix] Fix DeepSeek V3 MTP on B200 (#13548)
|
2025-11-18 16:30:26 -08:00 |
|
Baizhou Zhang
|
10969ae4be
|
[chore] Disable ccache for sgl-kernel release (#13541)
|
2025-11-18 14:28:58 -08:00 |
|
Baizhou Zhang
|
85ae508e8b
|
Add bfloat16 tuned fused moe config for Dpsk-MTP layer on B200 (#13455)
|
2025-11-17 17:44:31 -08:00 |
|
Baizhou Zhang
|
d64dd3e18e
|
[Tiny]Fix 1-gpu nightly test bugs (#13389)
|
2025-11-16 15:54:17 -08:00 |
|
Baizhou Zhang
|
3ccd7fa669
|
[CI] Fix B200 CI (#13387)
|
2025-11-16 15:13:57 -08:00 |
|
Baizhou Zhang
|
10285ec204
|
[Misc]Add date to cu13 dev image tag (#13316)
|
2025-11-14 19:42:50 -08:00 |
|
Baizhou Zhang
|
8ece99a9dc
|
[CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942)
|
2025-11-11 23:16:47 -08:00 |
|
Baizhou Zhang
|
99e25805f5
|
[Fix] Fix nan error for large scale ep (#12866)
|
2025-11-11 14:44:57 -08:00 |
|
Baizhou Zhang
|
5f02b918ec
|
[Fix] Fix trtllm-mla backend when chunked prefix cache is disabled (#12361)
|
2025-11-08 15:10:25 -08:00 |
|
Baizhou Zhang
|
e039ff382c
|
[CI] Fix huggingface access for test_flash_attention_4.py (#12846)
|
2025-11-07 20:07:06 -08:00 |
|
Baizhou Zhang
|
bb6a21cd99
|
[Fix]Tiny fix in Dockerfile (#12748)
|
2025-11-05 21:04:33 -08:00 |
|
Baizhou Zhang
|
3c0a6df82d
|
[chore] Fix triton installation for cu13 image (#12742)
|
2025-11-05 20:20:18 -08:00 |
|
Baizhou Zhang
|
9a954982de
|
[chore] SGLang tag management in Dockerfile (#12734)
|
2025-11-05 19:03:26 -08:00 |
|
Baizhou Zhang
|
9ec6031d70
|
[chore]Remove dockerfile from target file of bump kernel version (#12728)
|
2025-11-05 18:01:19 -08:00 |
|
Baizhou Zhang
|
7c45b8b4bb
|
[CI] Fix qwen3-vl lora nightly ci (#12708)
|
2025-11-05 11:00:13 -08:00 |
|
Baizhou Zhang
|
d22d044734
|
Revert "Enable memory saver for hybrid model" (#12648)
|
2025-11-04 16:22:06 -08:00 |
|
Baizhou Zhang
|
42889acbd0
|
[hotfix] Fix deepep w4a8 bug (#12642)
|
2025-11-04 13:55:59 -08:00 |
|
Baizhou Zhang
|
15efbcb4e7
|
[chore] Fix update_kernel_whl_index script for multiple cuda version (#12519)
|
2025-11-03 16:34:14 -08:00 |
|
Baizhou Zhang
|
6e29446e45
|
[hotfix] Remove flashinfer-jit-cache from pyproject (#12530)
|
2025-11-02 22:11:05 -08:00 |
|
Baizhou Zhang
|
9a512cf95b
|
[CI] Move some Lora/Deterministic CI tests to nightly (#12507)
|
2025-11-01 19:54:22 -07:00 |
|
Baizhou Zhang
|
2b7bf11bd2
|
[Hotfix] Remove extra comment in sgl-kernel README (#12500)
|
2025-11-01 12:22:45 -07:00 |
|
Baizhou Zhang
|
566ade0388
|
[CI] Build aarch64 kernels for sgl-kernel test (#12480)
|
2025-11-01 11:55:42 -07:00 |
|
Baizhou Zhang
|
5f98b7fe61
|
[CI] Fix kernel installation on aarch runners (#12475)
|
2025-10-31 14:25:27 -07:00 |
|
Baizhou Zhang
|
57cc5385c0
|
[CI] Add more bins for 1-gpu CI test (#12422)
|
2025-10-31 00:05:01 -07:00 |
|
Baizhou Zhang
|
b7fdde4bb4
|
[ci] Fix ci_install_deepep (#12375)
|
2025-10-30 11:39:14 -07:00 |
|
Baizhou Zhang
|
621dfb8886
|
Import flash_mla from sgl-kernel (#12135)
|
2025-10-29 23:54:21 -07:00 |
|
Baizhou Zhang
|
685c06451f
|
[ci] Try fixing broken CIs (#12317)
|
2025-10-29 01:13:51 -07:00 |
|
Baizhou Zhang
|
587deb15a7
|
[hotfix] Fix pytest not found in CI (#12311)
|
2025-10-29 11:07:36 +08:00 |
|
Baizhou Zhang
|
75c09e1ffe
|
[Fix] Fix cu130 sgl-kernel wheel renaming (#12173)
|
2025-10-26 22:44:05 -07:00 |
|
Baizhou Zhang
|
97828878d8
|
[Doc] Small update of DeepSeek v3.2 document (#12138)
|
2025-10-25 20:34:05 -07:00 |
|
Baizhou Zhang
|
bcecf27e7c
|
[Doc] Fix format for deepseek v3.2 document (#12130)
|
2025-10-25 15:07:50 -07:00 |
|
Baizhou Zhang
|
4b0ac1d52a
|
Update sgl-kernel version to 0.3.16.post4 (#12125)
|
2025-10-25 14:33:33 -07:00 |
|
Baizhou Zhang
|
8e987fa2a3
|
Update document index for DeepSeek-v32 docs (#12101)
|
2025-10-25 13:38:58 -07:00 |
|
Baizhou Zhang
|
ce86979355
|
[Fix] Set global args in cpu test (#12105)
|
2025-10-24 21:46:17 -07:00 |
|
 
|
729b242934
|
[Doc] Add documentation for DeepSeek V3.2 (#11877)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: ybyang <ybyang7@iflytek.com>
|
2025-10-24 19:06:22 -07:00 |
|
Baizhou Zhang
|
4ef981e2b6
|
Revert "[Fix] Fix lint to pass CI" (#12042)
|
2025-10-23 19:44:58 -07:00 |
|
Baizhou Zhang
|
69ed8b67a8
|
[Fix] Fix lint to pass CI (#12037)
|
2025-10-23 19:39:38 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Baizhou Zhangandgemini-code-assist[bot]
|
983ef22cf3
|
[Doc] Update deterministic inference flag in server_arguments.md (#11978)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-10-22 14:12:15 -07:00 |
|
Baizhou Zhang
|
ef4a8097b8
|
Rename flashmla kernel options of nsa backend for better readability (#11876)
|
2025-10-21 13:14:16 -07:00 |
|
Baizhou Zhang
|
ebff4ee648
|
Update sgl-kernel and remove fast hadamard depedency (#11844)
|
2025-10-21 13:13:54 -07:00 |
|
Baizhou Zhang
|
44f0ece9fc
|
[Doc] Update documents for FA4 (#11778)
|
2025-10-19 17:40:38 -07:00 |
|
Baizhou Zhang
|
cbb5fc2edc
|
[CI] Add CI test for DeepSeek V3.2 MTP (#11835)
|
2025-10-19 17:00:25 -07:00 |
|
Baizhou Zhang
|
20b8d2306c
|
Cleaning indexer for DeepSeek V3.2 (#11682)
|
2025-10-17 13:47:21 -07:00 |
|
Baizhou Zhang
|
b0d1d717e1
|
Revert "make radix cache deterministic" (#11728)
|
2025-10-16 14:36:15 -07:00 |
|
Baizhou Zhang
|
c224a4c6cc
|
Fix log for chunked prefix cache (#11624)
|
2025-10-14 11:49:33 -07:00 |
|
Baizhou Zhang
|
9f1f699a7a
|
[CI] Add Basic Test for DeepSeek V3.2 (#11308)
|
2025-10-13 11:41:02 -07:00 |
|
Baizhou Zhang
|
8b85926a6e
|
Remove tilelang dependency in Dockerfile (#11455)
|
2025-10-10 23:17:53 -07:00 |
|
Baizhou Zhang
|
292a867ad9
|
Add flashmla and fast hadamard transform to Dockerfile (#11235)
|
2025-10-05 21:31:28 -07:00 |
|
 Baizhou ZhangandYineng Zhang
|
aa1c5cf5bd
|
Add warnings and remove dependency for deterministic inference (#10724)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-09-22 10:56:02 -07:00 |
|
Baizhou Zhang
|
f111649580
|
Replace os.environ in layernorm.py (#10684)
|
2025-09-20 00:20:33 -07:00 |
|
 
|
8ecef73f12
|
[1/2] Support deterministic inference with flashinfer attention backend (#10645)
Co-authored-by: hebiao064 <hebiaobuaa@gmail.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
|
2025-09-19 23:34:29 -07:00 |
|
 Baizhou Zhangandzhyncs
|
3fa3c22ae2
|
Fix fast decode plan for flashinfer v0.4.0rc1 and upgrade sgl-kernel 0.3.11 (#10634)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-09-19 01:25:29 -07:00 |
|
Baizhou Zhang
|
8ad700f735
|
Cleaning codes for speculative attention mode (#10149)
|
2025-09-08 17:38:06 -07:00 |
|
Baizhou Zhang
|
beac202bfd
|
Add lora_path argument to bench_multiturn.py (#10092)
|
2025-09-05 19:20:42 -07:00 |
|
Baizhou Zhang
|
7de2ce45b2
|
Disable radix cache in test_lora_update.py for better stability (#9852)
|
2025-08-31 22:28:22 -07:00 |
|
Baizhou Zhang
|
75e6a7cde1
|
Support radix cache for Lora feature (#7216)
|
2025-08-11 10:14:11 -07:00 |
|
Baizhou Zhang
|
f2d68ded6d
|
Rename lora_path to lora_id in batches (#8437)
|
2025-08-03 21:08:28 -07:00 |
|
Baizhou Zhang
|
e7e5a3050a
|
Update batch size limitation of dsv3_router_gemm kernel to 16 (#8051)
|
2025-08-01 11:53:31 +08:00 |
|
Baizhou Zhang
|
91e3d1542e
|
Update Cutlass in sgl-kernel to v4.1 (#8392)
|
2025-07-27 00:36:15 -07:00 |
|
Baizhou Zhang
|
282eb59ff3
|
Add bf16 output option for dsv3_router_gemm kernel (#7999)
|
2025-07-20 09:49:37 +08:00 |
|
Baizhou Zhang
|
8cddfa56a1
|
Clean warning logs for gate_proj loading in Lora (#8172)
|
2025-07-19 15:56:50 -07:00 |
|
Baizhou Zhang
|
88f484ce4c
|
Apply dsv3 router gemm kernel for deepseek-r1 fp4 (#7677)
|
2025-07-02 12:30:18 -07:00 |
|
Baizhou Zhang
|
7248272ccc
|
Add dsv3 router gemm kernel (#7627)
|
2025-06-29 23:31:55 -07:00 |
|
Baizhou Zhang
|
d2679f5109
|
Fix ChunkCache object has no attribute 'disable' (#7217)
|
2025-06-15 20:55:15 -07:00 |
|
Baizhou Zhang
|
25a6a9aa22
|
Fix circular import in test_prefix_chunk_info.py (#7097)
|
2025-06-11 10:57:45 -07:00 |
|
Baizhou Zhang
|
2a5f0100e0
|
Fix GGuf and add back test_gguf.py (#7067)
|
2025-06-10 21:07:20 -07:00 |
|
Baizhou Zhang
|
3b014bc13d
|
Fix test_lora.py CI (#7061)
|
2025-06-10 12:24:46 -07:00 |
|
Baizhou Zhang
|
6716b41786
|
Update default settings for blackwell (#7023)
|
2025-06-09 20:37:47 -07:00 |
|
Baizhou Zhang
|
a979daac3b
|
Fallback to lower triton version for unfound fused moe configs (#7013)
|
2025-06-09 15:41:03 -07:00 |
|
Baizhou Zhang
|
971a0dfa32
|
Extend cuda graph capture bs for B200 (#6937)
|
2025-06-08 05:13:22 -07:00 |
|
Baizhou Zhang
|
c4ffbeca19
|
Add triton fused moe kernel config for E=257 on B200 (#6939)
|
2025-06-06 23:15:01 -07:00 |
|
Baizhou Zhang
|
6a47b73024
|
Remove contiguous before Flashinfer groupwise fp8 gemm (#6804)
|
2025-06-01 18:30:54 -07:00 |
|
Baizhou Zhang
|
73def253b5
|
Fix mem_fraction_static for AMD CI (#6748)
|
2025-05-29 12:37:30 -07:00 |
|