Commit Graph
100 Commits
Author SHA1 Message Date
Baizhou Zhang 9dfa01a435 [Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538) 2025-12-06 12:29:16 -08:00
Baizhou Zhang bc388471d2 [1/n] Fix hanging during DeepGemm Warmup (#14493) 2025-12-06 10:44:02 -08:00
Baizhou Zhang 42fcf5438f Revert "tiny remove deprecated endpoint call" (#14533) 2025-12-05 23:48:54 -08:00
Baizhou Zhang 80a575e4e8 Add YAMY1234 to CI Permission (#14475) 2025-12-04 21:25:49 -08:00
Baizhou ZhangandXinyuan Tong 7e78825d5a [Tiny]Small fixes in deepseek v32 doc (#14372)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2025-12-03 11:35:40 -08:00
Baizhou Zhang 4bcc5879af [Doc] Fix DeepSeek V32 Doc (#14336) 2025-12-02 21:06:55 -08:00
Baizhou Zhang 922054079c [Doc] Update DeepSeek-V3.2 document (#14321) 2025-12-02 18:19:39 -08:00
Baizhou Zhang 03888b9de5 [Minor] Upgrade cutedsl version in Dockerfile (#13968) 2025-12-01 17:15:26 -08:00
Baizhou Zhang eb5008846a [CI] Fix test_deepep_large.py (#14247) 2025-12-01 15:18:48 -08:00
Baizhou Zhang f1115cf58d Revert "[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs" (#14171) 2025-11-30 12:49:46 -08:00
Baizhou Zhang 7b03cc6482 [Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs (#14065) 2025-11-30 10:11:42 -08:00
Baizhou Zhang 051ad83347 [chore] Arrange NV packages in Dockerfile (#13749) 2025-11-27 18:08:27 -05:00
Baizhou Zhang 7ab548ef64 [2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960) 2025-11-27 09:00:26 -05:00
Baizhou Zhang 8a9b8b8457 Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015) 2025-11-26 10:45:23 -08:00
Baizhou Zhang 873382a910 [Tiny]Upgrade README for sgl-kernel (#13945) 2025-11-25 16:46:30 -08:00
Baizhou Zhang 808b6dfdea [Minor] Fix lint (#13938) 2025-11-25 10:57:23 -08:00
Baizhou Zhang 04b52fa8d6 [chore]Upgrade flashinfer to 0.5.3 (#13751) 2025-11-23 23:38:36 -08:00
Baizhou Zhangandfy1214 4683e244fe [1/2] Refactor DeepGeem requant for FP8 Linear on Blackwell (#13601)
Co-authored-by: fy1214
2025-11-23 16:07:56 -08:00
Baizhou Zhang c9bd1aca32 [CI] Tiny refactoring sgl-kernel tests (#13813) 2025-11-23 12:45:17 -08:00
Baizhou Zhang 8bfce9b08d [Tiny] Renaming environ for NVFP4 dispatch (#13756) 2025-11-22 00:05:20 -08:00
Baizhou Zhang 9f59194f29 [Fix] Fix DeepSeek V3 MTP on B200 (#13548) 2025-11-18 16:30:26 -08:00
Baizhou Zhang 10969ae4be [chore] Disable ccache for sgl-kernel release (#13541) 2025-11-18 14:28:58 -08:00
Baizhou Zhang 85ae508e8b Add bfloat16 tuned fused moe config for Dpsk-MTP layer on B200 (#13455) 2025-11-17 17:44:31 -08:00
Baizhou Zhang d64dd3e18e [Tiny]Fix 1-gpu nightly test bugs (#13389) 2025-11-16 15:54:17 -08:00
Baizhou Zhang 3ccd7fa669 [CI] Fix B200 CI (#13387) 2025-11-16 15:13:57 -08:00
Baizhou Zhang 10285ec204 [Misc]Add date to cu13 dev image tag (#13316) 2025-11-14 19:42:50 -08:00
Baizhou Zhang 8ece99a9dc [CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942) 2025-11-11 23:16:47 -08:00
Baizhou Zhang 99e25805f5 [Fix] Fix nan error for large scale ep (#12866) 2025-11-11 14:44:57 -08:00
Baizhou Zhang 5f02b918ec [Fix] Fix trtllm-mla backend when chunked prefix cache is disabled (#12361) 2025-11-08 15:10:25 -08:00
Baizhou Zhang e039ff382c [CI] Fix huggingface access for test_flash_attention_4.py (#12846) 2025-11-07 20:07:06 -08:00
Baizhou Zhang bb6a21cd99 [Fix]Tiny fix in Dockerfile (#12748) 2025-11-05 21:04:33 -08:00
Baizhou Zhang 3c0a6df82d [chore] Fix triton installation for cu13 image (#12742) 2025-11-05 20:20:18 -08:00
Baizhou Zhang 9a954982de [chore] SGLang tag management in Dockerfile (#12734) 2025-11-05 19:03:26 -08:00
Baizhou Zhang 9ec6031d70 [chore]Remove dockerfile from target file of bump kernel version (#12728) 2025-11-05 18:01:19 -08:00
Baizhou Zhang 7c45b8b4bb [CI] Fix qwen3-vl lora nightly ci (#12708) 2025-11-05 11:00:13 -08:00
Baizhou Zhang d22d044734 Revert "Enable memory saver for hybrid model" (#12648) 2025-11-04 16:22:06 -08:00
Baizhou Zhang 42889acbd0 [hotfix] Fix deepep w4a8 bug (#12642) 2025-11-04 13:55:59 -08:00
Baizhou Zhang 15efbcb4e7 [chore] Fix update_kernel_whl_index script for multiple cuda version (#12519) 2025-11-03 16:34:14 -08:00
Baizhou Zhang 6e29446e45 [hotfix] Remove flashinfer-jit-cache from pyproject (#12530) 2025-11-02 22:11:05 -08:00
Baizhou Zhang 9a512cf95b [CI] Move some Lora/Deterministic CI tests to nightly (#12507) 2025-11-01 19:54:22 -07:00
Baizhou Zhang 2b7bf11bd2 [Hotfix] Remove extra comment in sgl-kernel README (#12500) 2025-11-01 12:22:45 -07:00
Baizhou Zhang 566ade0388 [CI] Build aarch64 kernels for sgl-kernel test (#12480) 2025-11-01 11:55:42 -07:00
Baizhou Zhang 5f98b7fe61 [CI] Fix kernel installation on aarch runners (#12475) 2025-10-31 14:25:27 -07:00
Baizhou Zhang 57cc5385c0 [CI] Add more bins for 1-gpu CI test (#12422) 2025-10-31 00:05:01 -07:00
Baizhou Zhang b7fdde4bb4 [ci] Fix ci_install_deepep (#12375) 2025-10-30 11:39:14 -07:00
Baizhou Zhang 621dfb8886 Import flash_mla from sgl-kernel (#12135) 2025-10-29 23:54:21 -07:00
Baizhou Zhang 685c06451f [ci] Try fixing broken CIs (#12317) 2025-10-29 01:13:51 -07:00
Baizhou Zhang 587deb15a7 [hotfix] Fix pytest not found in CI (#12311) 2025-10-29 11:07:36 +08:00
Baizhou Zhang 75c09e1ffe [Fix] Fix cu130 sgl-kernel wheel renaming (#12173) 2025-10-26 22:44:05 -07:00
Baizhou Zhang 97828878d8 [Doc] Small update of DeepSeek v3.2 document (#12138) 2025-10-25 20:34:05 -07:00
Baizhou Zhang bcecf27e7c [Doc] Fix format for deepseek v3.2 document (#12130) 2025-10-25 15:07:50 -07:00
Baizhou Zhang 4b0ac1d52a Update sgl-kernel version to 0.3.16.post4 (#12125) 2025-10-25 14:33:33 -07:00
Baizhou Zhang 8e987fa2a3 Update document index for DeepSeek-v32 docs (#12101) 2025-10-25 13:38:58 -07:00
Baizhou Zhang ce86979355 [Fix] Set global args in cpu test (#12105) 2025-10-24 21:46:17 -07:00
729b242934 [Doc] Add documentation for DeepSeek V3.2 (#11877)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: ybyang <ybyang7@iflytek.com>
2025-10-24 19:06:22 -07:00
Baizhou Zhang 4ef981e2b6 Revert "[Fix] Fix lint to pass CI" (#12042) 2025-10-23 19:44:58 -07:00
Baizhou Zhang 69ed8b67a8 [Fix] Fix lint to pass CI (#12037) 2025-10-23 19:39:38 -07:00
Baizhou Zhangandgemini-code-assist[bot] 983ef22cf3 [Doc] Update deterministic inference flag in server_arguments.md (#11978)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-22 14:12:15 -07:00
Baizhou Zhang ef4a8097b8 Rename flashmla kernel options of nsa backend for better readability (#11876) 2025-10-21 13:14:16 -07:00
Baizhou Zhang ebff4ee648 Update sgl-kernel and remove fast hadamard depedency (#11844) 2025-10-21 13:13:54 -07:00
Baizhou Zhang 44f0ece9fc [Doc] Update documents for FA4 (#11778) 2025-10-19 17:40:38 -07:00
Baizhou Zhang cbb5fc2edc [CI] Add CI test for DeepSeek V3.2 MTP (#11835) 2025-10-19 17:00:25 -07:00
Baizhou Zhang 20b8d2306c Cleaning indexer for DeepSeek V3.2 (#11682) 2025-10-17 13:47:21 -07:00
Baizhou Zhang b0d1d717e1 Revert "make radix cache deterministic" (#11728) 2025-10-16 14:36:15 -07:00
Baizhou Zhang c224a4c6cc Fix log for chunked prefix cache (#11624) 2025-10-14 11:49:33 -07:00
Baizhou Zhang 9f1f699a7a [CI] Add Basic Test for DeepSeek V3.2 (#11308) 2025-10-13 11:41:02 -07:00
Baizhou Zhang 8b85926a6e Remove tilelang dependency in Dockerfile (#11455) 2025-10-10 23:17:53 -07:00
Baizhou Zhang 292a867ad9 Add flashmla and fast hadamard transform to Dockerfile (#11235) 2025-10-05 21:31:28 -07:00
Baizhou ZhangandYineng Zhang aa1c5cf5bd Add warnings and remove dependency for deterministic inference (#10724)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2025-09-22 10:56:02 -07:00
Baizhou Zhang f111649580 Replace os.environ in layernorm.py (#10684) 2025-09-20 00:20:33 -07:00
8ecef73f12 [1/2] Support deterministic inference with flashinfer attention backend (#10645)
Co-authored-by: hebiao064 <hebiaobuaa@gmail.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
2025-09-19 23:34:29 -07:00
Baizhou Zhangandzhyncs 3fa3c22ae2 Fix fast decode plan for flashinfer v0.4.0rc1 and upgrade sgl-kernel 0.3.11 (#10634)
Co-authored-by: zhyncs <me@zhyncs.com>
2025-09-19 01:25:29 -07:00
Baizhou Zhang 8ad700f735 Cleaning codes for speculative attention mode (#10149) 2025-09-08 17:38:06 -07:00
Baizhou Zhang beac202bfd Add lora_path argument to bench_multiturn.py (#10092) 2025-09-05 19:20:42 -07:00
Baizhou Zhang 7de2ce45b2 Disable radix cache in test_lora_update.py for better stability (#9852) 2025-08-31 22:28:22 -07:00
Baizhou Zhang 75e6a7cde1 Support radix cache for Lora feature (#7216) 2025-08-11 10:14:11 -07:00
Baizhou Zhang f2d68ded6d Rename lora_path to lora_id in batches (#8437) 2025-08-03 21:08:28 -07:00
Baizhou Zhang e7e5a3050a Update batch size limitation of dsv3_router_gemm kernel to 16 (#8051) 2025-08-01 11:53:31 +08:00
Baizhou Zhang 91e3d1542e Update Cutlass in sgl-kernel to v4.1 (#8392) 2025-07-27 00:36:15 -07:00
Baizhou Zhang 282eb59ff3 Add bf16 output option for dsv3_router_gemm kernel (#7999) 2025-07-20 09:49:37 +08:00
Baizhou Zhang 8cddfa56a1 Clean warning logs for gate_proj loading in Lora (#8172) 2025-07-19 15:56:50 -07:00
Baizhou Zhang 88f484ce4c Apply dsv3 router gemm kernel for deepseek-r1 fp4 (#7677) 2025-07-02 12:30:18 -07:00
Baizhou Zhang 7248272ccc Add dsv3 router gemm kernel (#7627) 2025-06-29 23:31:55 -07:00
Baizhou Zhang d2679f5109 Fix ChunkCache object has no attribute 'disable' (#7217) 2025-06-15 20:55:15 -07:00
Baizhou Zhang 25a6a9aa22 Fix circular import in test_prefix_chunk_info.py (#7097) 2025-06-11 10:57:45 -07:00
Baizhou Zhang 2a5f0100e0 Fix GGuf and add back test_gguf.py (#7067) 2025-06-10 21:07:20 -07:00
Baizhou Zhang 3b014bc13d Fix test_lora.py CI (#7061) 2025-06-10 12:24:46 -07:00
Baizhou Zhang 6716b41786 Update default settings for blackwell (#7023) 2025-06-09 20:37:47 -07:00
Baizhou Zhang a979daac3b Fallback to lower triton version for unfound fused moe configs (#7013) 2025-06-09 15:41:03 -07:00
Baizhou Zhang 971a0dfa32 Extend cuda graph capture bs for B200 (#6937) 2025-06-08 05:13:22 -07:00
Baizhou Zhang c4ffbeca19 Add triton fused moe kernel config for E=257 on B200 (#6939) 2025-06-06 23:15:01 -07:00
Baizhou Zhang 6a47b73024 Remove contiguous before Flashinfer groupwise fp8 gemm (#6804) 2025-06-01 18:30:54 -07:00
Baizhou Zhang 73def253b5 Fix mem_fraction_static for AMD CI (#6748) 2025-05-29 12:37:30 -07:00
Baizhou Zhang f2bd3515fb Tune memory arguments on B200 (#6718) 2025-05-29 00:03:22 -07:00
Baizhou Zhang 791b3bfabb [Feature] Support Flashinfer fp8 blockwise GEMM kernel on Blackwell (#6479) 2025-05-28 16:03:43 -07:00
Baizhou Zhang bdaefbbfbd Add environment flag for disabling message queue broadcaster (#6403) 2025-05-26 22:32:41 -07:00
Baizhou Zhang d4c038daed [Fix]Fix capture fail bug for DeepSeek (#6275) 2025-05-21 11:11:20 -07:00
Baizhou Zhang 299fd22f9e Fix throughput threshold for amd ci test (#6414) 2025-05-19 14:17:41 -07:00
Baizhou Zhang 839fb31e5f [Fix] Improve dependencies for Blackwell image (#6334) 2025-05-16 12:38:22 -07:00
Baizhou Zhang cfca4e0ed2 adding Triton configs for DeepSeekV3 FusedMoE kernel on Blackwell (#6111) 2025-05-07 23:39:10 -07:00