Commit Graph
43 Commits
Author SHA1 Message Date
Fan Yin 65c098592d [sgl-kernel] chore: update deepgemm version (#13402) 2025-12-19 00:20:24 -08:00
Fan Yinandybyang c72f0756d2 Fix: fix flashmla fp8 kv cache acc error (#13841)
Co-authored-by: ybyang <ybyang7@iflytek.com>
2025-11-30 13:38:19 -08:00
Fan YinandHydraQYH 412160f4c1 [sgl-kernel] fix b200 kernel ci (#13907)
Co-authored-by: HydraQYH <qyh820@outlook.com>
2025-11-30 10:15:37 -08:00
Fan YinandBaizhou Zhang 36b1bcd242 [chore] update torch version to 2.9 (#12969)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-11-25 14:47:34 -08:00
Fan Yin 2dec555d36 [10/n] decouple quantization impl from vllm dependency - fix import (#13524) 2025-11-21 01:45:51 +08:00
Fan Yin e9c0c55833 [sgl-kernel] clean up fa fetch in CMakeLists.txt (#12392) 2025-11-13 15:42:09 -08:00
Fan Yin 2966367a31 [sgl-kernel] support custom fp8 flashmla kernel (#13087) 2025-11-13 12:45:21 -08:00
Fan Yin 39cee0fed2 [sgl-kernel] upd deepgemm hash to rebased commit (#11960) 2025-10-30 07:19:00 -07:00
Fan Yin 433c622e80 [Docs] update sgl-kernel readme (#11379) 2025-10-24 19:31:14 -07:00
Fan YinandPeng Zhang 14a4d80e57 [8/n] decouple quantization impl from vllm dependency - gguf srt (#11964)
Co-authored-by: Peng Zhang <zhuangsen.zp@antgroup.com>
2025-10-23 18:12:00 -07:00
Fan YinandShangming Cai 1d097aac87 [Fix] Remove unused import from triton_kernels_moe.py (#11967)
Co-authored-by: Shangming Cai <171321666+shangmingcai@users.noreply.github.com>
2025-10-22 21:02:57 +08:00
Fan Yin 23afdfd1c2 [sgl-kernel] support flashmla libtorch (#11717) 2025-10-21 21:17:50 -07:00
Fan Yin 3289da5b41 [sgl-kernel] support hadamard (#11663) 2025-10-15 19:00:44 -07:00
Fan Yin 5464457251 [sgl-kernel] Optimize gguf test (#11667) 2025-10-15 15:45:53 -07:00
PGFLMG 8fdcd98efe [7/n] decouple quantization impl from vllm dependency - gguf kernel (#11019) 2025-10-11 14:04:57 -07:00
PGFLMG 1a599509cc chore: bump sgl-kernel v0.3.14.post1 (#11137) 2025-10-05 13:46:43 -07:00
PGFLMG 580051c5a8 chore: bump sgl-kernel v0.3.14 (#11067) 2025-09-30 02:53:24 -07:00
PGFLMG 7fe89f7cdb [sgl-kernel] fix: fix missing FetchContent_Populate for fmt (#9826) 2025-08-30 12:57:42 -07:00
aa3eba8eb4 [sgl-kernel] misc: update deepgemm version for sgl-kernel (#9340)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
2025-08-27 12:01:30 -07:00
PGFLMG a3d99d6dcd [Misc] feat: Deepgemm update for sgl-kernel (#8790) 2025-08-15 01:05:27 -07:00
PGFLMG b7cd743038 [Feat] QWen-1M context support[2/2]: Update block sparse attention backend (#5949) 2025-08-06 23:49:36 -07:00
PGFLMG ac6962ccd6 [Doc] Polish sgl-kernel readme for cu126 build error (#8704) 2025-08-02 17:03:07 +08:00
PGFLMG 83f2d9d4ed [QuickFix] fix gptq model initialize (#6429) 2025-05-19 21:17:10 -07:00
PGFLMGandzhyncs f6f96b0521 [sgl-kernel] fix: fix cu118 compile error (#6123)
Co-authored-by: zhyncs <me@zhyncs.com>
2025-05-08 14:26:51 -07:00
PGFLMGandzhyncs 08acdb5c3d [Feat] Scale up fa3 kernel to sm8x arch (#5912)
Co-authored-by: zhyncs <me@zhyncs.com>
2025-04-30 13:59:36 -07:00
PGFLMG 3ddf5b9d61 [Misc] use parallel build for cmake in sgl-kernel (#5919) 2025-04-30 08:56:46 -07:00
PGFLMGandsighingnow ee71ed8a41 [Feat] QWen-1M context support[1/2]: Update block sparse attention backend utils kernel (#5847)
Co-authored-by: sighingnow <sighingnow@gmail.com>
2025-04-28 11:03:17 -07:00
PGFLMGandzhyncs c08a717c77 [Feat] Update sgl-kernel flashinfer to latest main version (#5500)
Co-authored-by: zhyncs <me@zhyncs.com>
2025-04-17 12:43:23 -07:00
PGFLMG 4879e50c6d [Feat] Add sparse attn to sgl-kernel (#5327) 2025-04-12 11:36:36 -07:00
PGFLMG ed01b4515e [Misc] Clean sgl-kernel test (#5216) 2025-04-10 11:28:41 -07:00
yinfan98 d2e507df3c [Misc] clean up vllm in sgl-kernel test (#5189) 2025-04-09 01:22:13 -07:00
yinfan98 9798e72baa [Misc] Use pytest.mark.skipif in sgl-kernel test (#5137) 2025-04-07 21:35:14 -07:00
yinfan98 b8b6008f47 [Fix] fix fa3 build at cu118 (#5036) 2025-04-03 11:52:35 -07:00
yinfan98 c7457191a0 [Fix] revert clean m.def for cudagraph (#4944) 2025-03-31 02:08:55 -07:00
yinfan98andSleepcoo 37c66ec856 [feat] add fa3 in sgl-kernel (#4902)
Co-authored-by: Sleepcoo <Sleepcoo@gmail.com>
2025-03-30 12:57:10 -07:00
yinfan98 0d7fe866f9 [Misc] Clean m.def and add Development Tips (#4890) 2025-03-29 23:06:18 -07:00
yinfan98 8e7b31546c quick fix: add default for new kernel (#4898) 2025-03-29 12:31:59 -07:00
yinfan98 ddf8981d91 Delete test_deep_gemm.py (#4891) 2025-03-29 10:46:11 -07:00
yinfan98 05625b9792 [Docs] Update DeepGEMM at README.md (#4886) 2025-03-29 09:53:39 -07:00
yinfan98 4db29e82ec [Feat] support deepgemm for cmake (#4864) 2025-03-28 10:51:44 -07:00
yinfan98 ab7fba0ece Fix nightly ci Gsm8k & Fix flashinfer backend kvcache quant (#4147) 2025-03-06 11:50:07 -08:00
yinfan98 b4d34cd35d Fix nightly-test CI (#3826) 2025-03-02 23:14:45 -08:00
9286740eff feat: refactor sgl-kernel and use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#3130)
Co-authored-by: yinfan.1024 <yinfan.1024@bytedance.com>
Co-authored-by: yinfan98 <1106110035@qq.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2025-01-26 02:55:08 +08:00