This website requires JavaScript.
Explore
Help
Register
Sign In
minke.yu
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
2fef951fe8a86b27ab20a5fe2dfda20526e3ee69
sglang
/
benchmark
/
kernels
T
History
3 people
Caio Rocha
Caio Rocha
empyreus
c2eae96c56
MSCCL++ Integration (
#22734
)
...
Co-authored-by: Caio Rocha <
caiorocha@microsof.com
> Co-authored-by: empyreus <
rjsouza1995@gmail.com
>
2026-06-08 21:13:13 -07:00
..
all_gather
[AMD] Enable AITER custom all-gather on ROCm (
#25093
)
2026-06-02 15:57:37 -07:00
all_reduce
MSCCL++ Integration (
#22734
)
2026-06-08 21:13:13 -07:00
decoding_attention_triton
Fix benchmark import for should_use_tensor_core (
#17232
)
2026-01-16 17:48:36 -05:00
deepep
Add CLI args to conveniently support tuning more models (
#12922
)
2026-03-12 23:10:55 -07:00
deepseek
Fix Python 3.11 f-string lint error in deepgemm Blackwell benchmark (
#22108
)
2026-04-04 21:15:22 +08:00
elementwise
[Benchmark] use flashinfer bench_gpu_time instead of triton do_bench (
#20305
)
2026-03-12 04:04:30 +00:00
flashinfer_allreduce_fusion
[kernel slimming] Clean many useless sgl-kernel deprecated kernels (
#20277
)
2026-03-14 16:45:54 +08:00
fused_moe_triton
Add DeepSeekV4 fused MoE Triton autotune support (
#25569
)
2026-05-18 21:35:32 +08:00
lora_csgmv
Add offline auto-tuning for LoRA CSGMV kernel (
#20391
)
2026-04-10 13:10:43 -07:00
quantization
Use Cute-DSL NVFP4 quantization kernels (
#23745
)
2026-05-11 00:40:02 -07:00
scheduler_batch
[Benchmark] use flashinfer bench_gpu_time instead of triton do_bench (
#20305
)
2026-03-12 04:04:30 +00:00
sliding_window_attention_triton
[Benchmark] use flashinfer bench_gpu_time instead of triton do_bench (
#20305
)
2026-03-12 04:04:30 +00:00