 Jia GuoandClaude Opus 4.6
|
6da3aba6a5
|
perf: optimize PCG inductor path for FP8 models (#21734)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-14 17:51:27 +08:00 |
|
 Jia GuoandClaude Opus 4.6
|
bc16130a17
|
ci: skip full rerun when sgl-kernel wheel already built (#22534)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-13 20:32:55 -07:00 |
|
 Jia GuoandClaude Opus 4.6
|
a2b5111962
|
perf: skip KV cache in FA backend for embedding mode (#21971)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-13 16:27:52 -07:00 |
|
 Jia GuoandClaude Opus 4.6
|
5cb4ea1d4d
|
perf: enable inductor combo_kernels for horizontal fusion (#21977)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-11 01:01:14 +08:00 |
|
 Jia GuoandClaude Opus 4.6
|
ec01ef9092
|
Fix torch.compile/dynamo crash with Qwen3 QK-norm in piecewise CUDA g… (#19818)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-20 18:05:09 +08:00 |
|
Jia Guo
|
87549f8f0b
|
perf(mamba): use Triton conv1d for non-contiguous input to avoid .contiguous() copy (#20469)
|
2026-03-19 19:38:46 -07:00 |
|