[Kernel] Fix SM90 FP8 decode regression with benchmarked M/K/N routing (#37018)
Co-authored-by: John Doe <johndoe@example.com> Co-authored-by: BBuf <1182563586@qq.com>
This commit is contained in:
co-authored by
John Doe
BBuf
parent
00fbb6e8ac
commit
a1fe4e30a9
@@ -63,11 +63,10 @@ def _prefer_torch_rowwise_fp8(
|
|||||||
):
|
):
|
||||||
return False
|
return False
|
||||||
|
|
||||||
# Tuned on H100 over MiniMax-H3's complete dense shape set: four
|
# SM90 benchmarks show repeatable NVJet wins only for large prefill M with
|
||||||
# production sequence lengths and TP1/2/4/8 (64 shapes). This selector
|
# either a wide output or a broad down projection. Keep decode, narrow TP8
|
||||||
# chose the measured winner for every shape while retaining the AOT kernel
|
# projections, and shapes outside that measured envelope on the AOT kernel.
|
||||||
# for the smaller-K projections where NVJet loses.
|
return m >= 8192 and ((k >= 4096 and n >= 6144) or (k >= 7168 and n >= 5376))
|
||||||
return (k >= 5376 and n >= 3584) or (k >= 3584 and m >= 8192)
|
|
||||||
|
|
||||||
|
|
||||||
class Fp8ScaledMMOp(BaseFusedOp):
|
class Fp8ScaledMMOp(BaseFusedOp):
|
||||||
|
|||||||
Reference in New Issue
Block a user