[DeepSeek-V4.1] Bump FlashMLA to the fork's rebase head (v4.1 kernels) (#39171)

Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
This commit is contained in:
Baizhou Zhang
2026-09-13 15:43:19 -07:00
committed by GitHub
co-authored by Chunan Zeng
parent ff228d11fe
commit 5ebb16005d
3 changed files with 91 additions and 48 deletions
@@ -58,8 +58,11 @@ class BasicDecodeCorrectnessMixin:
# Language-agnostic gibberish detector. Healthy English output is
# >90% printable ASCII; multilingual token salad / Unicode noise
# from broken weight load drops well below 50%.
# Q/A framing, as in the probes above: a base model continues a bare
# instruction with arbitrary text whose language is not pinned down,
# so the ratio would measure the continuation, not output health.
out = self._decode_generate(
"Write a single sentence about a sunny day in the park.",
"Q: Write a single sentence about a sunny day in the park.\nA:",
self.sanity_max_new_tokens_long,
)
printable = sum(1 for c in out if 32 <= ord(c) < 127 or c in "\n\t")