[sgl-router] Add Kimi-K3 rendering with SGLang parity (#40390)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
Kan Wu
2026-09-21 17:38:49 +08:00
committed by GitHub
co-authored by Claude Fable 5.1
parent b86a30afba
commit 2016f5e7a1
16 changed files with 473 additions and 49 deletions
+11 -1
View File
@@ -176,7 +176,7 @@ HF Jinja template from `tokenizer_config.json` or a sibling
`chat_template.jinja`, or dynamo-render's built-in DeepSeek encoder (V4 family, V3.2)
for template-less models. Cache-aware routing hashes the rendered tokens so its
prefix queries match the blocks the engine caches. Models the engine encodes in
code but dynamo-render cannot tokenize here (Inkling, Kimi K3) route via raw prompt
code but dynamo-render cannot tokenize here (Inkling) route via raw prompt
text, as does any model whose template fails to load or render.
Plain text chat requests (string `content`, no tools, no template kwargs or
@@ -211,6 +211,16 @@ Detailed content-format parity coverage follows in #39133.
The Dynamo crates are pinned exactly and `Cargo.lock` is committed; CI builds
with `--locked`, so rendered bytes cannot change without a reviewed diff.
## Kimi-K3
Kimi-K3 renders through dynamo-render's native XTML formatter with SGLang's
request semantics (reasoning controls, tools, `response_format`, continuations)
and the checkpoint's chunked tiktoken encoding. `--tokenizer-path` accepts a
local `tiktoken.model` or an HF repo id, whose `tiktoken.model`, `config.json`
and `tokenizer_config.json` are downloaded when it has no `tokenizer.json`.
An explicit null `thinking_effort` with thinking enabled is not representable
in the pinned formatter and falls back to engine-side rendering.
## HTTP/2
There is nothing to configure. The router negotiates per connection inbound and