Commit Graph
100 Commits
Author SHA1 Message Date
Xinyuan Tong d1f14431fd GLM-5.3-Flash cookbook: HiCache for LL, fusion-flag drop, EAGLE, default-cell numbers, DCP4 overlay (#36544) 2026-08-28 03:18:25 +08:00
Xinyuan Tong 11de5e2281 docs(cookbook): add GB10 (DGX Spark) MXFP4 cells for Ling-3.0-flash (#36364) 2026-08-27 12:15:17 -07:00
20621aa14b [Model] Support Ling-3.0-flash (BailingMoeV3) (#33561)
Signed-off-by: JustinTong <justintong0323@gmail.com>
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com>
Co-authored-by: 翎悦 <vito.yy@antgroup.com>
Co-authored-by: 羽癫 <yudian.zy@antgroup.com>
Co-authored-by: tiwei.btw <tiwei.btw@antgroup.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: 文赋 <zibin.zb@antgroup.com>
Co-authored-by: JustinTong <justintong0323@gmail.com>
2026-08-26 17:27:23 -07:00
Xinyuan Tong a8d716ce8d dsa: widen the fp8 k-cache quant kernel's token_id to int64 (#30859) 2026-08-26 14:22:24 -07:00
Xinyuan Tong e27a7fac77 GLM-5.3-Flash cookbook: default Blackwell recipes to FP8 KV + TRT-LLM DSA (#36519) 2026-08-27 00:51:05 +08:00
Xinyuan Tong f8cc1f9525 GLM-5.3-Flash cookbook: FP8 KV + TRT-LLM DSA benchmark card (#36513) 2026-08-26 07:39:27 -07:00
Xinyuan Tong dfc40e0efe Add GLM-5.3-Flash cookbook (#36440) 2026-08-26 07:00:16 -07:00
34de1fb47f fix(test): stabilize nightly precision regression (#34668)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-08-25 20:04:52 -07:00
Xinyuan Tong 99c02d71b1 docs(cookbook): use auto parser resolution for Granite 4.2 (#36342) 2026-08-25 10:38:46 -07:00
Xinyuan Tong b760f7fb19 docs(cookbook): add IBM Granite 4.2 cookbook (#36286) 2026-08-25 22:54:55 +08:00
Xinyuan Tong 6e2f87d589 docs: mark Ling-3.0-flash DSPARK verified for all four quantizations on H200 (#36204) 2026-08-25 02:05:32 +08:00
Xinyuan Tong 05c584c44f docs: add DSPARK speculative decoding option to Ling-3.0-flash cookbook (#35861) 2026-08-22 02:50:55 +08:00
Xinyuan Tong 157d8ad27a Support Intern-S2-Mobius FP8 (#34908) 2026-08-19 10:58:01 -07:00
Xinyuan Tong e99ecb6eee [CI] Path-gate Rust workspace tests in lint (#34864) 2026-08-14 22:41:54 -07:00
Xinyuan Tongandhnyls2002 85cdf1178d [CI] Prune redundant CPU test overhead (#34309)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-08-13 19:51:17 -07:00
Xinyuan Tong d21eefc94f [Docs] Rename Qwen3.8-Max-DSpark to Qwen3.8-2.4T-A95B-DSpark (#34590) 2026-08-12 15:25:48 +00:00
Xinyuan Tong 687967c70d Fix tokenizer warning filtering for processors (#34500) 2026-08-11 22:26:44 -07:00
Xinyuan Tong d5d41d07ed [Docs] Add Ling-3.0-tiny INT4 recipes (#34395) 2026-08-12 03:16:03 +08:00
Xinyuan Tong 1c06c160f9 [Docs] Add Ling-3.0-flash INT4 and MXFP4 recipes (#34363) 2026-08-11 00:35:27 -07:00
Xinyuan Tong 13aeb91b6e [Fix] Update multimodal CUDA VMM helper import (#34358) 2026-08-10 22:41:21 -07:00
e54c153ba6 Add Intern-S2-Mobius cookbook (#33820)
Co-authored-by: Justin Tong <justintong0323@outlook.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-10 13:13:35 -07:00
Xinyuan Tong 77c90e7e54 Cookbook: add Ling-3.0-tiny (#34283) 2026-08-10 21:25:17 +08:00
Xinyuan Tong 0da25ee6f7 Docs: Ling-3.0-flash cookbook — serve native 256K, drop YaRN override (#33882) 2026-08-07 21:46:51 +00:00
Xinyuan Tong 31c1e5943f Facade DSA index-cache: MTP topk-reuse state + index-K storage (#28609) 2026-08-06 00:34:31 -07:00
Xinyuan TongandZijie Xia b3cdd016ba Add Ling-3.0-flash cookbook (#33556)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-05 22:53:34 +08:00
Xinyuan TongandAlex Nails a9c3b55435 [Refactor] Keep chat template validation out of ServerArgs dispatcher (#33392)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-04 14:45:47 -07:00
Xinyuan Tong 7cd79fda56 [CI] Skip absent inline suites when loading timeouts (#33410) 2026-08-03 14:21:25 -07:00
e2cf21b9e5 [Kimi K3] Add reasoning, tool-call, and OpenAI serving support (#33025)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: A-transformer <cl5743590921@gmail.com>
2026-08-01 14:57:23 -07:00
Xinyuan Tong ae84811666 [Docs] Add verified H200 and B200 DeepSeek-V4 Flash Official results (#33109) 2026-08-01 16:05:42 +08:00
Xinyuan Tong 4480e2a051 [Fix] Repair verify mask test fixture (#33087) 2026-07-31 14:48:30 -07:00
Xinyuan Tongandzijiexia 94743f934c [Docs] Add DeepSeek-V4 Flash Official (0731) recipe (#33083)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-31 19:14:53 +00:00
Xinyuan Tong 68d442945f Flush dropped reasoning at stream end when stream_reasoning=False (#32225) 2026-07-31 08:25:54 +08:00
Xinyuan TongandFAN YUCHEN ee236086db Fix invalid escape warnings in tool parsers (#28370)
Co-authored-by: FAN YUCHEN <2994114386@qq.com>
2026-07-28 23:31:47 +08:00
Xinyuan Tongandliyucheng09 fc8b328f5c [Model] Support standalone text-only Qwen3.5 checkpoints (#32401)
Co-authored-by: liyucheng09 <liyucheng09@gmail.com>
2026-07-28 14:08:39 +08:00
Xinyuan Tong 39955d5314 [MoE] Make DeepEP auto serve flashinfer_cutedsl FP4 (coerce to low_latency) + guard (#29523) 2026-07-24 06:59:18 +00:00
Xinyuan Tong afaa17a7f2 [Feature] Add --default-chat-template-kwargs server arg (#29579) 2026-07-13 11:34:39 -07:00
0663ebc783 [minimax-m3] Split 4/4: model + VL + glue + function-call + fp8 quant + generic infra (#28715)
Co-authored-by: Xinyuan Tong <xinyuan-tong@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-11 11:11:06 +08:00
Xinyuan Tong 6174f1cad8 [test] Set init-static attrs in mm_process_config mock fixtures (#30766) 2026-07-10 03:00:00 -07:00
Xinyuan Tong b76dd0be69 Fix Mistral GSM8K chat eval (#27757) 2026-07-09 21:08:48 -07:00
Xinyuan Tong 7132af28de Fix garbage output for bare-tekken Mistral checkpoints (e.g. Leanstral) (#30396) 2026-07-10 00:08:38 +05:30
Xinyuan Tong 074bb928f0 Move template manager files under parser; update CODEOWNERS (#26052) 2026-07-08 16:47:36 -07:00
Xinyuan Tong 937734d3ed ci(nightly): add force_baseline_update dispatch input for precision job (#30495) 2026-07-08 16:27:41 -07:00
Xinyuan TongandEazyReal 45019b56ce [Bugfix] Map reasoning_effort=low to Nemotron-3 Super low_effort + warn on unsupported levels (#30463)
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com>
2026-07-08 12:19:44 -07:00
Xinyuan Tong 6f22790943 cookbook: add Hunyuan 3 (Hy3) Day-0 page (#30201) 2026-07-06 13:30:47 +08:00
Xinyuan Tong 854b46be99 feat(parser): resolve special-token suffix at runtime for compatibility (#29920) 2026-07-05 00:13:46 +08:00
Xinyuan Tong 9588cacaa1 Remove transformers 5.12.1 dead-code workarounds (#29758) 2026-07-03 00:03:06 +08:00
Xinyuan Tong 9ba4b8f8ba sgl-kernel: bump sgl-attn for varlen num_splits OOM fix (#29551) 2026-07-01 21:13:53 -07:00
Xinyuan TongandAlison Shao 0c1a0be3b2 fix(precision): do not promote failed runs to the comparison baseline (#28190)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
2026-07-01 18:19:15 -07:00
Xinyuan Tong cf7c6ac234 fix(nightly-precision): pin flashinfer allreduce-fusion backend for TP-partial capture contract (#28925) 2026-07-01 17:26:19 -07:00
Xinyuan Tong 45314a9fcb [spec] Fix index_share_for_mtp_iteration being a no-op in EAGLE MTP draft (#29654) 2026-06-29 14:52:23 -07:00
Xinyuan Tong 38d4ffcd86 [cookbook] drop redundant serve flags (GLM-5.2) + fix M3 page-size note (#28731) 2026-06-29 13:34:13 +08:00
Xinyuan Tong ddc389cf09 [minimax-m3] Split 3/4: disagg K-only index-K transfer (#28714) 2026-06-28 13:54:42 +08:00
Xinyuan Tongandhzh0425 592f6c849b [minimax-m3] Split 2/4: mem-cache / HiCache / sparse KV pool (#28713)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-06-28 00:19:38 +08:00
Xinyuan Tong 30ea4c0f4b build(sgl-kernel): bump FlashMLA pin + fix cccl include for CUDA 13 (#29067) 2026-06-25 20:00:15 -07:00
Xinyuan Tong ed71fb8f95 fix(anthropic): detect-and-passthrough mid-conversation system messages (#28906) 2026-06-25 17:14:12 -07:00
Xinyuan Tong 0c6e8e9477 Expand parser auto detection coverage (#28449) 2026-06-23 12:26:37 -07:00
Xinyuan Tong de3ec2c437 [server_args] compute mem_fraction_static after dp chunked-prefill division (#28884) 2026-06-22 17:59:05 -07:00
6c212a5d6b [server_args] fix FA4 page_size auto-force for combined --attention-backend fa4 (#28825)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-22 15:35:10 -07:00
Xinyuan Tong 7c23d2255a [minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712) 2026-06-22 13:10:43 -07:00
Xinyuan Tong db12bfcdc8 [JIT] Add kpool_topk_transform JIT kernel (#28670) 2026-06-22 01:04:21 -07:00
Xinyuan TongandXinyuan Tong 441ae9a5ae [Lint] Fix black formatting of DeepSeek-R1-MXFP4 MI35x tests (#28885)
Co-authored-by: Xinyuan Tong <justintong0323@gmail.com>
2026-06-22 14:20:34 +08:00
018d0c21dc [Docs] Add Anthropic-compatible API documentation (#28522)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 04:01:09 +00:00
Xinyuan Tong 61a8b42c00 docs(minimax-m3): add MMMU-Pro accuracy to B200 benchmark card (#28668) 2026-06-18 11:40:56 -07:00
Xinyuan TongandZijie Xia 72ccfec594 docs(cookbook): verify GLM-5.2 single-node B300 (FP8 + BF16) (#28460)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-17 03:47:32 +00:00
Xinyuan Tong 33f205d8c5 docs(cookbook): fix GLM-5.2 thinking toggle kwarg + document reasoning effort (#28454) 2026-06-16 18:17:34 +00:00
Xinyuan Tong 00081a00d5 docs(cookbook): tune GLM-5.2 MTP to 5-1-6 and simplify launch flags (#28448) 2026-06-17 01:18:34 +08:00
Xinyuan Tong 0cb6183432 docs(cookbook): add GLM-5.2 deployment cookbook (#28437) 2026-06-16 21:49:25 +08:00
Xinyuan Tongandzijiexia 33f99831f8 docs(minimax-m3): refresh B200 benchmarks (tp8, piecewise) + add GPQA (#28207)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-15 00:15:38 -07:00
Xinyuan Tong 1a66059c4e [Spec] Restore index_share_for_mtp_iteration in EAGLE V2 draft worker (#28192) 2026-06-14 18:01:11 -07:00
Xinyuan Tong 000fc975c7 ci(docker): support layered overlay images in release-docker-dev (#28206) 2026-06-14 16:34:02 -07:00
Xinyuan Tong 47fabb52ed docs(minimax-m3): add high-concurrency throughput tip for H200 bf16 (#28150) 2026-06-13 13:09:59 -07:00
85712fa5b0 Fix Responses API request handling (#25881)
Co-authored-by: Kai-Hsun Chen <kaihsun@apache.org>
Co-authored-by: Kristin Cowalcijk <kristincowalcijk@gmail.com>
Co-authored-by: aerosta <63026763+aerosta@users.noreply.github.com>
Co-authored-by: glaziermag <glaziermag@users.noreply.github.com>
Co-authored-by: Blake Ledden <blake.ledden@gmail.com>
Co-authored-by: PanJason <pyyjason@gmail.com>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: kennyu <966806+kennyu@users.noreply.github.com>
2026-06-12 14:47:55 -07:00
+3 b3270264e4 Fix Anthropic Messages API compatibility (#25876)
Co-authored-by: Jairo David Campaña Rosero <jairocampana10001@gmail.com>
Co-authored-by: Karan Bansal <3264937+karanb192@users.noreply.github.com>
Co-authored-by: eason <85663565+mango766@users.noreply.github.com>
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: qingchanghan <17794466+qingchanghan@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ajay Anubolu <124525760+AjAnubolu@users.noreply.github.com>
Co-authored-by: Ravitez Dondeti <13931987+dondetir@users.noreply.github.com>
Co-authored-by: Ratish P <114130421+Ratish1@users.noreply.github.com>
Co-authored-by: Xiaoshuai Zhang <15795935+jetd1@users.noreply.github.com>
Co-authored-by: Ricardo-M-L <69202550+Ricardo-M-L@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuan.tong@radixark.ai>
2026-06-12 14:46:57 -07:00
Xinyuan Tong 9f6b2339f9 docs(minimax-m3): warm-steady-state benchmark numbers (#28062) 2026-06-12 08:42:23 -07:00
Xinyuan Tong dba617f2ec doc: update docs for new model (#28060) 2026-06-12 21:14:35 +08:00
Xinyuan Tong 66ab5c9c7a fix(gateway): make sgl-model-gateway a cargo workspace so maturin 1.14 accepts the parent README (#27997) 2026-06-11 23:37:29 -07:00
Xinyuan Tong 7425bebb6c docs(cookbook): restore Gemma 4 transformers commit pin (#27321) 2026-06-04 17:43:36 -07:00
Xinyuan Tong 45a66f4088 [Docs] Update unified Text/Vision/Audio model cookbook: install + sgl-eval accuracy (#27171) 2026-06-03 09:32:08 -07:00
Xinyuan Tong fa5c8a3101 [model] support encoder-free unified Text/Vision/Audio model (#27167) 2026-06-03 23:58:06 +08:00
Xinyuan Tong 79c844527c Upgrade xgrammar to 0.2.1 (#25676) 2026-05-29 11:40:07 +08:00
Xinyuan Tong bed20249f1 fix(tool_call): reland schema type normalization (#26433) 2026-05-28 14:31:18 +08:00
Xinyuan Tong 64c7c6851b fix(tool_call): normalize non-standard JSON Schema types in tool params (#23476) 2026-05-26 15:23:59 +08:00
Xinyuan Tong 40faf44f7a [auto-detect] match Ring-2.6/Ling XML kv tool-call format via vocab signature (#25366) 2026-05-20 23:34:52 -07:00
Xinyuan TongandXinyuan Tong 52eebc82ae [Docs] MiMo-V2.5 cookbook: B200 benchmarks + multi-layer EAGLE acceptance profile + long-context reference (#25359)
Co-authored-by: Xinyuan Tong <xinyuan.tong@radixark.ai>
2026-05-19 23:15:21 -07:00
Xinyuan Tong 0aedc5678b loader: yield filtered MTP weights lazily to avoid OOM hang on multi-layer EAGLE (#25748) 2026-05-20 12:33:54 +08:00
Xinyuan Tong aad00b0ed8 Upgrade transformers to 5.8.1 (#25451) 2026-05-19 22:20:30 +08:00
Xinyuan Tong 7cb4669a04 [Fix] DeepSeek-V3.2: build structural tag locally to encode both wrapper and invoke layers (#25233) 2026-05-15 14:32:50 -07:00
Xinyuan Tong 33f1d3915f [NEW MODEL] Add H200 validation for Ring-2.6-1T cookbook (#25370) 2026-05-15 11:47:15 -07:00
Xinyuan Tong c3daa77e9a [NEW MODEL] Add Ring-2.6-1T cookbook (#25360) 2026-05-14 23:32:22 -07:00
Xinyuan Tong 5b589ed2e7 feat(constrained): two-phase reasoning grammar + --enable-strict-thinking (#23953) 2026-05-07 14:21:51 -07:00
Xinyuan Tong af2a2ac618 fix(function_call): handle Kimi-K2.5 bare numeric tool call IDs (#23950) 2026-05-07 14:20:02 -07:00
Xinyuan Tong d8f9d32a05 feat(reasoning): auto-detect reasoning/tool-call parser from chat template (#23952) 2026-05-07 14:19:16 -07:00
Xinyuan Tong f1395af543 fix(openai): map reasoning.enabled to thinking AND enable_thinking (#23951) 2026-05-07 14:01:35 -07:00
Xinyuan Tong 1e404afec2 fix(req_pool): bump pool.size to match actual tensor row count after #24243 (#24439) 2026-05-05 16:58:26 -07:00
Xinyuan Tong 8d1b6f0c00 Add zRzRzRzRzRzRzR to CI permissions (#24432) 2026-05-05 12:37:50 -07:00
Xinyuan Tong 989a16187d [Bench] Fix bench_serving missing reasoning_content stream chunks (#23954) 2026-04-30 15:00:27 -07:00
Xinyuan Tong 1376761841 fix(moe): repair dead import in fused_moe_native after MoE refactor (#24069) 2026-04-29 11:14:52 -07:00
Xinyuan Tong 4cf109bbd1 debug followup (#24058) 2026-04-29 23:03:27 +08:00
Xinyuan Tong 1279ae0787 Bugfix (#24027) 2026-04-29 21:13:51 +08:00
Xinyuan Tong 832b4f59ed [Bench] fix MMMU answer-extraction regex dropping multi-line responses (#23864) 2026-04-29 14:48:49 +08:00