Commit Graph
1646 Commits
Author SHA1 Message Date
Xinyuan Tong afe90a8bc9 Point Ling-3.0-flash-VL cookbook install section at the model image (#38539) 2026-09-08 10:51:33 -07:00
Xinyuan Tong 482e9f257b Add Ling-3.0-flash-VL cookbook (#38434) 2026-09-08 23:00:24 +08:00
HZY a6b542813f fix(glm-5.2-nvfp4): bound Mooncake synchronous transfer batches (#32758) 2026-09-08 22:14:33 +08:00
Wuhen Duan dfd9b5c2a4 [NPU] Enable non-greedy MTP sampling (#32495) 2026-09-08 11:18:01 +08:00
fba967ed9c [diffusion] Support diffusion decoder parallel tiling for LTX-2.5 (#36026)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-08 10:01:06 +08:00
Siju Samuel 2358916d5a [Feature][Intel XPU] Add memory saver support for Intel XPU via upstream torch_memory_saver (#29935) 2026-09-08 09:38:51 +08:00
Kedar PotdarandPo-Han Huang f4bbf12423 docs(cookbook): Qwen3.5 FP8 on B200/B300 — trtllm-gen MoE + symm mem (#38374)
Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
2026-09-08 09:12:49 +08:00
f4b75b5c36 docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-07 15:33:03 -07:00
zijiexiaandClaude Opus 5 e4008de757 Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 21:30:17 +08:00
Mick ba6d3df69a [diffusion] doc: document verified GB300 and derived GB200 H3 recipes (#38296) 2026-09-07 16:32:29 +08:00
Mick 15d2cbcc90 [diffusion] CI: validate every repeated server request (#38185) 2026-09-07 14:36:50 +08:00
39a80354aa [MUSA] Add installation guide and Dockerfile (#36709)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-09-06 20:13:53 -05:00
faceless voidandgithub-actions[bot] 30d0eb2ca9 [NPU] Adapt DFlash2 speculative decoding to Ascend NPUs (#35629)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-07 09:08:39 +08:00
Mick f3d05644db [diffusion] docs+skill: document which components to stream under layerwise offload (#35674) 2026-09-06 23:12:36 +08:00
MickandClaude Fable 5 ade1da017f [diffusion] docs: verify the DGX Spark H3 recipe (#37456)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-06 23:10:27 +08:00
Mick 938dc5621d [diffusion] refactor: reuse plain state-dict loading without per-model classes (#38127) 2026-09-06 18:39:21 +08:00
Xiaoyu Zhang d61378af77 docs(diffusion): add per-model tuning decision table to performance guide (#38148) 2026-09-06 15:26:40 +08:00
bd16c22a04 [diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-05 18:12:30 +08:00
Mick 0ea8378085 [diffusion] feat: support request-scoped skip-softmax attention (#37959) 2026-09-05 13:50:17 +08:00
Shuwen WangandClaude Opus 5 4b44a1cde2 [Refactor] Let eviction policies take construction parameters (#37795)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 13:48:30 +08:00
09f542b23a [CI] Add /rerun-test --changed to rerun every test file a PR modifies (#37618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-09-04 22:38:11 -07:00
92a4d8b5ee Clean logging under --weight-loader-prefetch-checkpoints (#33930)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-09-04 20:05:53 -07:00
zijiexiaandClaude Opus 5 3b64169f9d [Cookbook] Kimi-K3: add measured B300 1x8 Unified 8k/1k speed numbers (#37878)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:44:00 -07:00
320bdd1ee2 [Docs] Document --retraction-policy, --return-hidden-states-mode, --language-model-only (#37989)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: mottopanikeiku <fcetin@hawk.iit.edu>
Co-authored-by: alp <falpercetin@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-05 05:26:28 +08:00
alpandXinyuan Tong 3e873c2110 [Docs] Clarify OpenAI chat template defaults (#32172)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-05 05:22:05 +08:00
Faradawn Yang c8ba8996c4 Update DeepSeek-V4 Pro for B200 FP4 agentic HiCache DSpark (#38026) 2026-09-04 11:09:24 -07:00
4349538c02 [model] add cosmos3 reasoner to llm only inference (#33572)
Signed-off-by: joeltg <joel@reflection.ai>
Signed-off-by: Joe Rowell <joe@poolside.ai>
Co-authored-by: Dawid Majchrowski <dmajchrowski@nvidia.com>
Co-authored-by: Kedi Wu <kediw@nvidia.com>
Co-authored-by: Kedi Wu <31940276+kediwu0331@users.noreply.github.com>
Co-authored-by: Joel Gustafson <joelgustafson@protonmail.com>
2026-09-04 22:11:58 +08:00
Brian 2216697f90 [Docs] Refresh TPU model list and link cookbooks (#37750) 2026-09-04 17:28:57 +08:00
Polisetty V R K Jyothendra Varma b168f905c8 [Intel GPU] Align XPU toml file for rust support (#31031)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
2026-09-04 14:48:30 +08:00
Xiaoyu Zhang 54c2c99feb [Diffusion] Fuse LingBot per-token gated residual and RMSNorm modulate (#37910) 2026-09-04 14:35:13 +08:00
Yu-Yun ChangandHAI cb32dbc9e0 [AMD] [Kimi-K3] Fuse the KDA input projection into a single GEMM on ROCm (#35176)
Co-authored-by: HAI <hixiao@gmail.com>
2026-09-03 23:30:08 -07:00
72078cd7f5 [XPU] Support GPT-OSS MXFP4 checkpoints on Intel XPU (#35751)
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 14:20:28 +08:00
Thomas Wang 225129fe44 [AMD] Update v4 amd cookbook 0903 (#37829) 2026-09-03 22:11:13 -07:00
59799a3687 [Simulator] Add high-fidelity CPU-based inference simulator (#33824)
Co-authored-by: zhouhaizhu.zhz <zhouhaizhu.zhz@alibaba-inc.com>
Co-authored-by: LinSiyuan814 <linsiyuan.lsy@alibaba-inc.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-09-04 11:12:11 +08:00
Martin Hickey 3ffacf949b [Docs] [BugFix] Sync --tool-call-parser and --reasoning-parser lists with the code (#37788)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
2026-09-03 13:17:33 -07:00
Jimmy ShongandClaude Fable 5.1 2da5802bfa [Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-03 11:43:01 -07:00
triple-muandmickqian bf71035d39 [diffusion] MiniMax-H3: tiered AdaLN plan cache (pinned-host tier + per-plan LRU) (#37266)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
2026-09-03 22:09:21 +08:00
2bb25dc18b [Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-03 20:08:41 +08:00
amote-i 354ed6d66b [NPU] [DOC] Refresh supported models and features on NPU (#37799) 2026-09-03 19:53:39 +08:00
kkandwunhuang dd091f43cd [AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-03 18:44:39 +08:00
Yash Akhauri 02d9b3060a [Docs] Update K2 Horizon MoE model names (#37723) 2026-09-02 23:32:44 -07:00
Yash Akhauri 98ef7d8ae6 docs: add K2 Horizon cookbook recipes and H200 results (#37655) 2026-09-03 11:52:05 +08:00
James Liu 4229088a48 feat(kernels): generalize persistent CuTe JIT cache (#33911) 2026-09-02 19:59:04 -07:00
Alex NailsandAlison Shao 28262c20df [CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-09-02 19:46:08 -07:00
Oguz Ulgen f15748d965 [bench] Support real-traffic replay with early-stop-aware steady-state metrics in bench_one_batch_server (#37469) 2026-09-02 17:20:25 -07:00
Xinyuan Tong 3421d4375b [Docs] GLM-5.3-Flash cookbook: drop stale EP caveat, add B300/H100/B200 FP8 speed data (#37576) 2026-09-02 17:16:23 -07:00
Faradawn Yang 9c70d22721 Update GLM-5.2 NVFP4 B200/B300 for AgentX HiCache (#35368) 2026-09-02 14:39:09 -07:00
Mick f6aed6ec53 [diffusion] doc: rewrite stale diffusion compatibility matrix (#36987) 2026-09-02 23:42:54 +08:00
Kevin MiandClaude Fable 5 f586654518 [diffusion] feat: support FastH3 (4-step VSA-distilled MiniMax-H3) with a VSA-H3 attention backend (#37480)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-02 21:39:54 +08:00
Liangsheng Yin ebfd8c60e5 [CI] Install sgl-eval from PyPI through the test extra (#37504) 2026-09-02 01:45:25 -07:00