 
|
f4b75b5c36
|
docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-07 15:33:03 -07:00 |
|
 zijiexiaandClaude Opus 5
|
e4008de757
|
Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-07 21:30:17 +08:00 |
|
Mick
|
ba6d3df69a
|
[diffusion] doc: document verified GB300 and derived GB200 H3 recipes (#38296)
|
2026-09-07 16:32:29 +08:00 |
|
Mick
|
15d2cbcc90
|
[diffusion] CI: validate every repeated server request (#38185)
|
2026-09-07 14:36:50 +08:00 |
|
 
|
39a80354aa
|
[MUSA] Add installation guide and Dockerfile (#36709)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-09-06 20:13:53 -05:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) faceless voidandgithub-actions[bot]
|
30d0eb2ca9
|
[NPU] Adapt DFlash2 speculative decoding to Ascend NPUs (#35629)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-09-07 09:08:39 +08:00 |
|
Mick
|
f3d05644db
|
[diffusion] docs+skill: document which components to stream under layerwise offload (#35674)
|
2026-09-06 23:12:36 +08:00 |
|
 MickandClaude Fable 5
|
ade1da017f
|
[diffusion] docs: verify the DGX Spark H3 recipe (#37456)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-06 23:10:27 +08:00 |
|
Mick
|
938dc5621d
|
[diffusion] refactor: reuse plain state-dict loading without per-model classes (#38127)
|
2026-09-06 18:39:21 +08:00 |
|
Xiaoyu Zhang
|
d61378af77
|
docs(diffusion): add per-model tuning decision table to performance guide (#38148)
|
2026-09-06 15:26:40 +08:00 |
|
 
|
bd16c22a04
|
[diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-05 18:12:30 +08:00 |
|
Mick
|
0ea8378085
|
[diffusion] feat: support request-scoped skip-softmax attention (#37959)
|
2026-09-05 13:50:17 +08:00 |
|
 Shuwen WangandClaude Opus 5
|
4b44a1cde2
|
[Refactor] Let eviction policies take construction parameters (#37795)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-05 13:48:30 +08:00 |
|
 
|
09f542b23a
|
[CI] Add /rerun-test --changed to rerun every test file a PR modifies (#37618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-04 22:38:11 -07:00 |
|
 
|
92a4d8b5ee
|
Clean logging under --weight-loader-prefetch-checkpoints (#33930)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-09-04 20:05:53 -07:00 |
|
 zijiexiaandClaude Opus 5
|
3b64169f9d
|
[Cookbook] Kimi-K3: add measured B300 1x8 Unified 8k/1k speed numbers (#37878)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-04 17:44:00 -07:00 |
|
  
|
320bdd1ee2
|
[Docs] Document --retraction-policy, --return-hidden-states-mode, --language-model-only (#37989)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: mottopanikeiku <fcetin@hawk.iit.edu>
Co-authored-by: alp <falpercetin@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-05 05:26:28 +08:00 |
|
 alpandXinyuan Tong
|
3e873c2110
|
[Docs] Clarify OpenAI chat template defaults (#32172)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-05 05:22:05 +08:00 |
|
Faradawn Yang
|
c8ba8996c4
|
Update DeepSeek-V4 Pro for B200 FP4 agentic HiCache DSpark (#38026)
|
2026-09-04 11:09:24 -07:00 |
|
   
|
4349538c02
|
[model] add cosmos3 reasoner to llm only inference (#33572)
Signed-off-by: joeltg <joel@reflection.ai>
Signed-off-by: Joe Rowell <joe@poolside.ai>
Co-authored-by: Dawid Majchrowski <dmajchrowski@nvidia.com>
Co-authored-by: Kedi Wu <kediw@nvidia.com>
Co-authored-by: Kedi Wu <31940276+kediwu0331@users.noreply.github.com>
Co-authored-by: Joel Gustafson <joelgustafson@protonmail.com>
|
2026-09-04 22:11:58 +08:00 |
|
Brian
|
2216697f90
|
[Docs] Refresh TPU model list and link cookbooks (#37750)
|
2026-09-04 17:28:57 +08:00 |
|
Polisetty V R K Jyothendra Varma
|
b168f905c8
|
[Intel GPU] Align XPU toml file for rust support (#31031)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
|
2026-09-04 14:48:30 +08:00 |
|
Xiaoyu Zhang
|
54c2c99feb
|
[Diffusion] Fuse LingBot per-token gated residual and RMSNorm modulate (#37910)
|
2026-09-04 14:35:13 +08:00 |
|
 Yu-Yun ChangandHAI
|
cb32dbc9e0
|
[AMD] [Kimi-K3] Fuse the KDA input projection into a single GEMM on ROCm (#35176)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-09-03 23:30:08 -07:00 |
|
 
|
72078cd7f5
|
[XPU] Support GPT-OSS MXFP4 checkpoints on Intel XPU (#35751)
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-04 14:20:28 +08:00 |
|
Thomas Wang
|
225129fe44
|
[AMD] Update v4 amd cookbook 0903 (#37829)
|
2026-09-03 22:11:13 -07:00 |
|
  
|
59799a3687
|
[Simulator] Add high-fidelity CPU-based inference simulator (#33824)
Co-authored-by: zhouhaizhu.zhz <zhouhaizhu.zhz@alibaba-inc.com>
Co-authored-by: LinSiyuan814 <linsiyuan.lsy@alibaba-inc.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-09-04 11:12:11 +08:00 |
|
Martin Hickey
|
3ffacf949b
|
[Docs] [BugFix] Sync --tool-call-parser and --reasoning-parser lists with the code (#37788)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
2026-09-03 13:17:33 -07:00 |
|
 Jimmy ShongandClaude Fable 5.1
|
2da5802bfa
|
[Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-03 11:43:01 -07:00 |
|
 triple-muandmickqian
|
bf71035d39
|
[diffusion] MiniMax-H3: tiered AdaLN plan cache (pinned-host tier + per-plan LRU) (#37266)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
|
2026-09-03 22:09:21 +08:00 |
|
 
|
2bb25dc18b
|
[Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-03 20:08:41 +08:00 |
|
amote-i
|
354ed6d66b
|
[NPU] [DOC] Refresh supported models and features on NPU (#37799)
|
2026-09-03 19:53:39 +08:00 |
|
 kkandwunhuang
|
dd091f43cd
|
[AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2026-09-03 18:44:39 +08:00 |
|
Yash Akhauri
|
02d9b3060a
|
[Docs] Update K2 Horizon MoE model names (#37723)
|
2026-09-02 23:32:44 -07:00 |
|
Yash Akhauri
|
98ef7d8ae6
|
docs: add K2 Horizon cookbook recipes and H200 results (#37655)
|
2026-09-03 11:52:05 +08:00 |
|
James Liu
|
4229088a48
|
feat(kernels): generalize persistent CuTe JIT cache (#33911)
|
2026-09-02 19:59:04 -07:00 |
|
 Alex NailsandAlison Shao
|
28262c20df
|
[CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-02 19:46:08 -07:00 |
|
Oguz Ulgen
|
f15748d965
|
[bench] Support real-traffic replay with early-stop-aware steady-state metrics in bench_one_batch_server (#37469)
|
2026-09-02 17:20:25 -07:00 |
|
Xinyuan Tong
|
3421d4375b
|
[Docs] GLM-5.3-Flash cookbook: drop stale EP caveat, add B300/H100/B200 FP8 speed data (#37576)
|
2026-09-02 17:16:23 -07:00 |
|
Faradawn Yang
|
9c70d22721
|
Update GLM-5.2 NVFP4 B200/B300 for AgentX HiCache (#35368)
|
2026-09-02 14:39:09 -07:00 |
|
Mick
|
f6aed6ec53
|
[diffusion] doc: rewrite stale diffusion compatibility matrix (#36987)
|
2026-09-02 23:42:54 +08:00 |
|
 Kevin MiandClaude Fable 5
|
f586654518
|
[diffusion] feat: support FastH3 (4-step VSA-distilled MiniMax-H3) with a VSA-H3 attention backend (#37480)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-02 21:39:54 +08:00 |
|
Liangsheng Yin
|
ebfd8c60e5
|
[CI] Install sgl-eval from PyPI through the test extra (#37504)
|
2026-09-02 01:45:25 -07:00 |
|
 
|
4b329482e8
|
[diffusion] feat: support cube sparse attention for minimax h3 (#34893)
Co-authored-by: zhenaozhenfu <zhenaozhenfu@minimaxi.com>
Co-authored-by: Reynor <reynor@minimaxi.com>
|
2026-09-02 15:21:33 +08:00 |
|
Mick
|
9175590aa0
|
[diffusion] refactor: admit explicit attention backends by capability (#37441)
|
2026-09-02 15:17:57 +08:00 |
|
Xiaoyu Zhang
|
1aa8299d1d
|
[Diffusion] Add cumulative extra-high quality tier (#37422)
|
2026-09-02 10:26:13 +08:00 |
|
Mick
|
dde0ecdb90
|
[diffusion] feat: support spargeattention (#37437)
|
2026-09-02 09:42:33 +08:00 |
|
zijiexia
|
6d34a4d3ce
|
[Cookbook] Verify DeepSeek-V4 Flash Vision on GB300 (#37492)
|
2026-09-01 17:11:41 -07:00 |
|
 Jimmy ShongandClaude Fable 5.1
|
ed82bea146
|
[Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-01 15:42:39 -07:00 |
|
 zijiexiaandClaude Fable 5
|
0f18d389b4
|
[Cookbook] Verify DeepSeek-V4 Flash Vision balanced and high-throughput on B200 (#37468)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-01 13:09:14 -07:00 |
|