Haoguang Cai
|
989e51ba9c
|
[Docs] Rename Tencent cookbook page titles to "Hy4 preview" / "Hy3 preview" (#36823)
|
2026-08-28 01:31:52 -07:00 |
|
 zijiexiaandClaude Fable 5
|
2960d69622
|
[Cookbook] Hy4-Preview follow-ups: runtime-accurate recipes + released-model info (#36808)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-28 00:19:07 -07:00 |
|
 zijiexiaandClaude Fable 5
|
1948b61ad4
|
[Cookbook] Add the Hy4-Preview model page (Tencent) (#36804)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-27 23:26:13 -07:00 |
|
 zijiexiaandClaude Fable 5
|
43b5a57dbb
|
[Docs] Feature GLM-5.3-Flash in the popular-models banner (#36784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-27 21:56:10 -07:00 |
|
 zijiexiaandClaude Opus 5
|
6ccfeb59bc
|
cookbook: add a Speculative card to the GLM-5.3-Flash playground (#36740)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-27 16:03:59 -07:00 |
|
Xinyuan Tong
|
d1f14431fd
|
GLM-5.3-Flash cookbook: HiCache for LL, fusion-flag drop, EAGLE, default-cell numbers, DCP4 overlay (#36544)
|
2026-08-28 03:18:25 +08:00 |
|
 zijiexiaandClaude Opus 5
|
46a544e0a0
|
[Docs] GLM-5.3-Flash: point at compute-mamba-ratio for the KDA/KV pool split (#36719)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-27 12:16:35 -07:00 |
|
Xinyuan Tong
|
11de5e2281
|
docs(cookbook): add GB10 (DGX Spark) MXFP4 cells for Ling-3.0-flash (#36364)
|
2026-08-27 12:15:17 -07:00 |
|
 
|
97ba99067d
|
Publish per-scheduler load on a dedicated socket for load-aware routers (#34608)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-08-27 19:33:18 +08:00 |
|
Julian Huang
|
536f570e66
|
docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state (#36611)
|
2026-08-27 02:49:03 -07:00 |
|
zijiexia
|
636a6f7dba
|
cookbook: fix GLM-5.3-Flash speculative flag, size Hopper memory, record GSM8K (#36660)
|
2026-08-27 02:19:47 -07:00 |
|
andyluo7
|
0f7b5b8b2a
|
[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)
|
2026-08-27 05:19:57 +00:00 |
|
 sglang-botandsglang-bot
|
e6c1a3cd4c
|
docs: sync LMSYS SGLang blog cards (#35416)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-08-27 03:56:54 +00:00 |
|
Xinyuan Tong
|
e27a7fac77
|
GLM-5.3-Flash cookbook: default Blackwell recipes to FP8 KV + TRT-LLM DSA (#36519)
|
2026-08-27 00:51:05 +08:00 |
|
Xinyuan Tong
|
f8cc1f9525
|
GLM-5.3-Flash cookbook: FP8 KV + TRT-LLM DSA benchmark card (#36513)
|
2026-08-26 07:39:27 -07:00 |
|
 XuFuandMick
|
924aeee59c
|
[diffusion] feat: support batching for cosmos3 action generation (#36301)
Signed-off-by: FxxxxU <fu18801374388@163.com>
Signed-off-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-26 22:15:03 +08:00 |
|
Xinyuan Tong
|
dfc40e0efe
|
Add GLM-5.3-Flash cookbook (#36440)
|
2026-08-26 07:00:16 -07:00 |
|
Yuhao Yang
|
8eaffdf382
|
docs: point the Qwen3.8-Flash-Next cookbook at model support PR #36497 (#36499)
|
2026-08-26 12:53:55 +00:00 |
|
 
|
c7b5e76fa9
|
Add Qwen3.8-Flash-Next cookbook (#36496)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-26 20:36:59 +08:00 |
|
Mick
|
c3c529c28e
|
[diffusion] docs: distinguish MiniMax H3 checkpoint variants (#36412)
|
2026-08-26 15:16:18 +08:00 |
|
amote-i
|
abe3aeb142
|
[NPU] [DOC] Update CANN version in NPU installation docs (#36325)
|
2026-08-26 14:51:43 +08:00 |
|
Ziang Li
|
3c9febc68b
|
[Spec][DSA] Add --speculative-dsa-topk-backend (#36313)
|
2026-08-25 23:35:03 -07:00 |
|
 Yihao WangandClaude Opus 5
|
d7baad0116
|
[diffusion] Keep the Cosmos3 Super DiT resident on high-memory GPUs (#36375)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-25 21:10:39 -07:00 |
|
Alison Shao
|
3ec22948c1
|
[CI] /rerun-failed-ci: rerun cancelled runs and target the newest run per workflow (#34057)
|
2026-08-25 20:15:54 -07:00 |
|
 
|
34de1fb47f
|
fix(test): stabilize nightly precision regression (#34668)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-08-25 20:04:52 -07:00 |
|
Mick
|
054f485d38
|
[diffusion] docs: add MiniMax H3 checkpoint format table (#36028)
|
2026-08-26 09:18:07 +08:00 |
|
 Baizhou ZhangandRyan Stewart
|
41e7612dee
|
[Model] Support Nemotron 3.5 Lightning speculative decoding (#36186)
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
|
2026-08-25 16:43:58 -07:00 |
|
Junpan Wu
|
4b4bf3d2a5
|
[Deepseek-V4] Enable shared-experts fusion on the flashinfer_mxfp4 (trtllm-gen) MoE path (#35505)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
|
2026-08-25 14:55:25 -07:00 |
|
Xinyuan Tong
|
99c02d71b1
|
docs(cookbook): use auto parser resolution for Granite 4.2 (#36342)
|
2026-08-25 10:38:46 -07:00 |
|
Xinyuan Tong
|
b760f7fb19
|
docs(cookbook): add IBM Granite 4.2 cookbook (#36286)
|
2026-08-25 22:54:55 +08:00 |
|
 MickandClaude Fable 5
|
c3947eeada
|
[diffusion] docs: desktop-safe 24 GB recipe and the DGX Spark tier (#36169)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-25 20:25:17 +08:00 |
|
 
|
a618d4c064
|
[AMD] Add Kimi-K2.7-Code-MXFP4 to cookbook (#36246)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-08-25 17:37:32 +08:00 |
|
Mick
|
191244b3f6
|
[diffusion] feat: support loading Comfy NVFP4-AWQ text encoders (#36046)
|
2026-08-25 15:43:30 +08:00 |
|
 Артем Савкинandronnie_zheng
|
61b67316d8
|
[NPU] [Diffusion] Support MiniMax H3 on Ascend NPU's (#33569)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-08-25 09:44:14 +03:00 |
|
jacky.cheng
|
9eee990ce1
|
[AMD] cookbook: add HiCache host-DRAM KV tier for Qwen3.5 MXFP4 on MI355X (#36245)
|
2026-08-25 12:55:52 +08:00 |
|
 
|
284ed9d1d3
|
[diffusion] feat: support LongCat-Image-Edit and LongCat-Image-Edit-Turbo (#35829)
Co-authored-by: 登辉 <yangdenghui.ydh@alibaba-inc.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-25 12:42:08 +08:00 |
|
Mick
|
67853c5804
|
[diffusion] feat: dispatch fp8 companions in mixed NVFP4 checkpoints (#36066)
|
2026-08-25 11:26:37 +08:00 |
|
Mick
|
f8f9226cd2
|
[diffusion] feat: support component-scoped quantization overrides (#36035)
|
2026-08-25 09:20:47 +08:00 |
|
Mick
|
ddea7b9156
|
[diffusion] feat: support mixed Comfy NVFP4 and INT8 layers (#36061)
|
2026-08-25 09:19:47 +08:00 |
|
Xinyuan Tong
|
6e2f87d589
|
docs: mark Ling-3.0-flash DSPARK verified for all four quantizations on H200 (#36204)
|
2026-08-25 02:05:32 +08:00 |
|
Jimmy Shong
|
5030637c65
|
[docs] Split the Qwen3.8-27B NVFP4 cells by lm_head precision (#36020)
|
2026-08-25 01:50:33 +08:00 |
|
Mick
|
30f9ed09d1
|
[diffusion] feat: support loading minimax h3 gguf text encoders (#36055)
|
2026-08-24 23:00:49 +08:00 |
|
Mick
|
9b0007ed19
|
[diffusion] feat: support loading comfy nvfp4 minimax h3 checkpoints (#36044)
|
2026-08-24 22:57:02 +08:00 |
|
Mick
|
76d1401881
|
[diffusion] feat: support mixed w4a4 and int8 checkpoints (#36040)
|
2026-08-24 20:35:11 +08:00 |
|
fzyzcjy
|
3b24d8981b
|
Report per-token weight-version spans in generation meta info (#35926)
|
2026-08-24 20:18:52 +08:00 |
|
Mick
|
bfeae4e79a
|
[diffusion] feat: support loading serialized convrot w4a4 checkpoints (#36039)
|
2026-08-24 19:13:35 +08:00 |
|
Xiaoyu Zhang
|
46b92b22e2
|
[diffusion] Accelerate LingBot Video RMSNorm in quality=high (#35969)
|
2026-08-24 18:02:00 +08:00 |
|
Mick
|
adc09a1f63
|
[diffusion] feat: support loading self-describing quanto int8 encoders (#36052)
|
2026-08-24 16:37:09 +08:00 |
|
Xiaoyu Zhang
|
9866fe910b
|
[diffusion] Speed up LingBot high-quality VAE decode (#36024)
|
2026-08-24 14:13:03 +08:00 |
|
Mick
|
8df3b9eff9
|
[diffusion] feat: support loading mixed w4a8 text encoders (#36037)
|
2026-08-24 13:31:27 +08:00 |
|