Commit Graph
138 Commits
Author SHA1 Message Date
Xiaoyu ZhangandZijie Xia ead1e490b5 [Doc] Add LongCat 2.0 FP8 cookbook (#30320)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-07-07 11:48:13 -07:00
Xinyuan Tong 6f22790943 cookbook: add Hunyuan 3 (Hy3) Day-0 page (#30201) 2026-07-06 13:30:47 +08:00
zijiexiaandClaude Opus 4.8 cba3801f52 docs: add PD disaggregation to GLM-5.2 cookbook playground (#29544)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 09:57:38 -07:00
Jimmy Shong 476c946543 [Doc] Cookbook: Laguna-XS-2.1 (DFlash low-latency + high-throughput) (#29884) 2026-07-02 20:05:33 +08:00
zijiexiaandClaude Opus 4.8 1c75243f5e docs: add Qwen3.6-27B-NVFP4 variant to cookbook (#29905)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 00:05:29 -07:00
Baizhou Zhang 677a11bfa9 [Doc] Tiny update dsv4 doc (#29827) 2026-07-01 01:47:35 -07:00
bb98629157 docs(cookbook): add AMD MI300X/MI325X/MI355X support for GLM-5.2 (#28471)
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-30 21:08:51 +00:00
Yihao WangandClaude Opus 4.8 a531d81c19 [diffusion][cache-dit] support Krea-2 + run-driven has_separate_cfg (#29688)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 08:22:05 -07:00
danielafrimiandDaniel Afrimi a2b5ce2ed1 Add stochastic rounding for FP16 Mamba SSM cache (#26929)
Signed-off-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
2026-06-29 01:47:09 -07:00
Xinyuan Tong 38d4ffcd86 [cookbook] drop redundant serve flags (GLM-5.2) + fix M3 page-size note (#28731) 2026-06-29 13:34:13 +08:00
Liangsheng Yin 909123ddb8 [misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591) 2026-06-28 18:38:28 -07:00
Mick cfd911ad6e docs: refine diffusion cookbook overview (#29507) 2026-06-28 00:10:21 +08:00
Aditya KamatandMick 1589603114 model: support baidu unlimited-ocr (#29186)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-27 23:36:19 +08:00
jonah-bermanandXinyuan Tong 2f34dbe372 Add native Exa-backed web_search support (#29342)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-27 14:41:09 +01:00
aeb4e98108 [diffusion] model: support JoyEcho multi-shot A/V generation support (#27420)
Co-authored-by: niehen6174 <niehen6174@users.noreply.github.com>
Co-authored-by: 1639206518@qq.com <niehen6174>
2026-06-26 15:46:31 +08:00
zijiexiaandClaude Opus 4.8 dd56a9f069 [Docs] Add NVFP4 quantization to GLM-5.2 cookbook (#29380)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 00:40:13 -07:00
Raiden MakotoandRaiden-Makoto 7f376644e0 [AMD] [GLM5] Mark EAGLE verified on MI300X/MI325X (gfx942) in GLM-5.1 cookbook (#29313)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-26 10:49:58 +08:00
Yi Zhong f83cbc2516 Add LFM2.5-230M to the LFM2.5 cookbook (#29321) 2026-06-26 08:27:57 +08:00
Xiaoyu Zhang 4d06d4c97f Sync Gemma4 hardware table with Blackwell recipes (#29266) 2026-06-25 22:57:33 +08:00
52c32035eb [diffusion] Add Qwen-Image ModelOpt NVFP4 support (#28928)
Co-authored-by: jingyu-ml <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 22:56:09 +08:00
Mick 890b38c211 [diffusion] doc: fix diffusion docs and cookbook drift (#29302) 2026-06-25 21:13:00 +08:00
Raiden MakotoandRaiden-Makoto 0075c8f02b [AMD] [GLM5] GLM-5.1 MXFP4 (MI355X) + enable EAGLE for gfx950 in cookbook (#29194)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-25 03:54:48 -07:00
Mohammad Miadh Angkad e4976683f4 [Docs] Fix broken links in cookbook (#29261) 2026-06-25 02:31:55 -07:00
Xiaoyu Zhang 7c9804ef21 Add MiMo V2.5 Blackwell vision FA4 recipe (#29253) 2026-06-25 13:47:32 +08:00
Xiaoyu Zhang efbe67d237 Tune Gemma4 26B-A4B B200 memory recipe (#29252) 2026-06-25 11:31:25 +08:00
Anusha Pant e97cc339e3 Add DeepSeek V4 Flash demo notebook (#28952) 2026-06-24 18:42:33 +00:00
Yihao Wang dd4caf9459 [Docs] fix SGLang-diffusion installation links (#29139) 2026-06-24 09:46:06 -07:00
amote-i 73d976e375 [NPU] [DOC] Fix TOC of Ascend NPU Docs (#29129) 2026-06-24 16:47:10 +08:00
Yihao Wang 83d32fbc2f remove lora section (#29056) 2026-06-23 16:26:51 +00:00
Yihao Wang b5e0965b07 [Diffusion][Cookbook] Add Krea-2 cookbook (#29051) 2026-06-23 08:02:08 -07:00
zijiexiaandClaude Opus 4.8 52a90c9a36 docs(minimax-m3): use published AMD ROCm images (#28777)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 07:48:16 +01:00
Thomas Wang 7e6587c94a [AMD] Update v4 cookbook to clean env vars (#28981) 2026-06-22 22:16:29 -07:00
Brayden ZhongandBrayden Zhong 34e5e38604 [Cookbook] Nemotron3-Ultra: Add mamba-backend and SSM dtype flags (#28675)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-22 09:50:22 -07:00
018d0c21dc [Docs] Add Anthropic-compatible API documentation (#28522)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 04:01:09 +00:00
Trevor Morris c0bb04b67f [NVIDIA] Support NVFP4 MoE for DeepSeek-V4 (#25820) 2026-06-21 19:35:14 -07:00
Jimmy Shong 7516f0db9f [cookbook] Laguna-M.1: add PD disaggregation section (#28737) 2026-06-19 19:49:49 -07:00
Xiaoyu Zhang 7e20e25848 [docs] Add B300 cookbook deployment options (#28697) 2026-06-18 21:52:45 -07:00
Jimmy Shong d962d18f15 docs: add --trust-remote-code to Laguna-M.1 / XS.2 cookbook configs (#28693) 2026-06-19 10:21:39 +08:00
Jimmy Shong f7632ef860 [Cookbook] Laguna-M.1: enable FP8 on Blackwell + drop provisional AIME numbers (#28664) 2026-06-18 09:23:53 -07:00
Jimmy Shong 0eded9e208 Add Laguna-M.1 cookbook (#28661) 2026-06-18 23:23:53 +08:00
Ryan Zzzandzhujunyu b55cf4382d docs: add DeepSeek-V4 compressed state dtype tip (#28613)
Co-authored-by: zhujunyu <zhujunyu.666@bytedance.com>
2026-06-17 23:22:22 -07:00
Mick 05b3fd0f44 [diffusion] chore: remove ltx2 snapshot mode (#28533) 2026-06-18 10:20:21 +08:00
Xinyuan TongandZijie Xia 72ccfec594 docs(cookbook): verify GLM-5.2 single-node B300 (FP8 + BF16) (#28460)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-17 03:47:32 +00:00
Thomas Wang 0d651e653b [AMD] Update v4 amd cookbook (#28423) 2026-06-16 18:15:16 -07:00
Xinyuan Tong 33f205d8c5 docs(cookbook): fix GLM-5.2 thinking toggle kwarg + document reasoning effort (#28454) 2026-06-16 18:17:34 +00:00
Xinyuan Tong 00081a00d5 docs(cookbook): tune GLM-5.2 MTP to 5-1-6 and simplify launch flags (#28448) 2026-06-17 01:18:34 +08:00
Xinyuan Tong 0cb6183432 docs(cookbook): add GLM-5.2 deployment cookbook (#28437) 2026-06-16 21:49:25 +08:00
Yi Zhong 3a0dd69f8e Minor refactorings to the LFM2.5 cookbook for accuracy (#28072) 2026-06-15 14:39:10 -07:00
Xinyuan Tongandzijiexia 33f99831f8 docs(minimax-m3): refresh B200 benchmarks (tp8, piecewise) + add GPQA (#28207)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-15 00:15:38 -07:00
Xinyuan Tong 47fabb52ed docs(minimax-m3): add high-concurrency throughput tip for H200 bf16 (#28150) 2026-06-13 13:09:59 -07:00