Commit Graph
259 Commits
Author SHA1 Message Date
Lianmin Zheng f18d38d040 Revert "[AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28213) 2026-06-14 13:34:06 -07:00
Mohammad Miadh Angkad 91c63aeb4d Fix stale CUDA graph benchmark and docs refs (#28041) 2026-06-13 21:51:42 -07:00
Jared Wen 5da3b37a9d [CI] add Precision Regression Test on Nightly Run CI (#26902) 2026-06-14 12:43:57 +08:00
Chang Min Bark 93b402580c feat: add decode clear steps env var (#28160) 2026-06-13 17:15:42 -07:00
3f4a338212 [AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs (#18182)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
2026-06-13 16:08:19 -07:00
Xinyuan Tong 47fabb52ed docs(minimax-m3): add high-concurrency throughput tip for H200 bf16 (#28150) 2026-06-13 13:09:59 -07:00
zijiexiaandClaude Opus 4.8 d988d5d681 feat(cookbook): generic config-declared flagSelects playground axis (#28128)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 11:46:34 -07:00
jianzhao-xu 29128f31fd [NPU] Best Practice Docs Splitting (#27663) 2026-06-13 22:16:20 +08:00
IBRAHIM IBRAHIMandClaude Opus 4.8 45f8d48994 docs: add llm-d page under Advanced Features (#28078)
Signed-off-by: IBRAHIM IBRAHIM <66755652+Ibrahim2595@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 18:40:31 -07:00
Yuwei AnandClaude Fable 5 6c3e429ba1 [Tiny] Cuda Graph Refactor Code Style Follow up (#28107)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 18:38:13 -07:00
amote-i c26669e37b [NPU] [DOC] Update server arguments to NPU support features page (#28083) 2026-06-12 17:16:52 -07:00
Brayden ZhongandBrayden Zhong 95867f0932 [Doc] Fix some inconsistencies in the Nemotron Cookbook (#28087)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-12 14:51:58 -07:00
Mohammad Miadh Angkadandzijiexia fa4273d2db [Docs] Add Kimi K2.7 Code cookbook (#28064)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-12 11:19:44 -07:00
Xinyuan Tong 9f6b2339f9 docs(minimax-m3): warm-steady-state benchmark numbers (#28062) 2026-06-12 08:42:23 -07:00
Xinyuan Tong dba617f2ec doc: update docs for new model (#28060) 2026-06-12 21:14:35 +08:00
ming_wang f308abc052 Revise the mimo-v2-flash best practice (#28016) 2026-06-12 17:12:33 +08:00
Yi Zhong 40894be3c3 add lfm2.5 to new cookbook. (#27409) 2026-06-11 20:08:19 -07:00
Liangsheng Yin c0480a88be [Spec] Retire Spec V1 (#27964) 2026-06-11 16:15:15 -07:00
Trevor Morris 0bac184425 [NVIDIA] Update Minimax-M2.5,M2.7 docs with flags for performance (#24465) 2026-06-11 14:58:44 -07:00
Brian Chao 7f57b344c9 [diffusion] feat: progressive resolution growing for Ideogram 4 via GPU DCT upsampling with up to 1.56× speedup (#27736) 2026-06-11 23:16:53 +08:00
Chi McIsaacandMick b2728bda9d [diffusion] feat: use fused w8a8 kernel for Ideogram4 weight-only linear as an opt-in (#27590)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-11 23:15:27 +08:00
ming_wang ce0ff154a5 add mimo best practice (#27665) 2026-06-11 10:15:31 +08:00
amote-i fc1fee528c [NPU] [DOC] replace <code> with backticks and remove obsolete params (#27677) 2026-06-11 09:33:31 +08:00
zijiexiaandClaude Fable 5 43835b5ba5 docs: cookbook benchmark accuracy labels come from the model config (no engine default) (#27842)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 18:27:38 -07:00
Baizhou Zhang 3c1b0fb226 [1/n] [CP] Simplify prefill context parallel server args (#27312) 2026-06-10 14:11:38 -07:00
zijiexiaandClaude Opus 4.8 99258b2f1e [Docs] Restore right-hand ToC on the DeepSeek-V4 cookbook page (#27830)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 11:58:02 -07:00
Mohammad Miadh Angkad 276c98c6cf [Docs] Add Kimi-K2.6 NVFP4 and update Kimi-K2.5 cookbook guidance (#27714)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-06-10 10:24:09 -07:00
billishyahaoandHAI 0ae27405d0 [AMD] Support eplb for moriep (#22985)
Co-authored-by: HAI <hixiao@gmail.com>
2026-06-10 10:23:51 -07:00
Mohammad Miadh Angkad 91ff7baa28 [Docs] Add GLM-5.1 NVFP4 to cookbook (#27708)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-06-10 10:23:15 -07:00
1cf8efdd08 docs: Diffusion Gemma cookbook (#27824)
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 09:33:01 -07:00
JoyFuture 53ed34cb88 Fix MiMo-V2.5-Pro DP-attention dp size in cookbook deployment snippet (#27668) 2026-06-10 22:08:28 +08:00
Ziang Li 01f10acd06 Implement online nvfp4 quantization (#26083) 2026-06-10 00:26:51 -07:00
Mick e8a437ef26 [diffusion] doc: update docs architecture (#27767) 2026-06-10 14:18:10 +08:00
2495c02c2c [Refactor] Cuda Graph Runner/Backend Refactor (#23906)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-09 21:36:57 -07:00
zijiexia 6565b7c464 [Docs] Update MegaMoE handling and rerun benchmarks (#27726) 2026-06-09 19:00:43 -07:00
bcd9c5a903 update pytorch-xpu to 2.12 (#27133)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
2026-06-10 09:25:30 +08:00
Thomas Wang 9ab7a64ee1 [AMD] Update amd qwen3.5 cookbook (#27660) 2026-06-09 16:26:33 -07:00
Liangsheng Yin 186f1e300a [CI] Move JIT kernel tests + benchmarks to test/registered/jit; add in-package guard (#27644) 2026-06-09 12:37:39 -07:00
Mick f6d53d6d16 docs: update SANA-WM cookbook serve examples (#27626) 2026-06-09 11:44:31 +08:00
bdf47315ce docs: add cookbook for SANA-WM (#27198)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-09 10:47:32 +08:00
801fe5e0f2 docs: sync LMSYS SGLang blog cards (#27517)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-08 20:31:30 +00:00
Brian Chao b0cd533a96 [diffusion] feat: progressive resolution growing for image and video models (#27524) 2026-06-08 20:44:07 +08:00
zijiexiaandClaude Opus 4.8 d1777d1f6d Cookbook renovation (#26885)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 22:04:54 -07:00
52a5c01eba [plugin] enable OOT platforms to provide custom quant configs (#25347)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <laldevashish@gmail.com>
2026-06-07 12:48:55 +08:00
Артем Савкин 9f28512737 [NPU] Update documentation for software version upgrades (#26731) 2026-06-06 15:06:13 +03:00
EduardDurech e9dbbd19e9 [model] Apertus Tool/Function and Reasoning parser (#25100) 2026-06-06 00:04:31 -07:00
Zaili Wangandgemini-code-assist[bot] caeb449cd6 [Doc][CPU]Update Cookbook with Xeon support info (#27248)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-05 22:39:40 -07:00
Mick bf66b7b6da [diffusion] model: support Ideogram4 NVFP4 (#27379) 2026-06-06 11:14:28 +08:00
6b180959a8 [SPEC][5/N] feat: batchsize-aware support for adaptive speculative_num_steps (#24055)
Co-authored-by: 坤钧 <maoyuhan.myh@antgroup.co>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
2026-06-05 15:43:02 -07:00
zijiexiaandClaude Opus 4.8 632a3d480e docs: add Tencent Hunyuan and Poolside cards to autoregressive cookbook (#27400)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:17:13 -07:00