 Xinyuan TongandZijie Xia
|
72ccfec594
|
docs(cookbook): verify GLM-5.2 single-node B300 (FP8 + BF16) (#28460)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-06-17 03:47:32 +00:00 |
|
chenxu214
|
71b090a8e7
|
[Ascend]GLM 5.2 deployment (#28433)
|
2026-06-17 09:55:50 +08:00 |
|
amote-i
|
74155d32bd
|
[NPU] [DOC] Update Ascend NPU docs: HDK 25.5.2, Triton 3.2.1.dev20260530 (#28311)
|
2026-06-17 09:38:31 +08:00 |
|
Thomas Wang
|
0d651e653b
|
[AMD] Update v4 amd cookbook (#28423)
|
2026-06-16 18:15:16 -07:00 |
|
Jyothirmai Kottu
|
b8b8992dde
|
docs: add Amazon SageMaker AI deployment guide (#28338)
|
2026-06-16 13:33:22 -07:00 |
|
Xinyuan Tong
|
33f205d8c5
|
docs(cookbook): fix GLM-5.2 thinking toggle kwarg + document reasoning effort (#28454)
|
2026-06-16 18:17:34 +00:00 |
|
Xinyuan Tong
|
00081a00d5
|
docs(cookbook): tune GLM-5.2 MTP to 5-1-6 and simplify launch flags (#28448)
|
2026-06-17 01:18:34 +08:00 |
|
Xinyuan Tong
|
0cb6183432
|
docs(cookbook): add GLM-5.2 deployment cookbook (#28437)
|
2026-06-16 21:49:25 +08:00 |
|
 Junlin Wuandronnie_zheng
|
2a8ea70059
|
✨ [llm][npu][quant] Add W8A8 MXFP8 quantization support for Qwen3 Dense on Ascend NPU (#22352)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-16 09:45:18 +03:00 |
|
 sglang-botandsglang-bot
|
407d3a91db
|
docs: sync LMSYS SGLang blog cards (#28364)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-06-15 18:30:13 -07:00 |
|
 zijiexiaandClaude Opus 4.8
|
7221be2cec
|
feat(cookbook): MTP --max-running-requests callout + skill sync (#28340)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-15 15:53:01 -07:00 |
|
Yi Zhong
|
3a0dd69f8e
|
Minor refactorings to the LFM2.5 cookbook for accuracy (#28072)
|
2026-06-15 14:39:10 -07:00 |
|
qinsir5522
|
81166f382d
|
Fix inaccuracies and add NPU constraints in ascend_npu_profiling.mdx. (#28283)
|
2026-06-15 21:12:18 +08:00 |
|
jianzhao-xu
|
f8d1d397b6
|
[NPU] fix ascend_docs (#28279)
|
2026-06-15 21:11:55 +08:00 |
|
McZyWu
|
8bdb007e58
|
[NPU] Docs op performance optimize (#28277)
|
2026-06-15 20:28:18 +08:00 |
|
loading66
|
f768344b1a
|
[DOCS][NPU]Supplementary Notes (#28295)
|
2026-06-15 20:06:05 +08:00 |
|
longxin9715
|
7bd1a9d163
|
Update documentation for Ascend NPU Guide (#28284)
|
2026-06-15 20:02:50 +08:00 |
|
amote-i
|
edd5eff519
|
[NPU] [DOC] fix issues in ascend_npu_support_new_models (#28296)
|
2026-06-15 19:58:33 +08:00 |
|
 Xinyuan Tongandzijiexia
|
33f99831f8
|
docs(minimax-m3): refresh B200 benchmarks (tp8, piecewise) + add GPQA (#28207)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-06-15 00:15:38 -07:00 |
|
Lianmin Zheng
|
f18d38d040
|
Revert "[AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28213)
|
2026-06-14 13:34:06 -07:00 |
|
Mohammad Miadh Angkad
|
91c63aeb4d
|
Fix stale CUDA graph benchmark and docs refs (#28041)
|
2026-06-13 21:51:42 -07:00 |
|
Jared Wen
|
5da3b37a9d
|
[CI] add Precision Regression Test on Nightly Run CI (#26902)
|
2026-06-14 12:43:57 +08:00 |
|
Chang Min Bark
|
93b402580c
|
feat: add decode clear steps env var (#28160)
|
2026-06-13 17:15:42 -07:00 |
|
 
|
3f4a338212
|
[AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs (#18182)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-06-13 16:08:19 -07:00 |
|
Xinyuan Tong
|
47fabb52ed
|
docs(minimax-m3): add high-concurrency throughput tip for H200 bf16 (#28150)
|
2026-06-13 13:09:59 -07:00 |
|
 zijiexiaandClaude Opus 4.8
|
d988d5d681
|
feat(cookbook): generic config-declared flagSelects playground axis (#28128)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-13 11:46:34 -07:00 |
|
jianzhao-xu
|
29128f31fd
|
[NPU] Best Practice Docs Splitting (#27663)
|
2026-06-13 22:16:20 +08:00 |
|
 IBRAHIM IBRAHIMandClaude Opus 4.8
|
45f8d48994
|
docs: add llm-d page under Advanced Features (#28078)
Signed-off-by: IBRAHIM IBRAHIM <66755652+Ibrahim2595@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-12 18:40:31 -07:00 |
|
 Yuwei AnandClaude Fable 5
|
6c3e429ba1
|
[Tiny] Cuda Graph Refactor Code Style Follow up (#28107)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-06-12 18:38:13 -07:00 |
|
amote-i
|
c26669e37b
|
[NPU] [DOC] Update server arguments to NPU support features page (#28083)
|
2026-06-12 17:16:52 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
95867f0932
|
[Doc] Fix some inconsistencies in the Nemotron Cookbook (#28087)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
|
2026-06-12 14:51:58 -07:00 |
|
 Mohammad Miadh Angkadandzijiexia
|
fa4273d2db
|
[Docs] Add Kimi K2.7 Code cookbook (#28064)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-06-12 11:19:44 -07:00 |
|
Xinyuan Tong
|
9f6b2339f9
|
docs(minimax-m3): warm-steady-state benchmark numbers (#28062)
|
2026-06-12 08:42:23 -07:00 |
|
Xinyuan Tong
|
dba617f2ec
|
doc: update docs for new model (#28060)
|
2026-06-12 21:14:35 +08:00 |
|
ming_wang
|
f308abc052
|
Revise the mimo-v2-flash best practice (#28016)
|
2026-06-12 17:12:33 +08:00 |
|
Yi Zhong
|
40894be3c3
|
add lfm2.5 to new cookbook. (#27409)
|
2026-06-11 20:08:19 -07:00 |
|
Liangsheng Yin
|
c0480a88be
|
[Spec] Retire Spec V1 (#27964)
|
2026-06-11 16:15:15 -07:00 |
|
Trevor Morris
|
0bac184425
|
[NVIDIA] Update Minimax-M2.5,M2.7 docs with flags for performance (#24465)
|
2026-06-11 14:58:44 -07:00 |
|
Brian Chao
|
7f57b344c9
|
[diffusion] feat: progressive resolution growing for Ideogram 4 via GPU DCT upsampling with up to 1.56× speedup (#27736)
|
2026-06-11 23:16:53 +08:00 |
|
 Chi McIsaacandMick
|
b2728bda9d
|
[diffusion] feat: use fused w8a8 kernel for Ideogram4 weight-only linear as an opt-in (#27590)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-11 23:15:27 +08:00 |
|
ming_wang
|
ce0ff154a5
|
add mimo best practice (#27665)
|
2026-06-11 10:15:31 +08:00 |
|
amote-i
|
fc1fee528c
|
[NPU] [DOC] replace <code> with backticks and remove obsolete params (#27677)
|
2026-06-11 09:33:31 +08:00 |
|
 zijiexiaandClaude Fable 5
|
43835b5ba5
|
docs: cookbook benchmark accuracy labels come from the model config (no engine default) (#27842)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-06-10 18:27:38 -07:00 |
|
Baizhou Zhang
|
3c1b0fb226
|
[1/n] [CP] Simplify prefill context parallel server args (#27312)
|
2026-06-10 14:11:38 -07:00 |
|
 zijiexiaandClaude Opus 4.8
|
99258b2f1e
|
[Docs] Restore right-hand ToC on the DeepSeek-V4 cookbook page (#27830)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-10 11:58:02 -07:00 |
|
Mohammad Miadh Angkad
|
276c98c6cf
|
[Docs] Add Kimi-K2.6 NVFP4 and update Kimi-K2.5 cookbook guidance (#27714)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-06-10 10:24:09 -07:00 |
|
 billishyahaoandHAI
|
0ae27405d0
|
[AMD] Support eplb for moriep (#22985)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-06-10 10:23:51 -07:00 |
|
Mohammad Miadh Angkad
|
91ff7baa28
|
[Docs] Add GLM-5.1 NVFP4 to cookbook (#27708)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-06-10 10:23:15 -07:00 |
|
 
|
1cf8efdd08
|
docs: Diffusion Gemma cookbook (#27824)
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-10 09:33:01 -07:00 |
|
JoyFuture
|
53ed34cb88
|
Fix MiMo-V2.5-Pro DP-attention dp size in cookbook deployment snippet (#27668)
|
2026-06-10 22:08:28 +08:00 |
|