Commit Graph
297 Commits
Author SHA1 Message Date
018d0c21dc [Docs] Add Anthropic-compatible API documentation (#28522)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 04:01:09 +00:00
be774d0acd [docs][cookbook] Laguna-M.1 playground: add HiCache; refresh EP / DP-Attention notes (#28774)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-22 03:58:42 +00:00
Trevor Morris c0bb04b67f [NVIDIA] Support NVFP4 MoE for DeepSeek-V4 (#25820) 2026-06-21 19:35:14 -07:00
amote-i 5deca2d39f [DOC] [NPU] Update features on Ascend NPU (#28643) 2026-06-22 09:50:51 +08:00
cctry 6d4ca9bc54 Cap SWA pool sizing with chunk cache (#28755) 2026-06-21 01:06:59 -07:00
Jimmy Shong 7516f0db9f [cookbook] Laguna-M.1: add PD disaggregation section (#28737) 2026-06-19 19:49:49 -07:00
Oguz Ulgen 3af991fb3e [AMD] Make breakable CUDA graph run on ROCm/HIP (#28173) 2026-06-19 07:16:00 -07:00
Xiaoyu Zhang 7e20e25848 [docs] Add B300 cookbook deployment options (#28697) 2026-06-18 21:52:45 -07:00
Clintandclintg6 fac11f3bc1 [AMD] Document Mori XGMI for Single-Node PD Disaggregation (#25094)
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
2026-06-18 19:25:49 -07:00
Jimmy Shong d962d18f15 docs: add --trust-remote-code to Laguna-M.1 / XS.2 cookbook configs (#28693) 2026-06-19 10:21:39 +08:00
Xinyuan Tong 61a8b42c00 docs(minimax-m3): add MMMU-Pro accuracy to B200 benchmark card (#28668) 2026-06-18 11:40:56 -07:00
Jimmy Shong f7632ef860 [Cookbook] Laguna-M.1: enable FP8 on Blackwell + drop provisional AIME numbers (#28664) 2026-06-18 09:23:53 -07:00
Jimmy Shong 0eded9e208 Add Laguna-M.1 cookbook (#28661) 2026-06-18 23:23:53 +08:00
zijiexiaandClaude Opus 4.8 3f66873304 [Docs] DeepSeek-V4 cookbook: drop --disable-flashinfer-autotune from GB300 Flash low-latency (#28590)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 00:05:21 -07:00
Ryan Zzzandzhujunyu b55cf4382d docs: add DeepSeek-V4 compressed state dtype tip (#28613)
Co-authored-by: zhujunyu <zhujunyu.666@bytedance.com>
2026-06-17 23:22:22 -07:00
Mohammad Miadh Angkad 8d4a22c5af [Docs] Add fp8 kv cache for tokenspeed mla docs (#28201) 2026-06-17 21:42:00 -07:00
Mick 05b3fd0f44 [diffusion] chore: remove ltx2 snapshot mode (#28533) 2026-06-18 10:20:21 +08:00
sglang-botandsglang-bot 1981464ba4 docs: sync LMSYS SGLang blog cards (#28589)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-18 01:23:58 +00:00
zijiexia 74e2e48c82 Introduce CpuDeviceMixin and CpuSRTPlatform (#26385) 2026-06-17 17:41:17 -07:00
Xinyuan TongandZijie Xia 72ccfec594 docs(cookbook): verify GLM-5.2 single-node B300 (FP8 + BF16) (#28460)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-17 03:47:32 +00:00
chenxu214 71b090a8e7 [Ascend]GLM 5.2 deployment (#28433) 2026-06-17 09:55:50 +08:00
amote-i 74155d32bd [NPU] [DOC] Update Ascend NPU docs: HDK 25.5.2, Triton 3.2.1.dev20260530 (#28311) 2026-06-17 09:38:31 +08:00
Thomas Wang 0d651e653b [AMD] Update v4 amd cookbook (#28423) 2026-06-16 18:15:16 -07:00
Jyothirmai Kottu b8b8992dde docs: add Amazon SageMaker AI deployment guide (#28338) 2026-06-16 13:33:22 -07:00
Xinyuan Tong 33f205d8c5 docs(cookbook): fix GLM-5.2 thinking toggle kwarg + document reasoning effort (#28454) 2026-06-16 18:17:34 +00:00
Xinyuan Tong 00081a00d5 docs(cookbook): tune GLM-5.2 MTP to 5-1-6 and simplify launch flags (#28448) 2026-06-17 01:18:34 +08:00
Xinyuan Tong 0cb6183432 docs(cookbook): add GLM-5.2 deployment cookbook (#28437) 2026-06-16 21:49:25 +08:00
Junlin Wuandronnie_zheng 2a8ea70059 [llm][npu][quant] Add W8A8 MXFP8 quantization support for Qwen3 Dense on Ascend NPU (#22352)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-16 09:45:18 +03:00
sglang-botandsglang-bot 407d3a91db docs: sync LMSYS SGLang blog cards (#28364)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-15 18:30:13 -07:00
zijiexiaandClaude Opus 4.8 7221be2cec feat(cookbook): MTP --max-running-requests callout + skill sync (#28340)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 15:53:01 -07:00
Yi Zhong 3a0dd69f8e Minor refactorings to the LFM2.5 cookbook for accuracy (#28072) 2026-06-15 14:39:10 -07:00
qinsir5522 81166f382d Fix inaccuracies and add NPU constraints in ascend_npu_profiling.mdx. (#28283) 2026-06-15 21:12:18 +08:00
jianzhao-xu f8d1d397b6 [NPU] fix ascend_docs (#28279) 2026-06-15 21:11:55 +08:00
McZyWu 8bdb007e58 [NPU] Docs op performance optimize (#28277) 2026-06-15 20:28:18 +08:00
loading66 f768344b1a [DOCS][NPU]Supplementary Notes (#28295) 2026-06-15 20:06:05 +08:00
longxin9715 7bd1a9d163 Update documentation for Ascend NPU Guide (#28284) 2026-06-15 20:02:50 +08:00
amote-i edd5eff519 [NPU] [DOC] fix issues in ascend_npu_support_new_models (#28296) 2026-06-15 19:58:33 +08:00
Xinyuan Tongandzijiexia 33f99831f8 docs(minimax-m3): refresh B200 benchmarks (tp8, piecewise) + add GPQA (#28207)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-15 00:15:38 -07:00
Lianmin Zheng f18d38d040 Revert "[AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28213) 2026-06-14 13:34:06 -07:00
Mohammad Miadh Angkad 91c63aeb4d Fix stale CUDA graph benchmark and docs refs (#28041) 2026-06-13 21:51:42 -07:00
Jared Wen 5da3b37a9d [CI] add Precision Regression Test on Nightly Run CI (#26902) 2026-06-14 12:43:57 +08:00
Chang Min Bark 93b402580c feat: add decode clear steps env var (#28160) 2026-06-13 17:15:42 -07:00
3f4a338212 [AMD][Quantization] Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs (#18182)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
2026-06-13 16:08:19 -07:00
Xinyuan Tong 47fabb52ed docs(minimax-m3): add high-concurrency throughput tip for H200 bf16 (#28150) 2026-06-13 13:09:59 -07:00
zijiexiaandClaude Opus 4.8 d988d5d681 feat(cookbook): generic config-declared flagSelects playground axis (#28128)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 11:46:34 -07:00
jianzhao-xu 29128f31fd [NPU] Best Practice Docs Splitting (#27663) 2026-06-13 22:16:20 +08:00
IBRAHIM IBRAHIMandClaude Opus 4.8 45f8d48994 docs: add llm-d page under Advanced Features (#28078)
Signed-off-by: IBRAHIM IBRAHIM <66755652+Ibrahim2595@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 18:40:31 -07:00
Yuwei AnandClaude Fable 5 6c3e429ba1 [Tiny] Cuda Graph Refactor Code Style Follow up (#28107)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 18:38:13 -07:00
amote-i c26669e37b [NPU] [DOC] Update server arguments to NPU support features page (#28083) 2026-06-12 17:16:52 -07:00
Brayden ZhongandBrayden Zhong 95867f0932 [Doc] Fix some inconsistencies in the Nemotron Cookbook (#28087)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-12 14:51:58 -07:00