Commit Graph
1436 Commits
Author SHA1 Message Date
MickandYiqi Yang d55f1c28e2 [diffusion] feat: load quantized H3 text encoder checkpoints (#34986)
Co-authored-by: Yiqi Yang <yangyiqi8787@gmail.com>
2026-08-18 09:10:54 +08:00
sglang-botandsglang-bot 9ffc2856fb docs: sync LMSYS SGLang blog cards (#35218)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-17 17:30:28 -07:00
Faradawn Yang 91144797c5 Update Qwen3.5 H200 FP8 for AgentX HiCache MTP (#35194) 2026-08-17 17:12:07 -07:00
Baizhou Zhang bc312d185d Clean deprecated DeepSeek V4 Environs (#34926) 2026-08-17 16:07:00 -07:00
Jimmy ShongandClaude Opus 5 b956e916ae docs(cookbook): add Qwen3.8-27B DGX Spark configs (#35121)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:04:17 -07:00
Yuhao Yang 861eca8e25 docs: add NVFP4 quantization option to Kimi-K3 deploy panel (#35168) 2026-08-17 11:01:47 -07:00
Lianmin Zheng af743371cc Clean up environ.py: remove dead env vars, unify deprecation handling, move examples to a unit test (#35060) 2026-08-17 06:53:34 -07:00
Mick e9ad8102a2 [diffusion] chore: reuse SRT SigLIP in Pi0.5 (#34992) 2026-08-17 19:33:36 +08:00
744740dbea [XPU] upgrade sglang xpu backend to PyTorch 2.13 (#31751)
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-17 18:29:15 +08:00
Jimmy ShongandClaude Fable 5 e03c53fc13 docs(cookbook): Qwen3.8-27B deployment grid rework (#35065)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 00:56:08 -07:00
Jimmy ShongandClaude Fable 5 07a28ec5cf docs: fix Qwen3.8-27B mamba ratio calculator for speculative decoding (#35064)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 22:54:57 -07:00
zijiexiaandClaude Opus 5 f019f0b064 [Docs] Feature MiniMax-H3 in the popular-models banner (#35068)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 02:09:26 +00:00
Lianmin Zheng f61f584347 Add explicit EPLB balancedness reporting modes (#34998) 2026-08-16 15:31:11 -07:00
Mick 2ee0d38a85 [diffusion] chore: refresh docs, retire stale knobs, and fix nightly attribution (#34663) 2026-08-16 15:41:08 +08:00
Mick a54de989c8 [diffusion] chore: speed up minimax-h3 vae decode on 2×h100 (#34817) 2026-08-16 15:38:21 +08:00
LinyuanLi 0da87024d3 [NPU] Add mxfp4-w4a8 MOE Quantization Support for NPU (#30318) 2026-08-16 14:03:17 +08:00
Mick e9fe58139f [diffusion] refactor: unify component residency controls (#34736) 2026-08-16 11:24:48 +08:00
Mick d106e8b23a [diffusion] chore: use native ernie prompt enhancer (#34951) 2026-08-16 09:59:51 +08:00
Mick eb6b773149 [diffusion] chore: use native qwen2.5-vl generation (#34896) 2026-08-16 09:57:39 +08:00
gilfordtingandClaude Fable 5 4a6dc267e1 [Spec] Support mamba-radix-cache-strategy extra_buffer_lazy with DFLASH (#34763)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 16:36:59 -07:00
cctryandYilong Zhao e5b3a48751 Add --http2-max-concurrent-streams server arg (#34796)
Co-authored-by: Yilong Zhao <74357408+happierpig@users.noreply.github.com>
2026-08-15 10:34:49 -07:00
Mick 4beb157e87 [diffusion] doc: define native diffusion model integration contract (#34952) 2026-08-15 23:37:34 +08:00
5c0ace30c0 [diffusion] model: support ltx-2.5 (#34471)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-15 23:36:02 +08:00
Mick 35cefd1c51 feat: add safeguards for remote media URLs (#34892) 2026-08-15 18:12:15 +08:00
Colin Z bc7e3ba66c [AMD][Quantization] Online MXFP4 quantization 4/N - NVFP4 to MXFP4 Online Requantization on AMD GPUs (#29328) 2026-08-14 21:59:39 -07:00
Baizhou Zhang 8b4faa3336 [Docs] Update Kimi-K3 installation options (#34886) 2026-08-14 16:19:23 -07:00
bfb224ff01 Add Reasoning-Aware Compression (RAC) pruning recipe for reasoning models (#32414)
Co-authored-by: Ryan Lucas <ryanluc@mit.edu>
Co-authored-by: Kayhan Behdin <kbehdin@linkedin.com>
Co-authored-by: Zhipeng Wang <zwanga@wustl.edu>
2026-08-14 15:13:45 -07:00
Han-Yin ChangandClaude Fable 5 22dde1dd5b [Docs] Fill GLM-5.2 H200 FP8 speed cells (low-latency, balanced); fix MTP notation (#31554)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:56:55 -07:00
sglang-botandsglang-bot b676793e5e docs: sync LMSYS SGLang blog cards (#32982)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-14 13:55:52 -07:00
70e291b70f [Docs] Add GB300 cells and benchmarks for Qwen3.8-27B (#34863)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:42:50 +00:00
29c6be15a4 [Docs] Add Qwen3.8-27B cookbook page (#34860)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:23:39 +00:00
Brayden ZhongandBrayden Zhong 5e65dd01a7 Remove the torchao integration (--torchao-config) (#34304)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-14 21:49:11 +08:00
amote-i fe0c18effd [NPU] [DOC] Add Qwen3.8-Max deployment tutorial on Ascend NPUs (#34836) 2026-08-14 20:02:04 +08:00
triple-muandMick a86edcdc0a [diffusion] feat: rebuild minimax-h3 adaln outputs on demand (#34650)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-14 15:33:51 +08:00
Mick 46d84f4b48 feat(cli): add extensible serve backend plugins (#34753) 2026-08-14 13:57:59 +08:00
zijiexiaandClaude Opus 5 463981922c [Cookbook] Add DeepSeek-V4-Pro-0813 (Pro Official) serving recipes (#34809)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:53:41 -07:00
Yuhao YangandClaude 6ad3f2d8fd docs: link dots3.note checkpoints, add H100 cells (#34797)
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-14 02:42:27 +00:00
Ziang Li 9d34c2809f [FlashInfer v0.6.16] Support FlashInfer CuTe DSL NVFP4 MoE quantization (#28354) 2026-08-13 17:33:46 -07:00
zijiexia abdef3c38e [Qwen] Update Docker image tag for MI300X to v0.5.17-rocm700-mi30x-20… (#34770) 2026-08-13 15:59:42 -07:00
Khoa PhamandCursor 652a2709d1 [Docs] Add decode context parallelism to advanced features (#34654)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-13 13:51:57 -07:00
c1142677a8 [do not merge] add new cookbooks (#34658)
Co-authored-by: Jianfei Wang <jianfei.wangg@outlook.com>
Co-authored-by: Jianfei Wang <905787410@qq.com>
2026-08-13 13:35:15 +00:00
fad376d3ee [CPU][QUANT] add amx cpu support for auto-round (#29593)
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Signed-off-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: Weiwei Zhang <WeiweiZhang1@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-13 15:50:59 +08:00
Carrie ChenandBrayden Zhong 6a5a9eccaa add flashinfer cute-dsl backend for mxfp8 gemm (#34042)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-13 08:50:01 +08:00
Lukas Humbel 40eaf34428 fix: make automatic NUMA binding configurable (#30394) 2026-08-12 17:40:28 -07:00
Jimmy Shong 198b7e9240 [Docs] Use Meta's canonical Muse Glimmer GGUF filename (#34626) 2026-08-12 16:07:55 -07:00
YAMY e6250c7c70 docs: update Qwen3.8 disaggregated serving configs (#34601) 2026-08-12 11:02:35 -07:00
Xinyuan Tong d21eefc94f [Docs] Rename Qwen3.8-Max-DSpark to Qwen3.8-2.4T-A95B-DSpark (#34590) 2026-08-12 15:25:48 +00:00
Yichi Zhang f28bc5a6de docs(cookbook): add BF16 recipes to Nemotron 3.5 Lightning (#34573) 2026-08-12 08:15:34 -07:00
zijiexiaandClaude Opus 5 8e7c07fae7 [Docs] Add Qwen3.8 cookbook (#34587)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 08:07:46 -07:00
ajith-sirra-amdandgiovanniguastiamd b5d1453ed2 [AMD] GLM 5.2 MXFP4 SGLANG COOKBOOK (#34379)
Signed-off-by: Sirra <asirra@amd.com>
Co-authored-by: giovanniguastiamd <giovanni.guasti@amd.com>
2026-08-12 03:39:25 -07:00