Commit Graph
1451 Commits
Author SHA1 Message Date
Siyuan Chen 38b74d294b Add docs for TP LMHead optimizaiton (#35283) 2026-08-19 14:59:35 -07:00
Jason Wiemels defb2a3100 feat(openai): Accept the input_audio content part in chat completions (#33606) 2026-08-19 13:37:50 -07:00
Xinyuan Tong 157d8ad27a Support Intern-S2-Mobius FP8 (#34908) 2026-08-19 10:58:01 -07:00
MickandClaude Opus 5 23f2320c95 [Docs] PaddleOCR-VL: update which stage of the pipeline this serves and show real output (#35458)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:27:58 +08:00
Arseniy Mironovandronnie_zheng c57ada81e1 [Diffusion] Use current_platform instead of hardcoded "cuda" in cosmos3 guardrails (#34612)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-19 15:26:28 +03:00
jacky.cheng 574274660f [AMD] cookbook: serve Qwen3.5 MXFP4 on MI355X with an fp8_e4m3 KV cache (#35445) 2026-08-19 18:49:32 +08:00
Xiaoyu ZhangandClaude Opus 5 9113fc6d93 [docs] Add a fused-kernels page for SGLang Diffusion (#35436)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 16:32:42 +08:00
MickandClaude Fable 5 e73201e462 [diffusion] feat: support cache-dit, cfg gating, attention backend override as per-request param (#35339)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:24:03 +08:00
MickandClaude Opus 5 77fc5c128e [perf] overlap page preprocessing, pack the vit, enable prefill CUDA graph for paddle-ocr (#35318)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 08:20:55 +08:00
MickandClaude Opus 5 3f26febaff [diffusion] fix: decouple encoder parallelism from the dit parallel layout (#34713)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 00:16:09 +08:00
LinyuanLi 9485c083bb [NPU] Add mxfp4-w4a4 MOE Quantization Support for NPU (#30319) 2026-08-18 19:06:05 +03:00
WenhaoZhang 63d783bbe0 [diffusion] optimization: INT8 Linear + pluggable DiT attention backends for MiniMax-H3 on consumer-level GPUs (#34581) 2026-08-18 21:44:39 +08:00
mohbasit fc0b95e7ba Profiling Enhancements [2/3]: detailed execution step annotations (#24911) 2026-08-18 01:09:07 -07:00
Thomas Wang 27596abdc0 [AMD] Update amd k3 cookbook for PR#34580 (#35263) 2026-08-17 23:19:36 -07:00
Baizhou Zhang 53621818e4 [Docs] Enable PD disaggregation for DSV4 low-latency recipes (#35224) 2026-08-17 20:07:27 -07:00
MickandYiqi Yang d55f1c28e2 [diffusion] feat: load quantized H3 text encoder checkpoints (#34986)
Co-authored-by: Yiqi Yang <yangyiqi8787@gmail.com>
2026-08-18 09:10:54 +08:00
sglang-botandsglang-bot 9ffc2856fb docs: sync LMSYS SGLang blog cards (#35218)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-17 17:30:28 -07:00
Faradawn Yang 91144797c5 Update Qwen3.5 H200 FP8 for AgentX HiCache MTP (#35194) 2026-08-17 17:12:07 -07:00
Baizhou Zhang bc312d185d Clean deprecated DeepSeek V4 Environs (#34926) 2026-08-17 16:07:00 -07:00
Jimmy ShongandClaude Opus 5 b956e916ae docs(cookbook): add Qwen3.8-27B DGX Spark configs (#35121)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:04:17 -07:00
Yuhao Yang 861eca8e25 docs: add NVFP4 quantization option to Kimi-K3 deploy panel (#35168) 2026-08-17 11:01:47 -07:00
Lianmin Zheng af743371cc Clean up environ.py: remove dead env vars, unify deprecation handling, move examples to a unit test (#35060) 2026-08-17 06:53:34 -07:00
Mick e9ad8102a2 [diffusion] chore: reuse SRT SigLIP in Pi0.5 (#34992) 2026-08-17 19:33:36 +08:00
744740dbea [XPU] upgrade sglang xpu backend to PyTorch 2.13 (#31751)
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-17 18:29:15 +08:00
Jimmy ShongandClaude Fable 5 e03c53fc13 docs(cookbook): Qwen3.8-27B deployment grid rework (#35065)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 00:56:08 -07:00
Jimmy ShongandClaude Fable 5 07a28ec5cf docs: fix Qwen3.8-27B mamba ratio calculator for speculative decoding (#35064)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 22:54:57 -07:00
zijiexiaandClaude Opus 5 f019f0b064 [Docs] Feature MiniMax-H3 in the popular-models banner (#35068)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 02:09:26 +00:00
Lianmin Zheng f61f584347 Add explicit EPLB balancedness reporting modes (#34998) 2026-08-16 15:31:11 -07:00
Mick 2ee0d38a85 [diffusion] chore: refresh docs, retire stale knobs, and fix nightly attribution (#34663) 2026-08-16 15:41:08 +08:00
Mick a54de989c8 [diffusion] chore: speed up minimax-h3 vae decode on 2×h100 (#34817) 2026-08-16 15:38:21 +08:00
LinyuanLi 0da87024d3 [NPU] Add mxfp4-w4a8 MOE Quantization Support for NPU (#30318) 2026-08-16 14:03:17 +08:00
Mick e9fe58139f [diffusion] refactor: unify component residency controls (#34736) 2026-08-16 11:24:48 +08:00
Mick d106e8b23a [diffusion] chore: use native ernie prompt enhancer (#34951) 2026-08-16 09:59:51 +08:00
Mick eb6b773149 [diffusion] chore: use native qwen2.5-vl generation (#34896) 2026-08-16 09:57:39 +08:00
gilfordtingandClaude Fable 5 4a6dc267e1 [Spec] Support mamba-radix-cache-strategy extra_buffer_lazy with DFLASH (#34763)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 16:36:59 -07:00
cctryandYilong Zhao e5b3a48751 Add --http2-max-concurrent-streams server arg (#34796)
Co-authored-by: Yilong Zhao <74357408+happierpig@users.noreply.github.com>
2026-08-15 10:34:49 -07:00
Mick 4beb157e87 [diffusion] doc: define native diffusion model integration contract (#34952) 2026-08-15 23:37:34 +08:00
5c0ace30c0 [diffusion] model: support ltx-2.5 (#34471)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-15 23:36:02 +08:00
Mick 35cefd1c51 feat: add safeguards for remote media URLs (#34892) 2026-08-15 18:12:15 +08:00
Colin Z bc7e3ba66c [AMD][Quantization] Online MXFP4 quantization 4/N - NVFP4 to MXFP4 Online Requantization on AMD GPUs (#29328) 2026-08-14 21:59:39 -07:00
Baizhou Zhang 8b4faa3336 [Docs] Update Kimi-K3 installation options (#34886) 2026-08-14 16:19:23 -07:00
bfb224ff01 Add Reasoning-Aware Compression (RAC) pruning recipe for reasoning models (#32414)
Co-authored-by: Ryan Lucas <ryanluc@mit.edu>
Co-authored-by: Kayhan Behdin <kbehdin@linkedin.com>
Co-authored-by: Zhipeng Wang <zwanga@wustl.edu>
2026-08-14 15:13:45 -07:00
Han-Yin ChangandClaude Fable 5 22dde1dd5b [Docs] Fill GLM-5.2 H200 FP8 speed cells (low-latency, balanced); fix MTP notation (#31554)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:56:55 -07:00
sglang-botandsglang-bot b676793e5e docs: sync LMSYS SGLang blog cards (#32982)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-14 13:55:52 -07:00
70e291b70f [Docs] Add GB300 cells and benchmarks for Qwen3.8-27B (#34863)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:42:50 +00:00
29c6be15a4 [Docs] Add Qwen3.8-27B cookbook page (#34860)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:23:39 +00:00
Brayden ZhongandBrayden Zhong 5e65dd01a7 Remove the torchao integration (--torchao-config) (#34304)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-14 21:49:11 +08:00
amote-i fe0c18effd [NPU] [DOC] Add Qwen3.8-Max deployment tutorial on Ascend NPUs (#34836) 2026-08-14 20:02:04 +08:00
triple-muandMick a86edcdc0a [diffusion] feat: rebuild minimax-h3 adaln outputs on demand (#34650)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-14 15:33:51 +08:00
Mick 46d84f4b48 feat(cli): add extensible serve backend plugins (#34753) 2026-08-14 13:57:59 +08:00