Commit Graph
1483 Commits
Author SHA1 Message Date
Jimmy ShongandClaude Fable 5 4cb5aebfe0 [docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spark grids on 1cf2b8c (#35825)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 13:03:18 +08:00
Xiaoyu Zhang 96bfd2476c [diffusion] Enable SANA-Video breakable CUDA graphs (#35729) 2026-08-22 12:57:09 +08:00
Xiaoyu Zhang 83e9ece672 [diffusion] Fuse SANA-Video interleaved RoPE (#35695) 2026-08-22 12:56:45 +08:00
Mick 5290327025 [diffusion] feat: resolve hub component subfolders (#35939) 2026-08-22 12:01:44 +08:00
R0CKSTARandAlex Nails d90318b3e2 [MLX] Upgrade to Torch 2.13/MLX 0.32+ and redesign the Torch-MLX tensor bridge (#32984)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-21 18:51:42 -07:00
Baizhou Zhang 3b5909de0e [DeepSeek V4] Add W4A4 MegaMoE server flag (#35918) 2026-08-21 18:44:18 -07:00
Mick b26695a26e [diffusion] feat: reject unsupported quantized component checkpoints (#35873) 2026-08-22 09:33:43 +08:00
zijiexiaandClaude Opus 5 fe8f9d7457 [Docs] Add --prerelease=allow to cookbook uv install commands (#35920)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:51:48 -07:00
YAMY 834400705f perf: overlap Qwen shared expert with DeepEP routed experts (#34938) 2026-08-21 15:39:44 -07:00
sglang-botandsglang-bot 4d42deff0a chore: bump docs install version to 0.5.18 (#35911)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-21 13:24:03 -07:00
Xinyuan Tong 05c584c44f docs: add DSPARK speculative decoding option to Ling-3.0-flash cookbook (#35861) 2026-08-22 02:50:55 +08:00
Davis Wertheimer 70983bd7db Add SGLang Granite SWA support via existing Granite models (#35794)
Signed-off-by: Davis Wertheimer <davis.wertheimer@ibm.com>
2026-08-21 11:13:15 -07:00
Mick 8658d00764 [diffusion] feat: support loading peft lora (#35868) 2026-08-21 22:58:59 +08:00
li_maxandMick 0447ade326 [diffusion] fix: fall back to a component's default attention backend (#35796)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-21 22:56:44 +08:00
Mick 5a46d657b7 [diffusion] refactor: resolve lora weight sources deterministically (#35774) 2026-08-21 21:05:40 +08:00
Xiaoyu Zhang 39d4d65a51 [diffusion] Accelerate SANA-Video linear attention in quality=high (#35728) 2026-08-21 18:05:43 +08:00
Xiaoyu Zhang a5c52a9358 [diffusion] Enable LongCat breakable CUDA graphs (#35724) 2026-08-21 17:59:16 +08:00
Jimmy Shong 3efa057449 [docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c (#35786) 2026-08-20 21:54:51 -07:00
MickandClaude Opus 5 e0cf75d9bd [doc] standardize diffusion cookbook model pages (#34247)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 10:25:40 +08:00
Xiaoyu Zhang 7e80e889a2 [diffusion] Fuse LTX-2.5 decoder 3D RoPE (#35698) 2026-08-21 10:13:09 +08:00
Jimmy Shong 14795dcb1a [docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev image tag (#35767) 2026-08-20 23:43:55 +00:00
Yihao Wang ba8e601358 [docs] fix note formatting in sglang-d documentation (#35761) 2026-08-20 16:39:39 -07:00
Jimmy Shong 1a138e13b9 [docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753) 2026-08-20 23:34:44 +00:00
Jiajun Li a4ef828207 fix(openai): avoid duplicate routed expert in response when return_meta_info = True (#35323) 2026-08-20 15:21:58 -07:00
Liangsheng Yin 0149f56e84 [CI] Gate /rerun-test on commenter trust and remove /rerun-stage (#35750) 2026-08-20 15:05:52 -07:00
Jimmy Shong d9f6861359 [docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663) 2026-08-20 13:26:55 -07:00
Chao Shi 2ef0fe4669 TP/PP Consensus checker (#34406) 2026-08-21 01:36:03 +08:00
Mick be373395b4 [diffusion] feat: support out-of-tree models and pipelines (#35713) 2026-08-21 00:33:34 +08:00
Mick 7f8f030000 [diffusion] feat: let every layerwise component be configurable (#35688) 2026-08-20 22:38:05 +08:00
Mick 82c6fc2db9 [diffusion] quant: support pruned safetensors checkpoints for minimax-h3 (#35418) 2026-08-20 19:34:14 +08:00
21c88f8625 [diffusion] quant: support gguf (#35370)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-20 15:46:34 +08:00
Mohammad Miadh Angkad ba433bb462 [Docs] Update contribution guide (#35419) 2026-08-19 22:31:24 -07:00
Siyuan Chen 38b74d294b Add docs for TP LMHead optimizaiton (#35283) 2026-08-19 14:59:35 -07:00
Jason Wiemels defb2a3100 feat(openai): Accept the input_audio content part in chat completions (#33606) 2026-08-19 13:37:50 -07:00
Xinyuan Tong 157d8ad27a Support Intern-S2-Mobius FP8 (#34908) 2026-08-19 10:58:01 -07:00
MickandClaude Opus 5 23f2320c95 [Docs] PaddleOCR-VL: update which stage of the pipeline this serves and show real output (#35458)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:27:58 +08:00
Arseniy Mironovandronnie_zheng c57ada81e1 [Diffusion] Use current_platform instead of hardcoded "cuda" in cosmos3 guardrails (#34612)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-19 15:26:28 +03:00
jacky.cheng 574274660f [AMD] cookbook: serve Qwen3.5 MXFP4 on MI355X with an fp8_e4m3 KV cache (#35445) 2026-08-19 18:49:32 +08:00
Xiaoyu ZhangandClaude Opus 5 9113fc6d93 [docs] Add a fused-kernels page for SGLang Diffusion (#35436)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 16:32:42 +08:00
MickandClaude Fable 5 e73201e462 [diffusion] feat: support cache-dit, cfg gating, attention backend override as per-request param (#35339)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:24:03 +08:00
MickandClaude Opus 5 77fc5c128e [perf] overlap page preprocessing, pack the vit, enable prefill CUDA graph for paddle-ocr (#35318)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 08:20:55 +08:00
MickandClaude Opus 5 3f26febaff [diffusion] fix: decouple encoder parallelism from the dit parallel layout (#34713)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 00:16:09 +08:00
LinyuanLi 9485c083bb [NPU] Add mxfp4-w4a4 MOE Quantization Support for NPU (#30319) 2026-08-18 19:06:05 +03:00
WenhaoZhang 63d783bbe0 [diffusion] optimization: INT8 Linear + pluggable DiT attention backends for MiniMax-H3 on consumer-level GPUs (#34581) 2026-08-18 21:44:39 +08:00
mohbasit fc0b95e7ba Profiling Enhancements [2/3]: detailed execution step annotations (#24911) 2026-08-18 01:09:07 -07:00
Thomas Wang 27596abdc0 [AMD] Update amd k3 cookbook for PR#34580 (#35263) 2026-08-17 23:19:36 -07:00
Baizhou Zhang 53621818e4 [Docs] Enable PD disaggregation for DSV4 low-latency recipes (#35224) 2026-08-17 20:07:27 -07:00
MickandYiqi Yang d55f1c28e2 [diffusion] feat: load quantized H3 text encoder checkpoints (#34986)
Co-authored-by: Yiqi Yang <yangyiqi8787@gmail.com>
2026-08-18 09:10:54 +08:00
sglang-botandsglang-bot 9ffc2856fb docs: sync LMSYS SGLang blog cards (#35218)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-17 17:30:28 -07:00
Faradawn Yang 91144797c5 Update Qwen3.5 H200 FP8 for AgentX HiCache MTP (#35194) 2026-08-17 17:12:07 -07:00