Commit Graph
221 Commits
Author SHA1 Message Date
a6cf05817f dsv4.1: remaining model and runtime integration (#38798)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-09-18 02:55:30 -07:00
Ziang Li c46bf5e990 [MoE] Disable FlashInfer fused finalize by default for numerical accuracy (#40105) 2026-09-18 00:15:57 -07:00
amote-i 6952538980 [NPU] [DOC] delete unsupported api --disable-hybrid-swa-memory for npu (#40058) 2026-09-18 14:31:05 +08:00
Chi McIsaacandMick Qian 6215aecd51 [diffusion] feat: add metrics support (#19084)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-18 13:22:03 +08:00
Arseniy Mironov a1b4ec02ae [NPU][Diffusion] FA MXFP8 and modelslim w4a4f8 and w8a8f8 support for Wan2.2 and FLUX (#39438) 2026-09-17 13:13:05 +03:00
yl3469andShuwen Wang 1a90ae6727 Add Agentic-Aware Tail-Optimized LRU eviction to the unified radix cache (#34012)
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
2026-09-17 17:52:59 +08:00
MickandMick Qian c525ed8f02 [diffusion] fix: preserve explicit attention backends during autotune (#39882)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-17 13:28:55 +08:00
Liangsheng YinandBBuf 464fffbec8 dsv4.1: chat encoding and tool parsing (#39665)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-16 20:49:42 -07:00
amote-i 575163ff32 [NPU] [DOC] delete unsupported models in npu docs (#39555) 2026-09-15 14:55:17 +08:00
amote-i bdf8886ad3 [NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389) 2026-09-15 10:33:14 +08:00
6220f45d8e [PD] Add /v1/responses support to the HTTP PD router (#36141)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-13 22:48:13 +08:00
Shangming CaiandXinyuan Tong 14a131ad5b [PD][OpenAI] Gate /v1/responses persistence behind --enable-response-store, default off (#39122)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-13 21:20:47 +08:00
cebca698e2 [Qwen3.8] Enable NVIDIA NVFP4 on DGX Spark with file-backed PLE and PDL router fix (#39126)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: rdxa <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Manrique <nanomlm@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
2026-09-13 16:23:41 +08:00
desmond-intel 6dc7b3421b Inference Support Mamba 2 and 1 (#34556) 2026-09-12 20:51:18 +08:00
Khoa PhamandClaude Fable 5.1 6ba96d329f [DCP] Resolve --dcp-comm-backend to fi_a2a/a2a by default for every model (#39165)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 23:48:42 -07:00
ff1ce11348 [diffusion] model: support VDN-H3 with a hybrid_window_attn_h3 backend (#37903)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Haocheng Xi <xihc@berkeley.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-09-12 11:36:32 +08:00
Vignesh Sethuraman 0d1bea77da [AMD] Allow aiter attention backend for Gemma-4 (#38758) 2026-09-11 18:48:26 -07:00
MickandMick Qian e016de462c [diffusion] CI: expose nightly server telemetry coverage (#38782)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-11 23:07:37 +08:00
Siju Samuel 67d3a2ea57 [XPU] Make checkpoint_engine worker device-agnostic (#32382) 2026-09-11 09:56:39 +08:00
Mohammad Miadh AngkadandMohammad Angkad 52c191da52 [Deps] Retire the CUDA 12 lane (#38404)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-10 16:58:09 -07:00
MickandMick Qian c9c26d56b2 [diffusion] docs: sync snapshot and minimax-h3 subblock features (#38784)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-10 17:23:43 +08:00
1b77f498a0 [NVIDIA] Support flashinfer Mega Moe (#31470)
Co-authored-by: djns99 <40156487+djns99@users.noreply.github.com>
Co-authored-by: 云挚 <ningyunxiao.nyx@antgroup.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-10 00:22:47 -07:00
Brayden ZhongandBrayden Zhong c0b790cf7f Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-09-10 15:12:01 +08:00
MickandMick Qian ce555ed82a [diffusion] refactor: refactor utility ownership and document helper placement (#38699)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-10 09:11:01 +08:00
Alison Shao a7e00b7576 [CI] Answer unrecognized slash commands instead of skipping silently (#38736) 2026-09-09 18:01:18 -07:00
Alison Shao 2948a62a6f [CI] Add /run-full-ci and /run-extra-ci slash commands (#38734) 2026-09-09 13:56:38 -07:00
Even Zhou dba34cc964 [NPU] Bump memfabric and sgl-kernel-npu versions in docs and pyproject_npu.toml (#38437) 2026-09-09 19:53:49 +08:00
MickandMick Qian 00a9028e87 [diffusion] feat: add explicit snapshot-offload component residency (#38535)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-09 13:44:54 +08:00
Baizhou Zhang e54ff1efb9 [CP V1 Deprecation 5/5] Update prefill CP documentation (#36230) 2026-09-08 19:27:49 -07:00
MickandMick Qian a8e45f16cc [diffusion] feat: support mixed INT8 embeddings and Comfy NVFP4 encoders for minimax-h3 (#38506)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-09 08:36:43 +08:00
MickandMick Qian 65400bb420 [diffusion] model: support MiniMax-H3 singularity hybrid checkpoints (#38455)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-09 08:35:39 +08:00
Cheng Wan db272201a2 [Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement (#38375) 2026-09-08 16:42:12 -07:00
Xinyuan Tong 482e9f257b Add Ling-3.0-flash-VL cookbook (#38434) 2026-09-08 23:00:24 +08:00
HZY a6b542813f fix(glm-5.2-nvfp4): bound Mooncake synchronous transfer batches (#32758) 2026-09-08 22:14:33 +08:00
Wuhen Duan dfd9b5c2a4 [NPU] Enable non-greedy MTP sampling (#32495) 2026-09-08 11:18:01 +08:00
Siju Samuel 2358916d5a [Feature][Intel XPU] Add memory saver support for Intel XPU via upstream torch_memory_saver (#29935) 2026-09-08 09:38:51 +08:00
Mick 15d2cbcc90 [diffusion] CI: validate every repeated server request (#38185) 2026-09-07 14:36:50 +08:00
39a80354aa [MUSA] Add installation guide and Dockerfile (#36709)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-09-06 20:13:53 -05:00
Mick f3d05644db [diffusion] docs+skill: document which components to stream under layerwise offload (#35674) 2026-09-06 23:12:36 +08:00
Mick 938dc5621d [diffusion] refactor: reuse plain state-dict loading without per-model classes (#38127) 2026-09-06 18:39:21 +08:00
Xiaoyu Zhang d61378af77 docs(diffusion): add per-model tuning decision table to performance guide (#38148) 2026-09-06 15:26:40 +08:00
bd16c22a04 [diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-05 18:12:30 +08:00
Mick 0ea8378085 [diffusion] feat: support request-scoped skip-softmax attention (#37959) 2026-09-05 13:50:17 +08:00
Shuwen WangandClaude Opus 5 4b44a1cde2 [Refactor] Let eviction policies take construction parameters (#37795)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 13:48:30 +08:00
09f542b23a [CI] Add /rerun-test --changed to rerun every test file a PR modifies (#37618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-09-04 22:38:11 -07:00
320bdd1ee2 [Docs] Document --retraction-policy, --return-hidden-states-mode, --language-model-only (#37989)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: mottopanikeiku <fcetin@hawk.iit.edu>
Co-authored-by: alp <falpercetin@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-05 05:26:28 +08:00
alpandXinyuan Tong 3e873c2110 [Docs] Clarify OpenAI chat template defaults (#32172)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-05 05:22:05 +08:00
4349538c02 [model] add cosmos3 reasoner to llm only inference (#33572)
Signed-off-by: joeltg <joel@reflection.ai>
Signed-off-by: Joe Rowell <joe@poolside.ai>
Co-authored-by: Dawid Majchrowski <dmajchrowski@nvidia.com>
Co-authored-by: Kedi Wu <kediw@nvidia.com>
Co-authored-by: Kedi Wu <31940276+kediwu0331@users.noreply.github.com>
Co-authored-by: Joel Gustafson <joelgustafson@protonmail.com>
2026-09-04 22:11:58 +08:00
Brian 2216697f90 [Docs] Refresh TPU model list and link cookbooks (#37750) 2026-09-04 17:28:57 +08:00
Polisetty V R K Jyothendra Varma b168f905c8 [Intel GPU] Align XPU toml file for rust support (#31031)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
2026-09-04 14:48:30 +08:00