b44e248682
[AMD] [GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950 ( #38546 )
...
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: andyluo7 <andy.luo@amd.com >
Co-authored-by: Kevin Mi <mikevin920@yahoo.com >
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
2026-09-21 21:07:37 -07:00
Cheng Wan
acac4dd9d9
[Refactor] Clean up parallel runtime comments ( #40632 )
2026-09-21 14:32:22 -07:00
Liangsheng Yin
76a9065bef
[Fix] Raise on undelivered embeddings in send_with_url, fix broken tests ( #40502 )
2026-09-20 17:35:27 -07:00
42875bcd2a
fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global ( #38932 )
...
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com >
2026-09-20 19:28:54 -04:00
DarkSharpness and BBuf
d1acbe0746
[DSV4.1] Big fused wo_a quant ( #39957 )
...
Co-authored-by: BBuf <1182563586@qq.com >
2026-09-19 19:46:53 +08:00
Cheng Wan
567d5925fe
Fix mxfp4 padding test stubbing an accessor the module no longer imports ( #40308 )
2026-09-19 00:43:56 -07:00
Liangsheng Yin
d507accadc
[Test] Drop dead and strictly-subsumed CI test registrations ( #40264 )
2026-09-18 17:49:13 -07:00
Cheng Wan
afe71f4b9e
Read process groups through the runtime context ( #40068 )
2026-09-18 17:40:32 -07:00
Sam (Kesen Li)
d346b214fb
feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding ( #36340 )
2026-09-18 14:50:46 -07:00
a6cf05817f
dsv4.1: remaining model and runtime integration ( #38798 )
...
Co-authored-by: BBuf <1182563586@qq.com >
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai >
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com >
Co-authored-by: Zhichen Zeng <zczeng@uw.edu >
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
2026-09-18 02:55:30 -07:00
Liangsheng Yin
1b200ffaaa
[Quant] Serve 32-wide-K ue8m0 block-FP8 linears through the FlashInfer MXFP8 GEMMs ( #40039 )
2026-09-18 02:51:12 -07:00
cebca698e2
[Qwen3.8] Enable NVIDIA NVFP4 on DGX Spark with file-backed PLE and PDL router fix ( #39126 )
...
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com >
Co-authored-by: rdxa <rdxa@rdxa-int-spark-01.yvb.moe >
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
Co-authored-by: Manrique <nanomlm@gmail.com >
Co-authored-by: yhyang201 <yhyang201@gmail.com >
2026-09-13 16:23:41 +08:00
Tuan Nguyen Gia
7c195b9151
[AMD] Fix Quark load of MiniMax-M3 MXFP4 index_qkv_proj ( #37254 )
2026-09-12 00:50:41 -07:00
Brayden Zhong and Brayden Zhong
c0b790cf7f
Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention ( #32114 )
...
Co-authored-by: Brayden Zhong <brayden@radixark.ai >
2026-09-10 15:12:01 +08:00
Mohammad Miadh Angkad and Mohammad Angkad
a711785475
[Kernel] Drop the vendored dense BF16 GEMM port in favor of FlashInfer 0.6.18 ( #38124 )
...
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
2026-09-07 12:59:33 -07:00
214313ee79
Fuse Nemotron latent MoE projection and shared add ( #30430 )
...
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com >
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-09-07 10:26:19 +08:00
sglang-bot and sglang-bot
6252993afe
chore: update CI test est_time values ( #38238 )
...
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com >
2026-09-06 17:49:41 -07:00
Beihao Zhou
1e6f18bfeb
[MoE Refactor] Migrate SM100 trtllm-gen mxfp4 MoE onto MoeRunner ( #32405 )
2026-09-05 13:48:10 +00:00
kk and wunhuang
a6001478f4
[AMD] Perf Kimi-K3 MoE optimization ( #33838 )
...
Co-authored-by: wunhuang <wunhuang@amd.com >
2026-09-03 02:28:28 -07:00
Alex Nails and Alison Shao
28262c20df
[CI][RFC] Replace black-jupyter with ruff-format ( #37210 )
...
Co-authored-by: Alison Shao <a.shao@wustl.edu >
2026-09-02 19:46:08 -07:00
Lee Nau and Yangmin Li
6e41f1ad29
[Fix] Preserve FP32 in SM107 MXFP8 fallback ( #37489 )
...
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
2026-09-02 17:55:14 -07:00
26f760d5c0
[CPU] Support FP8 KV cache ( #32733 )
...
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com >
Co-authored-by: mingfeima <mingfei.ma@intel.com >
2026-09-02 10:20:54 +08:00
Yuan Luo and luoyuan.luo
5b04408784
[MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE ( #34967 )
...
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com >
2026-08-31 20:04:41 -07:00
Cheng Wan
7e751153eb
[Config] Round 5.1: the published-side readers ask the bags, and a platform fact gets one address ( #37086 )
2026-08-30 02:18:33 -07:00
+8
5f216fc33f
qwen 3.8 rebase ( #35758 )
...
Co-authored-by: cherichy <cherichy@outlook.com >
Co-authored-by: guangyunh-nv <guangyunh@nvidia.com >
Co-authored-by: jiahanc <jiahanc@nvidia.com >
Co-authored-by: jinyangyuan-nvidia <joyuan@nvidia.com >
Co-authored-by: Cheng Hang <chang@nvidia.com >
Co-authored-by: Yicheng Qiang <yqiang@nvidia.com >
Co-authored-by: Sam Li <lsam@nvidia.com >
Co-authored-by: Tom-Zheng <tizheng@nvidia.com >
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
Co-authored-by: xiaoweiw-nv <xiaoweiw@nvidia.com >
Co-authored-by: Zheng Li <lizheng.cs@zju.edu.cn >
Co-authored-by: yizhang2077 <1109276519@qq.com >
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com >
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai >
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-28 20:41:34 -07:00
Cheng Wan and Claude Opus 5
fd40a331bf
config: a parallel size has one spelling; a patched scope declares its own ( #36621 )
...
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
2026-08-27 12:56:42 -07:00
20621aa14b
[Model] Support Ling-3.0-flash (BailingMoeV3) ( #33561 )
...
Signed-off-by: JustinTong <justintong0323@gmail.com >
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com >
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com >
Co-authored-by: 翎悦 <vito.yy@antgroup.com >
Co-authored-by: 羽癫 <yudian.zy@antgroup.com >
Co-authored-by: tiwei.btw <tiwei.btw@antgroup.com >
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
Co-authored-by: 文赋 <zibin.zb@antgroup.com >
Co-authored-by: JustinTong <justintong0323@gmail.com >
2026-08-26 17:27:23 -07:00
27c36368b6
fix(moe): guard FP8 delegate activation params ( #36275 )
...
Signed-off-by: jikuixie <jikuixie@gmail.com >
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Co-authored-by: shyeh25 <206795756+shyeh25@users.noreply.github.com >
2026-08-26 12:12:21 +00:00
elvischenv
46d9427b91
Fix MXFP8 MoE weight sizing for non-gated models ( #36097 )
2026-08-25 22:09:09 +08:00
guzekai01
716a6bf10c
feat(humming): support native W4AFP8 checkpoint schemas ( #32033 )
2026-08-24 18:59:27 +08:00
Sahithi Chigurupati and Mohammad Miadh Angkad
44db041700
[NVIDIA] Fix SM107 MXFP8 activation prep ( #35405 )
...
Signed-off-by: Sahithi Chigurupati <chigurupati.sahithi@gmail.com >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-08-23 18:17:03 +08:00
Jimmy Shong
710267dc4c
[Quant] Load compressed-tensors kv_cache_scheme scales ( #35455 )
2026-08-20 19:17:59 +08:00
Jimmy Shong
5375babbac
[Quant] Load compressed-tensors quantized lm_head instead of value-casting it ( #35228 )
2026-08-19 15:37:45 -07:00
YAMY
5f12839591
[Fix] Support Kimi-K3 ModelOpt mixed NVFP4/FP8 checkpoint ( #35077 )
2026-08-19 08:13:45 -07:00
Enrique Shockwave
92b1d382c7
[Fix] Correct dense FP8 Marlin bias ordering ( #35020 )
2026-08-17 03:43:36 +00:00
Mohammad Miadh Angkad and Mohammad Angkad
6ab4b99bc2
[Quantization] Fix GPTQ scheme attachment broken by LinearBase.scheme default ( #34962 )
...
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
2026-08-16 00:48:02 -07:00
Spandan Tiwari
f68517f644
[AMD][Quantization][Bugfix] Fix bug related to fp8 max on gfx95x for per-token-group quant (ROCm) ( #30900 )
2026-08-15 19:54:16 -07:00
Colin Z
bc7e3ba66c
[AMD][Quantization] Online MXFP4 quantization 4/N - NVFP4 to MXFP4 Online Requantization on AMD GPUs ( #29328 )
2026-08-14 21:59:39 -07:00
Carrie Chen and Brayden Zhong
6a5a9eccaa
add flashinfer cute-dsl backend for mxfp8 gemm ( #34042 )
...
Co-authored-by: Brayden Zhong <brayden@radixark.ai >
2026-08-13 08:50:01 +08:00
fde9ad2531
[Feature] Add Muse Glimmer model support ( #34262 )
...
Co-authored-by: sglang-bot <232288953+sglang-bot@users.noreply.github.com >
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai >
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com >
Co-authored-by: hnyls2002 <lsyincs@gmail.com >
Co-authored-by: Alex Nails <alex.nails@radixark.ai >
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
2026-08-11 15:41:52 -07:00
Cheng Wan
e216c2bc59
config: the runner and scheduler read resolved config from the bags ( #34095 )
2026-08-09 14:45:11 -07:00
Ziang Li and Brayden Zhong
4ad990ba7d
[ModelOpt FP4] Support online MoE weight quantization ( #33115 )
...
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca >
2026-08-06 11:01:55 -07:00
Cheng Wan
4ea227fa91
config: the draft runner carries its own attention backend
...
`build_draft_tp_worker` built a `ServerArgs` variant whose only job was to make
four config reads answer with the draft's backend instead of the target's, and
published it for the duration of the build so the bags agreed. The backend is a
per-runner fact — target and draft coexist in one process — so it moves onto the
runner, and the variant and the construction-time publish both go away.
`ModelRunner` takes `draft_attention_backend` and resolves the runner's effective
value once (`resolve_draft_attention_backend`: the algorithm's resolved backend,
else `--speculative-draft-attention-backend`, else None for a target runner);
`TpModelWorker` threads it to both runner constructions.
`resolve_attention_backend_strs` reads it off the runner, and `ModelRunner`
stamps the resolved pair *before* building backends so a backend can read it
while it constructs — which is what the FlashInfer KV-access check needs now that
it no longer asks the config. `configure_kv_cache_dtype` and the draft backend
factory read the runner too.
One latent bug falls out: the non-hybrid branch of the backend build ignored the
resolved pair and re-read `server_args.attention_backend`, which is why the
variant had to set that field as well as the split pair. It now uses the value
that was resolved for the runner.
`draft_server_args_overrides` and the `preserve_config()` publish switch are
deleted; with them goes the last production `ServerArgs.derive` outside
pre-publish config building, and the last construction-time publish. The
chunked-prefix gate the target resolved simply stays in the bags, since nothing
re-projects them.
2026-08-05 19:32:24 -07:00
Liangsheng Yin
1a045669e4
[CI] Merge tokenizer worker tests and drop redundant triton attention e2e ( #33641 )
2026-08-05 11:55:11 -07:00
Liangsheng Yin
198a3bc29b
[Test] Route GEMM backend UTs through real layer modules and weight loaders ( #33615 )
2026-08-04 20:53:26 -07:00
Liangsheng Yin
76dc89f5aa
[Test] Replace NVFP4 MoE runner backend e2e matrix with a layer-level unit test ( #33611 )
2026-08-04 16:03:29 -07:00
Liangsheng Yin
a0b3f1dde6
[Test] Replace GEMM backend e2e matrices with layer-level unit tests ( #33596 )
2026-08-04 15:50:41 -07:00
+26
abddb1c7e9
[Kimi] Support kimi-k3 ( #32541 )
...
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com >
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
Co-authored-by: Mick <mickjagger19@icloud.com >
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com >
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com >
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com >
Co-authored-by: Chunan Zeng <zcnrex@gmail.com >
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai >
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai >
Co-authored-by: Zijie Xia <37504505+zijiexia@users.noreply.github.com >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: zhangxiaohao <1024393531@qq.com >
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
Co-authored-by: Julien Lin <jullin@nvidia.com >
Co-authored-by: Hao Phan <htphan@nvidia.com >
Co-authored-by: Thomas Wang <1am9trash@gmail.com >
Co-authored-by: RolaoDenthu <xinyisong0111@gmail.com >
Co-authored-by: pigeonsoup <32922982+pigeonsoup@users.noreply.github.com >
Co-authored-by: HaiShaw <hixiao@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com >
Co-authored-by: Lee Nau <lee.nau@gmail.com >
Co-authored-by: HMING <126185151+Hearum@users.noreply.github.com >
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com >
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com >
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai >
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai >
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai >
Co-authored-by: Hanming Lu <hanminglu@meta.com >
Co-authored-by: Xinyi Song <xinyis10@illinois.edu >
2026-08-04 13:22:49 -07:00
YAMY and Sam Li
5fe97637df
Support DeepGEMM for standard MoE dispatch ( #33128 )
...
Co-authored-by: Sam Li <lsam@nvidia.com >
2026-08-02 21:48:13 -07:00
Mohammad Miadh Angkad
a55e1764a2
Enable GPT-OSS FlashInfer MXFP4 on SM120 ( #32668 )
2026-07-30 00:04:23 +00:00