Mohammad Miadh Angkad and Mohammad Angkad
a711785475
[Kernel] Drop the vendored dense BF16 GEMM port in favor of FlashInfer 0.6.18 ( #38124 )
...
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
2026-09-07 12:59:33 -07:00
214313ee79
Fuse Nemotron latent MoE projection and shared add ( #30430 )
...
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com >
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-09-07 10:26:19 +08:00
sglang-bot and sglang-bot
6252993afe
chore: update CI test est_time values ( #38238 )
...
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com >
2026-09-06 17:49:41 -07:00
Beihao Zhou
1e6f18bfeb
[MoE Refactor] Migrate SM100 trtllm-gen mxfp4 MoE onto MoeRunner ( #32405 )
2026-09-05 13:48:10 +00:00
kk and wunhuang
a6001478f4
[AMD] Perf Kimi-K3 MoE optimization ( #33838 )
...
Co-authored-by: wunhuang <wunhuang@amd.com >
2026-09-03 02:28:28 -07:00
Alex Nails and Alison Shao
28262c20df
[CI][RFC] Replace black-jupyter with ruff-format ( #37210 )
...
Co-authored-by: Alison Shao <a.shao@wustl.edu >
2026-09-02 19:46:08 -07:00
Lee Nau and Yangmin Li
6e41f1ad29
[Fix] Preserve FP32 in SM107 MXFP8 fallback ( #37489 )
...
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
2026-09-02 17:55:14 -07:00
26f760d5c0
[CPU] Support FP8 KV cache ( #32733 )
...
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com >
Co-authored-by: mingfeima <mingfei.ma@intel.com >
2026-09-02 10:20:54 +08:00
Yuan Luo and luoyuan.luo
5b04408784
[MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE ( #34967 )
...
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com >
2026-08-31 20:04:41 -07:00
Cheng Wan
7e751153eb
[Config] Round 5.1: the published-side readers ask the bags, and a platform fact gets one address ( #37086 )
2026-08-30 02:18:33 -07:00
+8
5f216fc33f
qwen 3.8 rebase ( #35758 )
...
Co-authored-by: cherichy <cherichy@outlook.com >
Co-authored-by: guangyunh-nv <guangyunh@nvidia.com >
Co-authored-by: jiahanc <jiahanc@nvidia.com >
Co-authored-by: jinyangyuan-nvidia <joyuan@nvidia.com >
Co-authored-by: Cheng Hang <chang@nvidia.com >
Co-authored-by: Yicheng Qiang <yqiang@nvidia.com >
Co-authored-by: Sam Li <lsam@nvidia.com >
Co-authored-by: Tom-Zheng <tizheng@nvidia.com >
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
Co-authored-by: xiaoweiw-nv <xiaoweiw@nvidia.com >
Co-authored-by: Zheng Li <lizheng.cs@zju.edu.cn >
Co-authored-by: yizhang2077 <1109276519@qq.com >
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com >
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai >
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-28 20:41:34 -07:00
Cheng Wan and Claude Opus 5
fd40a331bf
config: a parallel size has one spelling; a patched scope declares its own ( #36621 )
...
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
2026-08-27 12:56:42 -07:00
20621aa14b
[Model] Support Ling-3.0-flash (BailingMoeV3) ( #33561 )
...
Signed-off-by: JustinTong <justintong0323@gmail.com >
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com >
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com >
Co-authored-by: 翎悦 <vito.yy@antgroup.com >
Co-authored-by: 羽癫 <yudian.zy@antgroup.com >
Co-authored-by: tiwei.btw <tiwei.btw@antgroup.com >
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
Co-authored-by: 文赋 <zibin.zb@antgroup.com >
Co-authored-by: JustinTong <justintong0323@gmail.com >
2026-08-26 17:27:23 -07:00
27c36368b6
fix(moe): guard FP8 delegate activation params ( #36275 )
...
Signed-off-by: jikuixie <jikuixie@gmail.com >
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Co-authored-by: shyeh25 <206795756+shyeh25@users.noreply.github.com >
2026-08-26 12:12:21 +00:00
elvischenv
46d9427b91
Fix MXFP8 MoE weight sizing for non-gated models ( #36097 )
2026-08-25 22:09:09 +08:00
guzekai01
716a6bf10c
feat(humming): support native W4AFP8 checkpoint schemas ( #32033 )
2026-08-24 18:59:27 +08:00
Sahithi Chigurupati and Mohammad Miadh Angkad
44db041700
[NVIDIA] Fix SM107 MXFP8 activation prep ( #35405 )
...
Signed-off-by: Sahithi Chigurupati <chigurupati.sahithi@gmail.com >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-08-23 18:17:03 +08:00
Jimmy Shong
710267dc4c
[Quant] Load compressed-tensors kv_cache_scheme scales ( #35455 )
2026-08-20 19:17:59 +08:00
Jimmy Shong
5375babbac
[Quant] Load compressed-tensors quantized lm_head instead of value-casting it ( #35228 )
2026-08-19 15:37:45 -07:00
YAMY
5f12839591
[Fix] Support Kimi-K3 ModelOpt mixed NVFP4/FP8 checkpoint ( #35077 )
2026-08-19 08:13:45 -07:00
Enrique Shockwave
92b1d382c7
[Fix] Correct dense FP8 Marlin bias ordering ( #35020 )
2026-08-17 03:43:36 +00:00
Mohammad Miadh Angkad and Mohammad Angkad
6ab4b99bc2
[Quantization] Fix GPTQ scheme attachment broken by LinearBase.scheme default ( #34962 )
...
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
2026-08-16 00:48:02 -07:00
Spandan Tiwari
f68517f644
[AMD][Quantization][Bugfix] Fix bug related to fp8 max on gfx95x for per-token-group quant (ROCm) ( #30900 )
2026-08-15 19:54:16 -07:00
Colin Z
bc7e3ba66c
[AMD][Quantization] Online MXFP4 quantization 4/N - NVFP4 to MXFP4 Online Requantization on AMD GPUs ( #29328 )
2026-08-14 21:59:39 -07:00
Carrie Chen and Brayden Zhong
6a5a9eccaa
add flashinfer cute-dsl backend for mxfp8 gemm ( #34042 )
...
Co-authored-by: Brayden Zhong <brayden@radixark.ai >
2026-08-13 08:50:01 +08:00
fde9ad2531
[Feature] Add Muse Glimmer model support ( #34262 )
...
Co-authored-by: sglang-bot <232288953+sglang-bot@users.noreply.github.com >
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai >
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com >
Co-authored-by: hnyls2002 <lsyincs@gmail.com >
Co-authored-by: Alex Nails <alex.nails@radixark.ai >
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
2026-08-11 15:41:52 -07:00
Cheng Wan
e216c2bc59
config: the runner and scheduler read resolved config from the bags ( #34095 )
2026-08-09 14:45:11 -07:00
Ziang Li and Brayden Zhong
4ad990ba7d
[ModelOpt FP4] Support online MoE weight quantization ( #33115 )
...
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca >
2026-08-06 11:01:55 -07:00
Cheng Wan
4ea227fa91
config: the draft runner carries its own attention backend
...
`build_draft_tp_worker` built a `ServerArgs` variant whose only job was to make
four config reads answer with the draft's backend instead of the target's, and
published it for the duration of the build so the bags agreed. The backend is a
per-runner fact — target and draft coexist in one process — so it moves onto the
runner, and the variant and the construction-time publish both go away.
`ModelRunner` takes `draft_attention_backend` and resolves the runner's effective
value once (`resolve_draft_attention_backend`: the algorithm's resolved backend,
else `--speculative-draft-attention-backend`, else None for a target runner);
`TpModelWorker` threads it to both runner constructions.
`resolve_attention_backend_strs` reads it off the runner, and `ModelRunner`
stamps the resolved pair *before* building backends so a backend can read it
while it constructs — which is what the FlashInfer KV-access check needs now that
it no longer asks the config. `configure_kv_cache_dtype` and the draft backend
factory read the runner too.
One latent bug falls out: the non-hybrid branch of the backend build ignored the
resolved pair and re-read `server_args.attention_backend`, which is why the
variant had to set that field as well as the split pair. It now uses the value
that was resolved for the runner.
`draft_server_args_overrides` and the `preserve_config()` publish switch are
deleted; with them goes the last production `ServerArgs.derive` outside
pre-publish config building, and the last construction-time publish. The
chunked-prefix gate the target resolved simply stays in the bags, since nothing
re-projects them.
2026-08-05 19:32:24 -07:00
Liangsheng Yin
1a045669e4
[CI] Merge tokenizer worker tests and drop redundant triton attention e2e ( #33641 )
2026-08-05 11:55:11 -07:00
Liangsheng Yin
198a3bc29b
[Test] Route GEMM backend UTs through real layer modules and weight loaders ( #33615 )
2026-08-04 20:53:26 -07:00
Liangsheng Yin
76dc89f5aa
[Test] Replace NVFP4 MoE runner backend e2e matrix with a layer-level unit test ( #33611 )
2026-08-04 16:03:29 -07:00
Liangsheng Yin
a0b3f1dde6
[Test] Replace GEMM backend e2e matrices with layer-level unit tests ( #33596 )
2026-08-04 15:50:41 -07:00
+26
abddb1c7e9
[Kimi] Support kimi-k3 ( #32541 )
...
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com >
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
Co-authored-by: Mick <mickjagger19@icloud.com >
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com >
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com >
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com >
Co-authored-by: Chunan Zeng <zcnrex@gmail.com >
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai >
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai >
Co-authored-by: Zijie Xia <37504505+zijiexia@users.noreply.github.com >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: zhangxiaohao <1024393531@qq.com >
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
Co-authored-by: Julien Lin <jullin@nvidia.com >
Co-authored-by: Hao Phan <htphan@nvidia.com >
Co-authored-by: Thomas Wang <1am9trash@gmail.com >
Co-authored-by: RolaoDenthu <xinyisong0111@gmail.com >
Co-authored-by: pigeonsoup <32922982+pigeonsoup@users.noreply.github.com >
Co-authored-by: HaiShaw <hixiao@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com >
Co-authored-by: Lee Nau <lee.nau@gmail.com >
Co-authored-by: HMING <126185151+Hearum@users.noreply.github.com >
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com >
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com >
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai >
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com >
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai >
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai >
Co-authored-by: Hanming Lu <hanminglu@meta.com >
Co-authored-by: Xinyi Song <xinyis10@illinois.edu >
2026-08-04 13:22:49 -07:00
YAMY and Sam Li
5fe97637df
Support DeepGEMM for standard MoE dispatch ( #33128 )
...
Co-authored-by: Sam Li <lsam@nvidia.com >
2026-08-02 21:48:13 -07:00
Mohammad Miadh Angkad
a55e1764a2
Enable GPT-OSS FlashInfer MXFP4 on SM120 ( #32668 )
2026-07-30 00:04:23 +00:00
Hert4 and Mohammad Miadh Angkad
f69af7b7ad
[Bugfix] compressed-tensors: mixed-precision checkpoints silently load unquantized ( #32736 )
...
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-07-29 09:55:33 -07:00
Gabriel Wu
faf6894093
Implement SM120 DeepSeek V4 flashinfer_mxfp4 moe runner backend + TP2 ( #30272 )
2026-07-18 03:01:06 -07:00
Sam (Kesen Li)
ec6a3163b7
[Feature] Add FP4 KV Cache Design and support SM120 GPUs ( #21601 )
2026-07-17 14:49:43 -07:00
Po-Han Huang (NVIDIA)
cfc3d0555e
Fix ModelOpt NVFP4 scalar scales for merged linears ( #29151 )
2026-07-13 16:14:20 -07:00
Liangsheng Yin
23390589f7
[misc] Remove unit test cases that fail the admission criteria (round 3) ( #30713 )
2026-07-09 19:42:51 -07:00
Spandan Tiwari
40a522203c
[Quantization][Bugfix]: Join multi-arg RuntimeError in Quark _check_scheme_supported ( #25694 )
2026-07-09 15:07:14 -07:00
Spandan Tiwari
48d98b7c68
[Quantization][bugfix] Correct E8M0 NaN-sentinel detection in e8m0_to_f32 ( #25519 )
2026-07-09 15:02:53 -07:00
xutizhou
d364cd8ead
Support DSV4 shared expert fusion for DeepEP and MegaMOE ( #27349 )
2026-07-02 23:18:25 -07:00
eb75d990f7
[Bugfix] compressed-tensors WNA16 MoE: don't assume a "Linear" config group ( #29761 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Jiminator <jimmysh341@gmail.com >
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com >
2026-07-01 18:08:09 +00:00
Jimmy Shong
e745b3af22
[Fix] compressed-tensors block FP8: requantize weight scales to UE8M0 for DeepGEMM on Blackwell ( #28662 )
2026-06-26 21:41:18 +00:00
Mohammad Miadh Angkad
212c30d008
[MoE Refactor] Centralize FlashInfer CUTLASS MoE runner ( #28211 )
2026-06-25 13:40:33 -07:00
Trevor Morris
20f4272109
fix: Fix DSR1 perf regression due to unnecessarily falling back to triton gemm ( #28073 )
2026-06-15 09:45:09 -04:00
Prajj and prajjwal1
441b75ee69
[quantization] NVFP4 MoE: split fused w13 gate/up global scales ( #27588 )
...
Co-authored-by: prajjwal1 <prajjwal1@protonmail.com >
2026-06-14 21:18:36 -07:00
Yuan Luo and luoyuan.luo
0d9a2a9de3
[MoE Refactor] Migrate SM90 Cutlass W4A16 to MoeRunner ( #26489 )
...
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com >
2026-05-30 02:02:56 -07:00