Commit Graph
1008 Commits
Author SHA1 Message Date
Sundara Raman Ramachandran 4927975427 [Score API] Add return_pooled_hidden_states to Scoring API for SequenceClassification / RewardModel (#22427) 2026-04-15 14:58:56 -07:00
Lianmin Zheng adb310b976 Cleanup server_args.py and minor code tidying (#22820) 2026-04-14 18:52:41 -07:00
Colin ZandHAI b10f852118 GLM-5/5.1 MXFP4 Checkpoint Inference Compatibility Fix (#22543)
Co-authored-by: HAI <hixiao@gmail.com>
2026-04-13 23:56:48 -07:00
yuki-brook 1ec018f27a [Feature] Add SiMM as sglang HiCache Storage backend (#18016) 2026-04-13 17:12:37 -07:00
Kurt ShusterandYusheng Su ff13dfee45 [lora][moe] Virtual experts for LoRA MoE (#22122)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
2026-04-13 21:19:30 +00:00
Mohammad Miadh Angkad 4dbd59850b Add bfloat16 KV cache validation for HiSparse (#22505) 2026-04-13 12:41:42 +08:00
Ziang Li 5593539942 [RL] Refactor NVFP4 shuffling/swizzling to in-place replacement (#22204) 2026-04-12 19:08:45 -07:00
Kurt ShusterandYusheng Su f81b6df3a3 [lora] Fix partial MoE rank loading, VL lm_head, strict loading, deepseek on-demand (#21864)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
2026-04-12 16:25:02 -07:00
Mohammad Miadh Angkad bcc0c65aa8 [DSA] Hopper FP8 FlashMLA KV padding (#22372) 2026-04-12 02:19:17 -07:00
Kurt Shuster 0e0091c6c8 [server] Add --quantization unquant to explicitly opt out of quantization (#21863) 2026-04-12 02:17:22 -07:00
Wenyao Gao 4dfc8e1c3f VLM: support passing --mm-process-config for all models (#18467) 2026-04-12 17:08:05 +08:00
f855a0bde6 Introduce CUDA graph debug mode with breakable CUDA graph (#19102)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-11 00:36:56 -07:00
oriandzhiguo.qin f7a1740101 [MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine) (#22051)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
2026-04-10 14:18:39 -07:00
Lee Nau c554dc5c64 Add dedicated FlashInferCuteDslMoE layer for standard-path FP4 MoE (#21339) 2026-04-10 01:35:56 -07:00
Kangyan-ZhouandClaude Opus 4.6 89553ff82b [Observability] Add Prometheus metrics endpoint for gRPC mode (#20801)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 20:04:54 -07:00
LHXuuu 42ffb168b3 [EPD][VLM] Support Kimi K25 EPD (#22269)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
2026-04-10 10:58:42 +08:00
Baizhou Zhang 60acdc31f2 [Fix] Fix several bugs on DSA models (#22430) 2026-04-09 12:46:23 -07:00
Baizhou Zhang 606aa11ea8 [DSA] Enable all reduce fusion for DSA models (#22390) 2026-04-09 12:42:44 -07:00
billishyahao 1df9f4e2f6 [AMD] Add prealloc token env for mori-ep (#22329) 2026-04-09 09:34:35 -07:00
57ffc55fb6 feat: [1/2] [DeepEP] Fuse shared expert into MoE dispatch under EP (#20089)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: AichenF <aichenf@nvidia.com>
2026-04-09 01:48:28 -07:00
YAMY c26b8b4a4b [GDN] Remove FlashInfer GDN decode + no_buffer guard and default to FlashInfer on SM100+ (#21861) 2026-04-08 11:59:15 -07:00
f08726fd56 [Feature] Add DFLASH speculative decoding support (#22077)
Co-authored-by: Jian Chen <141193260+jianc99@users.noreply.github.com>
Co-authored-by: Zhijian Liu <5782437+zhijian-liu@users.noreply.github.com>
Co-authored-by: Richard Gong <8001209+gongy@users.noreply.github.com>
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com>
Co-authored-by: xm:D <38322020+xiaomin-d@users.noreply.github.com>
2026-04-07 14:48:51 -07:00
Ke Bao be42fbbbd7 Support HTTP2 server (#21700) 2026-04-08 00:42:52 +08:00
shuwenn ec5742f4ab fix: Auto-correct page_size for Mamba no_buffer radix cache mode (#20538) 2026-04-08 00:19:31 +08:00
Xingyu Liu 98f38b14df Add registration API for external linear attention backend (#21983)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
2026-04-07 02:47:40 -07:00
2813cb6d9a [New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
2026-04-06 20:24:44 -07:00
Liangsheng Yin e4b1366a46 [Spec][Ngram] Support multiple SAMs with dynamic HTTP API (#22203) 2026-04-06 18:49:22 -07:00
Ratish P 7f2fcc0b08 [VLM]: allow Qwen3.5 models for encoder disaggregation (#21849) 2026-04-07 02:07:24 +08:00
Khoa Pham 12272b6791 [Spec][Ngram] 6/N: Load an external corpus and construct a Suffix Automaton (#21425) 2026-04-06 00:11:14 -07:00
Baizhou Zhang bf984ae65d Revert "[Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang" (#22098) 2026-04-04 02:17:19 -07:00
Ethan (Yusheng) Su ff8e47edf9 [5/n] Lora support cuda graph (#21647) 2026-04-04 00:31:46 -07:00
faceless void de9859073f Add --stream-response-default-include-usage server flag (#16711) 2026-04-03 21:36:00 -07:00
Mohammad Miadh AngkadandBaizhou Zhang 8cb337c8ea [Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang (#21906)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-03 14:19:13 -07:00
Zhangheng ed3435e37f [HiSparse]: Optimize server args checking-HiSparse is temporarily only available for DSA models. (#22065) 2026-04-04 02:23:56 +08:00
56ac9c9932 [Fix] Add _MOE_TP to graph_capture for MoE models with ep>1 (#21907)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-03 02:33:16 -07:00
Ricardo-M-L 24f52e66d3 fix: remove duplicate words in comments (#22007) 2026-04-03 00:05:39 -07:00
Baizhou Zhang efa7b2d5d3 Revert "[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine)" (#22002) 2026-04-02 20:42:13 -07:00
oriandR0CKSTAR 939cf398a9 [MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine) (#17985)
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
2026-04-02 15:04:31 -07:00
Baizhou Zhang fbc1f92453 [DSA] Set trtllm kernels as nsa default for Blackwell (#21914) 2026-04-02 00:22:27 -07:00
Khoa Pham f836658077 [Spec][Ngram] 4/N: Remove max_match_window_size and min_match_window_size, matching all suffixes of the Trie (#21225) 2026-04-01 22:09:46 -07:00
Noa Neria 8d9145d97e Direct model loading from object storage with Runai Model Streamer (#17948)
Signed-off-by: Noa Neria <noa@run.ai>
2026-04-01 18:41:22 -07:00
YAMY 821a8a99fb [Disagg] GPU staging buffer with dynamic ring allocator for heterogeneous TP KV transfer (#19890) 2026-04-01 14:09:18 -07:00
Baizhou Zhang 5e12c4e08e [DSA] Support trtllm sparse mla kernel for prefill batches (#21783) 2026-04-01 13:55:05 -07:00
Yuhao Yang 1aabe44b64 [VLM] remove AsyncMMDataProcessor wrapper (#21651) 2026-04-01 17:39:50 +08:00
Zhiqiang Xie 9eb75211b1 style refinement for hisparse (#21198) 2026-04-01 01:03:17 -07:00
Brayden Zhong 6a9b09847c CUTLASS NVFP4 GEMM improvement of SM120 (#21314) 2026-04-01 09:04:34 +08:00
Baizhou ZhangandClaude Opus 4.6 f60f2ccc10 [Fix] Fall back to triton MOE for GPT-OSS on Blackwell with driver >= 595 (#21780)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 15:52:10 -07:00
weireweire 9191b02eda Fix cuda graph max bs capture upper bound (#21005) 2026-03-31 15:20:56 -07:00
Ethan (Yusheng) SuandBaizhou Zhang 3c91ebdf55 [2/n] lora - Shared outer experts and support qwen3_30b_a3b_instruct (#21466)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-31 14:06:23 -07:00
JD 20d07c4384 Fix remote weight info nnode>1 and dp>1 (#17389) 2026-03-31 21:17:18 +08:00