Sundara Raman Ramachandran
|
4927975427
|
[Score API] Add return_pooled_hidden_states to Scoring API for SequenceClassification / RewardModel (#22427)
|
2026-04-15 14:58:56 -07:00 |
|
Lianmin Zheng
|
adb310b976
|
Cleanup server_args.py and minor code tidying (#22820)
|
2026-04-14 18:52:41 -07:00 |
|
 Colin ZandHAI
|
b10f852118
|
GLM-5/5.1 MXFP4 Checkpoint Inference Compatibility Fix (#22543)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-04-13 23:56:48 -07:00 |
|
yuki-brook
|
1ec018f27a
|
[Feature] Add SiMM as sglang HiCache Storage backend (#18016)
|
2026-04-13 17:12:37 -07:00 |
|
 Kurt ShusterandYusheng Su
|
ff13dfee45
|
[lora][moe] Virtual experts for LoRA MoE (#22122)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
|
2026-04-13 21:19:30 +00:00 |
|
Mohammad Miadh Angkad
|
4dbd59850b
|
Add bfloat16 KV cache validation for HiSparse (#22505)
|
2026-04-13 12:41:42 +08:00 |
|
Ziang Li
|
5593539942
|
[RL] Refactor NVFP4 shuffling/swizzling to in-place replacement (#22204)
|
2026-04-12 19:08:45 -07:00 |
|
 Kurt ShusterandYusheng Su
|
f81b6df3a3
|
[lora] Fix partial MoE rank loading, VL lm_head, strict loading, deepseek on-demand (#21864)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
|
2026-04-12 16:25:02 -07:00 |
|
Mohammad Miadh Angkad
|
bcc0c65aa8
|
[DSA] Hopper FP8 FlashMLA KV padding (#22372)
|
2026-04-12 02:19:17 -07:00 |
|
Kurt Shuster
|
0e0091c6c8
|
[server] Add --quantization unquant to explicitly opt out of quantization (#21863)
|
2026-04-12 02:17:22 -07:00 |
|
Wenyao Gao
|
4dfc8e1c3f
|
VLM: support passing --mm-process-config for all models (#18467)
|
2026-04-12 17:08:05 +08:00 |
|
  
|
f855a0bde6
|
Introduce CUDA graph debug mode with breakable CUDA graph (#19102)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-11 00:36:56 -07:00 |
|
 oriandzhiguo.qin
|
f7a1740101
|
[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine) (#22051)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
|
2026-04-10 14:18:39 -07:00 |
|
Lee Nau
|
c554dc5c64
|
Add dedicated FlashInferCuteDslMoE layer for standard-path FP4 MoE (#21339)
|
2026-04-10 01:35:56 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
89553ff82b
|
[Observability] Add Prometheus metrics endpoint for gRPC mode (#20801)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-09 20:04:54 -07:00 |
|
LHXuuu
|
42ffb168b3
|
[EPD][VLM] Support Kimi K25 EPD (#22269)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
|
2026-04-10 10:58:42 +08:00 |
|
Baizhou Zhang
|
60acdc31f2
|
[Fix] Fix several bugs on DSA models (#22430)
|
2026-04-09 12:46:23 -07:00 |
|
Baizhou Zhang
|
606aa11ea8
|
[DSA] Enable all reduce fusion for DSA models (#22390)
|
2026-04-09 12:42:44 -07:00 |
|
billishyahao
|
1df9f4e2f6
|
[AMD] Add prealloc token env for mori-ep (#22329)
|
2026-04-09 09:34:35 -07:00 |
|
 
|
57ffc55fb6
|
feat: [1/2] [DeepEP] Fuse shared expert into MoE dispatch under EP (#20089)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: AichenF <aichenf@nvidia.com>
|
2026-04-09 01:48:28 -07:00 |
|
YAMY
|
c26b8b4a4b
|
[GDN] Remove FlashInfer GDN decode + no_buffer guard and default to FlashInfer on SM100+ (#21861)
|
2026-04-08 11:59:15 -07:00 |
|
    
|
f08726fd56
|
[Feature] Add DFLASH speculative decoding support (#22077)
Co-authored-by: Jian Chen <141193260+jianc99@users.noreply.github.com>
Co-authored-by: Zhijian Liu <5782437+zhijian-liu@users.noreply.github.com>
Co-authored-by: Richard Gong <8001209+gongy@users.noreply.github.com>
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com>
Co-authored-by: xm:D <38322020+xiaomin-d@users.noreply.github.com>
|
2026-04-07 14:48:51 -07:00 |
|
Ke Bao
|
be42fbbbd7
|
Support HTTP2 server (#21700)
|
2026-04-08 00:42:52 +08:00 |
|
shuwenn
|
ec5742f4ab
|
fix: Auto-correct page_size for Mamba no_buffer radix cache mode (#20538)
|
2026-04-08 00:19:31 +08:00 |
|
Xingyu Liu
|
98f38b14df
|
Add registration API for external linear attention backend (#21983)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
|
2026-04-07 02:47:40 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)    
|
2813cb6d9a
|
[New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
|
2026-04-06 20:24:44 -07:00 |
|
Liangsheng Yin
|
e4b1366a46
|
[Spec][Ngram] Support multiple SAMs with dynamic HTTP API (#22203)
|
2026-04-06 18:49:22 -07:00 |
|
Ratish P
|
7f2fcc0b08
|
[VLM]: allow Qwen3.5 models for encoder disaggregation (#21849)
|
2026-04-07 02:07:24 +08:00 |
|
Khoa Pham
|
12272b6791
|
[Spec][Ngram] 6/N: Load an external corpus and construct a Suffix Automaton (#21425)
|
2026-04-06 00:11:14 -07:00 |
|
Baizhou Zhang
|
bf984ae65d
|
Revert "[Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang" (#22098)
|
2026-04-04 02:17:19 -07:00 |
|
Ethan (Yusheng) Su
|
ff8e47edf9
|
[5/n] Lora support cuda graph (#21647)
|
2026-04-04 00:31:46 -07:00 |
|
faceless void
|
de9859073f
|
Add --stream-response-default-include-usage server flag (#16711)
|
2026-04-03 21:36:00 -07:00 |
|
 Mohammad Miadh AngkadandBaizhou Zhang
|
8cb337c8ea
|
[Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang (#21906)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-04-03 14:19:13 -07:00 |
|
Zhangheng
|
ed3435e37f
|
[HiSparse]: Optimize server args checking-HiSparse is temporarily only available for DSA models. (#22065)
|
2026-04-04 02:23:56 +08:00 |
|
 
|
56ac9c9932
|
[Fix] Add _MOE_TP to graph_capture for MoE models with ep>1 (#21907)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-04-03 02:33:16 -07:00 |
|
Ricardo-M-L
|
24f52e66d3
|
fix: remove duplicate words in comments (#22007)
|
2026-04-03 00:05:39 -07:00 |
|
Baizhou Zhang
|
efa7b2d5d3
|
Revert "[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine)" (#22002)
|
2026-04-02 20:42:13 -07:00 |
|
 oriandR0CKSTAR
|
939cf398a9
|
[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine) (#17985)
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
|
2026-04-02 15:04:31 -07:00 |
|
Baizhou Zhang
|
fbc1f92453
|
[DSA] Set trtllm kernels as nsa default for Blackwell (#21914)
|
2026-04-02 00:22:27 -07:00 |
|
Khoa Pham
|
f836658077
|
[Spec][Ngram] 4/N: Remove max_match_window_size and min_match_window_size, matching all suffixes of the Trie (#21225)
|
2026-04-01 22:09:46 -07:00 |
|
Noa Neria
|
8d9145d97e
|
Direct model loading from object storage with Runai Model Streamer (#17948)
Signed-off-by: Noa Neria <noa@run.ai>
|
2026-04-01 18:41:22 -07:00 |
|
YAMY
|
821a8a99fb
|
[Disagg] GPU staging buffer with dynamic ring allocator for heterogeneous TP KV transfer (#19890)
|
2026-04-01 14:09:18 -07:00 |
|
Baizhou Zhang
|
5e12c4e08e
|
[DSA] Support trtllm sparse mla kernel for prefill batches (#21783)
|
2026-04-01 13:55:05 -07:00 |
|
Yuhao Yang
|
1aabe44b64
|
[VLM] remove AsyncMMDataProcessor wrapper (#21651)
|
2026-04-01 17:39:50 +08:00 |
|
Zhiqiang Xie
|
9eb75211b1
|
style refinement for hisparse (#21198)
|
2026-04-01 01:03:17 -07:00 |
|
Brayden Zhong
|
6a9b09847c
|
CUTLASS NVFP4 GEMM improvement of SM120 (#21314)
|
2026-04-01 09:04:34 +08:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
f60f2ccc10
|
[Fix] Fall back to triton MOE for GPT-OSS on Blackwell with driver >= 595 (#21780)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-31 15:52:10 -07:00 |
|
weireweire
|
9191b02eda
|
Fix cuda graph max bs capture upper bound (#21005)
|
2026-03-31 15:20:56 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
3c91ebdf55
|
[2/n] lora - Shared outer experts and support qwen3_30b_a3b_instruct (#21466)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 14:06:23 -07:00 |
|
JD
|
20d07c4384
|
Fix remote weight info nnode>1 and dp>1 (#17389)
|
2026-03-31 21:17:18 +08:00 |
|