 Jacob0226andClaude Opus 4.6
|
7e4e1dcd7a
|
[AMD] Fuse RMSNorm + FP8 per-token quant for GLM-4.7-FP8 (#21403)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 22:45:31 -07:00 |
|
Khoa Pham
|
aeeff58cd4
|
[Spec][Ngram] Clean up unused stateless batchMatch (#22487)
|
2026-04-10 21:52:56 -07:00 |
|
 Khoa PhamandClaude Opus 4.6
|
04bd8e1218
|
[Spec][Ngram] Return token counts in list_external_corpora API (#22471)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:50:02 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
3c46ff2ac5
|
fix: restore CPU flash_attn test to use sgl_kernel directly (#22573)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:39:20 -07:00 |
|
Zhangheng
|
f2af00d05a
|
[HiSparse-pd] Add device-buffer budget and fix logical pool admission in decode side (#22453)
|
2026-04-11 12:30:38 +08:00 |
|
 Alex NailsandClaude Opus 4.6
|
8eac618a8d
|
[tokenizer] lazy text accumulation + use deltas directly for streaming (#22548)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:26:04 -07:00 |
|
Liangsheng Yin
|
c7f93a2ce7
|
[metrics] Add PoolStats.update_scheduler_stats to deduplicate metrics assignment (#22559)
|
2026-04-10 21:04:18 -07:00 |
|
Bi Xue
|
d30b3efa84
|
[sgl] _ATTN_TP and _ATTN_CP use message queue for broadcast on CPU (#22205)
|
2026-04-10 20:52:49 -07:00 |
|
Xinyuan Tong
|
7c6db40540
|
Fix tool call constrained decoding and parsing for models with native formats (#21593)
|
2026-04-10 20:37:23 -07:00 |
|
Liangsheng Yin
|
c2821dfbe9
|
[mem] Introduce PoolStats dataclass; unify pool metrics and token_usage (#22554)
|
2026-04-10 20:35:50 -07:00 |
|
Liangsheng Yin
|
6cd183ff6b
|
Remove redundant test_page_size.py (#22571)
|
2026-04-10 20:35:04 -07:00 |
|
Yuhao Yang
|
16f306fd85
|
[VLM] GPU Image Preprocessing for Kimi-K2.5 (#22368)
|
2026-04-11 11:13:30 +08:00 |
|
Yilong Zhao
|
58f863956c
|
cuda graph: adjust capture time num-non-padded-tokens to align capture with replay (#22404)
|
2026-04-11 10:27:50 +08:00 |
|
Mick
|
0b4f5c9fcb
|
[diffusion] CI: improve readability and fix bug of early-return (#22507)
|
2026-04-11 10:08:44 +08:00 |
|
Qiaolin Yu
|
f41c810a2d
|
[misc] update CI_PERMISSIONS.json (#22570)
|
2026-04-10 18:58:55 -07:00 |
|
Bingxu Chen
|
213027951a
|
[AMD] Upgrade Aiter (#22264)
|
2026-04-10 18:40:43 -07:00 |
|
  
|
265696b176
|
chore: update CI test est_time values (#22565)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-04-10 18:15:02 -07:00 |
|
Cheng Wan
|
b5e4ae7b1a
|
fix: match est_time updates by backend, not just suite (#22563)
|
2026-04-10 17:54:50 -07:00 |
|
 Alison ShaoandAlison Shao
|
75223c5404
|
[Diffusion][CI] Fix nunchaku unit test broken by #22365 (#22560)
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
|
2026-04-10 17:49:56 -07:00 |
|
Cheng Wan
|
0011d2aec0
|
fix: track est_time per suite instead of per backend (#22557)
|
2026-04-10 16:58:40 -07:00 |
|
Liangsheng Yin
|
b4a1d8fd71
|
[mem] Fix idle token_usage missing mamba_usage; add FIXME for naming (#22555)
|
2026-04-10 16:20:33 -07:00 |
|
Alex Nails
|
0af9166474
|
[tokenizer] improve non streaming request processing + some small fixes. (#20310)
|
2026-04-10 15:46:12 -07:00 |
|
Sahithi Chigurupati
|
451320596f
|
[CI] Add GB200 nightly perf regression pipeline (#22461)
|
2026-04-10 15:12:24 -07:00 |
|
 Cheng WanandClaude Opus 4.6
|
3f39b3d811
|
feat: add weekly workflow to update CI test est_time values (#22545)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 15:03:37 -07:00 |
|
 oriandzhiguo.qin
|
f7a1740101
|
[MUSA][9/N] Add FA3 attention backend support through MATE (MUSA AI Tensor Engine) (#22051)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
|
2026-04-10 14:18:39 -07:00 |
|
 Minglei Zhuandzminglei
|
6af34b95b6
|
perf: precompute FA3 scheduler_metadata to eliminate per-layer prepare_varlen_num_blocks (#21104)
Co-authored-by: zminglei <zminglei@linkedin.com>
|
2026-04-10 13:57:54 -07:00 |
|
Zhongdongming Dai
|
4ace144fae
|
feat: update ModelExpress metadata API to SourceIdentity-based schema (#21222)
|
2026-04-10 13:45:05 -07:00 |
|
 satyamk7054andSatyam Kumar
|
6d8330bdb7
|
Update CI_PERMISSIONS.json (#22465)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2026-04-10 13:43:50 -07:00 |
|
 Cheng WanandClaude Opus 4.6
|
6d95602ea3
|
Reduce GPU memory for MoE parallel groups (#22515)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 13:23:23 -07:00 |
|
 satyamk7054andSatyam Kumar
|
059b287e25
|
Add offline auto-tuning for LoRA CSGMV kernel (#20391)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2026-04-10 13:10:43 -07:00 |
|
 Qiaolin Yuand0xNullPath
|
d8831355a3
|
Fix multi_layer_eagle_worker_v2 draft extend selection, add chain style multi layer mtp test (#22340)
Co-authored-by: 0xNullPath <luyan@nvidia.com>
|
2026-04-10 12:44:52 -07:00 |
|
Trevor Morris
|
7dbd0dd9f0
|
MiniMax-M2.5 - Support dp attention, dp reduce scatter, FP4 all gather, AR fusion in prepare_attn (#20067)
|
2026-04-10 12:41:27 -07:00 |
|
KrishnanPrash
|
a937ec31be
|
fix: server crash when stop_token_ids contains null (#22175)
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
|
2026-04-10 11:42:23 -07:00 |
|
 Jia GuoandClaude Opus 4.6
|
5cb4ea1d4d
|
perf: enable inductor combo_kernels for horizontal fusion (#21977)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-11 01:01:14 +08:00 |
|
Tarushii Goel
|
2ba94136ce
|
[sgl] improve mamba_track_indices perf in specdec (#22380)
|
2026-04-11 00:39:53 +08:00 |
|
Bi Xue
|
f652135d52
|
[sgl] fix using symmetric memory issues for attention_tp (#22286)
|
2026-04-11 00:26:18 +08:00 |
|
Ratish P
|
8227187d47
|
[SKILL]: add component accuracy guidance to the diffusion add-model skill (#22460)
|
2026-04-10 23:08:31 +08:00 |
|
Ratish P
|
cf5ad12612
|
[diffusion][CI]: route multimodal component accuracy through run_suite (#21960)
|
2026-04-10 23:06:03 +08:00 |
|
 kingkingleeljjandClaude Opus 4.6
|
84194c25c1
|
[BugFix] fix the bug of minimax_m2.5 model that causes repeated outputs when using tp16 (#20967)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 22:21:19 +08:00 |
|
Xiaoyu Zhang
|
1ff51555f2
|
[Diffusion] modelopt diffusion fp8 support for flux1/flux2 and wan2.2 (#22365)
|
2026-04-10 20:56:57 +08:00 |
|
Yujun Dong
|
8ba9646044
|
Make GDN support non-continuous B/A Tensor input to fix the accuracy regression of Qwen3.5-27B (#22312)
Signed-off-by: cs-cat <118669451+cs-cat@users.noreply.github.com>
|
2026-04-10 18:58:13 +08:00 |
|
Jincong Chen
|
0668a7f51a
|
[Perf] Remove two operations in gdn_backend extend verify path (#22444)
|
2026-04-10 17:53:57 +08:00 |
|
Shangming Cai
|
1c76f322df
|
[HiCache] Add CP support for HiCache (#20977)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-10 17:52:51 +08:00 |
|
 Cheng WanandClaude Opus 4.6
|
37107bee6f
|
[Observability] Add pending token count to prefill log and get_load (#22480)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 02:05:21 -07:00 |
|
Lee Nau
|
c554dc5c64
|
Add dedicated FlashInferCuteDslMoE layer for standard-path FP4 MoE (#21339)
|
2026-04-10 01:35:56 -07:00 |
|
Mick
|
7c6b5c095c
|
[diffusion] fix: fix flux2 i2i accuracy (#22423)
|
2026-04-10 16:16:51 +08:00 |
|
Liangsheng Yin
|
6cf7f210bf
|
Add page_size to admission token budget check (#22495)
|
2026-04-10 01:16:04 -07:00 |
|
 Jacob0226andClaude Opus 4.6
|
dd41764487
|
[AMD][HIP] NSA: bf16 passthrough from RMSNorm to eliminate FP8 dequantization (#22258)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 01:08:32 -07:00 |
|
Yuhao Yang
|
f5fd5ab622
|
add whisper test (#22302)
|
2026-04-10 15:34:53 +08:00 |
|
 jianan-guandMa Mingfei
|
2ab141547d
|
[CPU] Add apply_routed_scaling_factor_on_output support for biased_grouped_topk fusion (#22413)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-10 15:16:05 +08:00 |
|