  
|
f855a0bde6
|
Introduce CUDA graph debug mode with breakable CUDA graph (#19102)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-11 00:36:56 -07:00 |
|
 YC Yen-Ching TsengandHAI
|
3ce72252de
|
[AMD] Fix Timeout: stage-b-test-2-gpu-large-amd,stage-b-test-1-gpu-large-amd (#22228)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-04-10 22:55:44 -07:00 |
|
 Khoa PhamandClaude Opus 4.6
|
04bd8e1218
|
[Spec][Ngram] Return token counts in list_external_corpora API (#22471)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:50:02 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
3c46ff2ac5
|
fix: restore CPU flash_attn test to use sgl_kernel directly (#22573)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:39:20 -07:00 |
|
 Alex NailsandClaude Opus 4.6
|
8eac618a8d
|
[tokenizer] lazy text accumulation + use deltas directly for streaming (#22548)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:26:04 -07:00 |
|
Xinyuan Tong
|
7c6db40540
|
Fix tool call constrained decoding and parsing for models with native formats (#21593)
|
2026-04-10 20:37:23 -07:00 |
|
Liangsheng Yin
|
c2821dfbe9
|
[mem] Introduce PoolStats dataclass; unify pool metrics and token_usage (#22554)
|
2026-04-10 20:35:50 -07:00 |
|
Liangsheng Yin
|
6cd183ff6b
|
Remove redundant test_page_size.py (#22571)
|
2026-04-10 20:35:04 -07:00 |
|
  
|
265696b176
|
chore: update CI test est_time values (#22565)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-04-10 18:15:02 -07:00 |
|
Alex Nails
|
0af9166474
|
[tokenizer] improve non streaming request processing + some small fixes. (#20310)
|
2026-04-10 15:46:12 -07:00 |
|
 satyamk7054andSatyam Kumar
|
059b287e25
|
Add offline auto-tuning for LoRA CSGMV kernel (#20391)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2026-04-10 13:10:43 -07:00 |
|
 Qiaolin Yuand0xNullPath
|
d8831355a3
|
Fix multi_layer_eagle_worker_v2 draft extend selection, add chain style multi layer mtp test (#22340)
Co-authored-by: 0xNullPath <luyan@nvidia.com>
|
2026-04-10 12:44:52 -07:00 |
|
Trevor Morris
|
7dbd0dd9f0
|
MiniMax-M2.5 - Support dp attention, dp reduce scatter, FP4 all gather, AR fusion in prepare_attn (#20067)
|
2026-04-10 12:41:27 -07:00 |
|
Yujun Dong
|
8ba9646044
|
Make GDN support non-continuous B/A Tensor input to fix the accuracy regression of Qwen3.5-27B (#22312)
Signed-off-by: cs-cat <118669451+cs-cat@users.noreply.github.com>
|
2026-04-10 18:58:13 +08:00 |
|
Lee Nau
|
c554dc5c64
|
Add dedicated FlashInferCuteDslMoE layer for standard-path FP4 MoE (#21339)
|
2026-04-10 01:35:56 -07:00 |
|
Yuhao Yang
|
f5fd5ab622
|
add whisper test (#22302)
|
2026-04-10 15:34:53 +08:00 |
|
 jianan-guandMa Mingfei
|
2ab141547d
|
[CPU] Add apply_routed_scaling_factor_on_output support for biased_grouped_topk fusion (#22413)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-10 15:16:05 +08:00 |
|
Ethan (Yusheng) Su
|
6d79c60995
|
[Lora] Lora kimi support (#22381)
|
2026-04-09 22:31:53 -07:00 |
|
Liangsheng Yin
|
722e25a621
|
Fix SWA eviction boundary and page-align chunked prefill (#22470)
|
2026-04-09 22:09:43 -07:00 |
|
 
|
45b0182205
|
[CI] Update est_time for 64 tests based on actual elapsed times (#22305)
Co-authored-by: Alison Shao <alison.shao@Mac.lan>
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
|
2026-04-09 20:31:37 -07:00 |
|
Ethan (Yusheng) Su
|
28ef6de091
|
[Lora] Lora quat info re-factor and support deepseekv3 mla lora (#22323)
|
2026-04-09 14:19:58 -07:00 |
|
 Lawrence WuandKangyan-Zhou
|
8eb235ab51
|
fix: do not strip whitespace from GLM tool call values (#20543)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-09 11:14:15 -07:00 |
|
YMbmzy
|
8a67fb20ea
|
[Speculative] Support penalty for spec v2 overlap scheduling (#22049)
|
2026-04-09 01:59:04 -07:00 |
|
 
|
19bbaeb3ee
|
[HiSparse]: Add HiSpares-DSA Model's nightly CI (#22425)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-04-09 01:00:55 -07:00 |
|
Sundara Raman Ramachandran
|
a64905a7b8
|
[CICD] [prefill-only] Consolidate prefill-only model E2E tests (#22405)
|
2026-04-09 00:54:34 -07:00 |
|
 Liangsheng YinandKe Bao
|
8ff01d6841
|
[Test] Add CPU unit tests for MemoryPoolConfigurator (#22420)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-04-09 00:39:19 -07:00 |
|
Michael
|
ef6bfc1197
|
[AMD] Add GLM-5.1-FP8 nightly accuracy and performance benchmarks for MI30x and MI35x (#22336)
|
2026-04-08 22:57:43 -07:00 |
|
Liangsheng Yin
|
edfddda192
|
Move runai model loader test to nightly suite (#22418)
|
2026-04-08 21:39:32 -07:00 |
|
 jsheng_LinkedinandClaude Opus 4.6
|
6838a23226
|
[Feature] Add token embedding overrides for sparse embedding replacement (#20960)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-08 20:51:36 -07:00 |
|
Khoa Pham
|
f127d67823
|
[Spec][Ngram] Misc enhance support for multiple SAMs (#22294)
|
2026-04-08 19:56:23 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
46c2b77627
|
[CI] Add GLM-5.1 nightly tests and update Qwen3.5 model (#22399)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-08 17:04:57 -07:00 |
|
Yihao Wang
|
a5ed507a16
|
[refactor] [asr] add transcription adapter for extensible ASR models support (#22181)
|
2026-04-09 01:19:37 +08:00 |
|
 Xiaoyu ZhangandMick
|
b5b2dbe05f
|
[Diffusion] Add diffusion NVFP4 scaled-mm correctness test (#22127)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-08 22:07:24 +08:00 |
|
Xiaoyu Zhang
|
ea119adc90
|
Refactor auto benchmark unit tests and fix CI bug (#22270)
|
2026-04-08 21:54:41 +08:00 |
|
 Alison ShaoandAlison Shao
|
2ad5e6df12
|
[CI] Relax gpt-oss 4GPU accuracy threshold from 0.60 to 0.58 (#22237)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-08 02:20:23 -07:00 |
|
Sundara Raman Ramachandran
|
712c8c5051
|
[Score API] Add SequenceClassification Model support (#22118)
|
2026-04-08 01:30:58 -07:00 |
|
Baizhou Zhang
|
213af1d4f7
|
Add CI tests for GLM-5 (#22285)
|
2026-04-08 01:05:36 -07:00 |
|
Michael
|
db60a620db
|
[AMD] Add GLM-5-FP8 nightly performance benchmarks for MI30x and MI35x (#21710)
|
2026-04-07 22:43:14 -07:00 |
|
 Alison ShaoandAlison Shao
|
36f05810c9
|
[CI] Move manual-only nightly tests out of test/registered/ (#22298)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-07 21:03:52 -07:00 |
|
 Alex NailsandClaude Opus 4.6
|
493ec91cbe
|
[CI] Fix stage-b-test-1-gpu-large (0) timeout by reordering LoRA tests and using tokenizer from cache (#22292)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 20:00:44 -07:00 |
|
 Qiaolin YuandLiangsheng Yin
|
117508dcd7
|
Switch eagle_infer_beta to EAGLE3 (#22303)
Co-authored-by: Liangsheng Yin <hnyls2002@users.noreply.github.com>
|
2026-04-07 18:43:48 -07:00 |
|
Kangyan-Zhou
|
dd73e9a62e
|
Revert "[CI] Update nightly test models for H200/B200 (#22288)" (#22297)
|
2026-04-07 17:04:06 -07:00 |
|
 
|
f6fc39569a
|
[CI] Migrate mgsm_en eval to gsm8k to remove openaipublic dependency (#21931)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-07 16:29:20 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
e6652309c4
|
[CI] Update nightly test models for H200/B200 (#22288)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 15:44:52 -07:00 |
|
    
|
f08726fd56
|
[Feature] Add DFLASH speculative decoding support (#22077)
Co-authored-by: Jian Chen <141193260+jianc99@users.noreply.github.com>
Co-authored-by: Zhijian Liu <5782437+zhijian-liu@users.noreply.github.com>
Co-authored-by: Richard Gong <8001209+gongy@users.noreply.github.com>
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com>
Co-authored-by: xm:D <38322020+xiaomin-d@users.noreply.github.com>
|
2026-04-07 14:48:51 -07:00 |
|
YC Yen-Ching Tseng
|
e14876742a
|
[AMD] Fix test_kimi_k25_mxfp4.py : stage-c-test-large-8-gpu-amd-mi35x (linux-mi35x-gpu-8, 1) (#22188)
|
2026-04-07 13:48:37 -07:00 |
|
Liangsheng Yin
|
cc35714b03
|
[tiny] migrate /get_server_info; print accept length in accuracy tests (#22282)
|
2026-04-07 13:08:35 -07:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
Ke Bao
|
be42fbbbd7
|
Support HTTP2 server (#21700)
|
2026-04-08 00:42:52 +08:00 |
|
Ke Bao
|
fae90abf6e
|
Move ring test to nightly (#22267)
|
2026-04-07 21:56:39 +08:00 |
|