 Jacob0226andClaude Opus 4.6
|
dd41764487
|
[AMD][HIP] NSA: bf16 passthrough from RMSNorm to eliminate FP8 dequantization (#22258)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-10 01:08:32 -07:00 |
|
Yuhao Yang
|
f5fd5ab622
|
add whisper test (#22302)
|
2026-04-10 15:34:53 +08:00 |
|
 jianan-guandMa Mingfei
|
2ab141547d
|
[CPU] Add apply_routed_scaling_factor_on_output support for biased_grouped_topk fusion (#22413)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-10 15:16:05 +08:00 |
|
Polisetty V R K Jyothendra Varma
|
599cce4d82
|
[Intel GPU] import flash_attn functions from sgl_kernel only (#22438)
|
2026-04-10 15:10:00 +08:00 |
|
 
|
5ba7d4e523
|
[HiSparse]: Update HiSparse's user-guide (#22499)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-04-10 15:06:43 +08:00 |
|
 
|
18f41ac427
|
[Reland] DeepSeek-R1-0528-w4a8: DeepEP Low Latency Dispatch Adopts FP8 Communication (#22316)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
|
2026-04-10 14:56:05 +08:00 |
|
Tarushii Goel
|
0334d4b7e8
|
[sgl] Fix mamba tracking calculation in spec dec (#22239)
|
2026-04-10 14:46:16 +08:00 |
|
Ethan (Yusheng) Su
|
6d79c60995
|
[Lora] Lora kimi support (#22381)
|
2026-04-09 22:31:53 -07:00 |
|
Liangsheng Yin
|
722e25a621
|
Fix SWA eviction boundary and page-align chunked prefill (#22470)
|
2026-04-09 22:09:43 -07:00 |
|
Ke Bao
|
e77bfba24d
|
Fix NCCL AllGather hanging issue for Qwen3 Next MTP (#22458)
|
2026-04-10 11:40:54 +08:00 |
|
 Alison ShaoandAlison Shao
|
b853e2c41b
|
[CI] Remove Slack notification from ci-auto-bisect workflow (#22483)
Co-authored-by: Alison Shao <alison.shao@Mac.lan>
|
2026-04-09 20:32:09 -07:00 |
|
 
|
45b0182205
|
[CI] Update est_time for 64 tests based on actual elapsed times (#22305)
Co-authored-by: Alison Shao <alison.shao@Mac.lan>
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
|
2026-04-09 20:31:37 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
89553ff82b
|
[Observability] Add Prometheus metrics endpoint for gRPC mode (#20801)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-09 20:04:54 -07:00 |
|
LHXuuu
|
42ffb168b3
|
[EPD][VLM] Support Kimi K25 EPD (#22269)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
|
2026-04-10 10:58:42 +08:00 |
|
 Yibo CaiandMa Mingfei
|
4644d28213
|
[sgl-kernel/cpu] fix build error on non-x86 platform (#22245)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-10 09:58:07 +08:00 |
|
ishandhanani
|
3aaaf53f59
|
[Docker] Fix CI docker target after Dockerfile restructure (#22478)
|
2026-04-09 18:53:42 -07:00 |
|
jacky.cheng
|
d283808457
|
[AMD] Replace triton rotary_emb with aiter rotary_emb for Wan2.2 denoise (#22422)
|
2026-04-09 18:21:02 -07:00 |
|
Shu Wang
|
5638d40f3a
|
[nvidia] Gemma4 nvfp4 fix (#22079)
|
2026-04-10 08:44:34 +08:00 |
|
Tarushii Goel
|
cebd9c2a1e
|
[sgl] add ability to return logprobs in MultiLayerEagleWorkerV2 (#22241)
|
2026-04-09 16:20:55 -07:00 |
|
ishandhanani
|
aa103eab8d
|
[Docker] Optimize Dockerfile for BuildKit layer caching (#22160)
|
2026-04-09 15:34:57 -07:00 |
|
 Mohammad Miadh AngkadandDavid Wang
|
c3833ba929
|
Enable DFLASH support for additional model backends (#22358)
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
|
2026-04-09 14:36:12 -07:00 |
|
Ethan (Yusheng) Su
|
28ef6de091
|
[Lora] Lora quat info re-factor and support deepseekv3 mla lora (#22323)
|
2026-04-09 14:19:58 -07:00 |
|
Baizhou Zhang
|
60acdc31f2
|
[Fix] Fix several bugs on DSA models (#22430)
|
2026-04-09 12:46:23 -07:00 |
|
Baizhou Zhang
|
606aa11ea8
|
[DSA] Enable all reduce fusion for DSA models (#22390)
|
2026-04-09 12:42:44 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
9d905efa2c
|
[Docker] Fix Trivy CVEs, cubin download 403s, and kernels command order (#22322)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-09 12:26:22 -07:00 |
|
 Lawrence WuandKangyan-Zhou
|
8eb235ab51
|
fix: do not strip whitespace from GLM tool call values (#20543)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-09 11:14:15 -07:00 |
|
Lishan H
|
8b991d98a1
|
[feature] asr: add chunk-based streaming ASR for Qwen3-ASR (#22089)
|
2026-04-10 01:49:03 +08:00 |
|
Ke Bao
|
021bd77c91
|
Add skills for debugging hanging issues (#22463)
|
2026-04-10 01:37:44 +08:00 |
|
billishyahao
|
1df9f4e2f6
|
[AMD] Add prealloc token env for mori-ep (#22329)
|
2026-04-09 09:34:35 -07:00 |
|
Jonah Bernard
|
8216b921a1
|
Add MLX profiling to bench_one_batch.py (#22159)
|
2026-04-09 20:45:21 +08:00 |
|
Xiaoyu Zhang
|
7603b226ce
|
Upgrade sglang-torch-profiler-analysis SKILLS (#22440)
|
2026-04-09 18:23:03 +08:00 |
|
Liangsheng Yin
|
9fed58805f
|
[Doc] Clarify SWA HybridSWAPoolConfigurator comments on all-SWA vs hybrid semantics (#22443)
|
2026-04-09 03:02:16 -07:00 |
|
YMbmzy
|
8a67fb20ea
|
[Speculative] Support penalty for spec v2 overlap scheduling (#22049)
|
2026-04-09 01:59:04 -07:00 |
|
Thomas Wang
|
628df31d08
|
[AMD] Use aiter CK layernorm2d for LayerNorm to reduce NSA indexer kernel launches (#22424)
|
2026-04-09 01:55:29 -07:00 |
|
 
|
57ffc55fb6
|
feat: [1/2] [DeepEP] Fuse shared expert into MoE dispatch under EP (#20089)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: AichenF <aichenf@nvidia.com>
|
2026-04-09 01:48:28 -07:00 |
|
amote-i
|
7965573eb4
|
fix issues for npu docs (#22307)
|
2026-04-09 16:27:34 +08:00 |
|
Liwansi
|
8ec0934f8f
|
[NPU]add Qwen3-32b and Qwen3-8b low latency md (#22429)
|
2026-04-09 16:18:34 +08:00 |
|
 
|
19bbaeb3ee
|
[HiSparse]: Add HiSpares-DSA Model's nightly CI (#22425)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-04-09 01:00:55 -07:00 |
|
Sundara Raman Ramachandran
|
a64905a7b8
|
[CICD] [prefill-only] Consolidate prefill-only model E2E tests (#22405)
|
2026-04-09 00:54:34 -07:00 |
|
Mick
|
9709192ce9
|
[diffusion] feat: support FLUX.2-small-decoder (#22414)
|
2026-04-09 15:53:14 +08:00 |
|
 Liangsheng YinandKe Bao
|
8ff01d6841
|
[Test] Add CPU unit tests for MemoryPoolConfigurator (#22420)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-04-09 00:39:19 -07:00 |
|
Liangsheng Yin
|
de441ac6bb
|
[core] Introduce MemoryPoolConfigurator class hierarchy (#22389)
|
2026-04-09 15:29:19 +08:00 |
|
Evgueni Petrov
|
b9c316917b
|
fix AttributeError: 'LazyValue' object has no attribute 'keys' in eplb_manager.py for qwen3 moe (#21822)
|
2026-04-09 00:13:29 -07:00 |
|
Michael
|
ef6bfc1197
|
[AMD] Add GLM-5.1-FP8 nightly accuracy and performance benchmarks for MI30x and MI35x (#22336)
|
2026-04-08 22:57:43 -07:00 |
|
Nicolas Castet
|
e379befbac
|
Add symmetric debug mode to print stack trace of comm ops with unregistered tensors (#18569)
|
2026-04-08 22:34:58 -07:00 |
|
Bingxu Chen
|
6b96f8341d
|
[AMD] Fix multimodal diffusion test crash on ROCm by falling back to SDPA (#22335)
|
2026-04-08 22:32:49 -07:00 |
|
Xiaoyu Zhang
|
30b738d3a6
|
[SKILL] add torch profiler analysis workflow (#22353)
|
2026-04-09 12:53:48 +08:00 |
|
Liangsheng Yin
|
edfddda192
|
Move runai model loader test to nightly suite (#22418)
|
2026-04-08 21:39:32 -07:00 |
|
Mick
|
355fcbcc17
|
[diffusion] fix: fix cache dit refresh none mask (#22374)
|
2026-04-09 11:58:24 +08:00 |
|
 jsheng_LinkedinandClaude Opus 4.6
|
6838a23226
|
[Feature] Add token embedding overrides for sparse embedding replacement (#20960)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-08 20:51:36 -07:00 |
|