 Xun SunandShangming Cai
|
9eb2dccbb7
|
[Elastic EP] Fix recovery lifecycle and add manual coverage (#31744)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-07-25 14:32:47 +08:00 |
|
 
|
f9c14e6bd4
|
[FEAT] Support fast engine recovery through weight cache (#27139)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-25 14:31:21 +08:00 |
|
 SovietPowerandShangming Cai
|
6a046fad09
|
[PD] Prevent decode scheduler from blocking on ZMQ sends to a stalled prefill peer (#31144)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-07-25 12:53:41 +08:00 |
|
 
|
ebcb74abd4
|
feat(hicache): Add shared memory allocator for host KV cache (#29326)
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-07-25 12:05:37 +08:00 |
|
 Alex NailsandClaude Opus 4.8
|
b83041c3cc
|
Migrate CompressedTensorsW4A4Nvfp4MoE TRT-LLM path onto MoeRunner (#32248)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-24 20:57:38 -07:00 |
|
Liangsheng Yin
|
cff20a2fbb
|
[Spec] Consolidate the grammar sync decision into ScheduleBatch.grammar_needs_sync (#32353)
|
2026-07-24 20:42:38 -07:00 |
|
Yihao Wang
|
95865de24f
|
[diffusion] CI: read consistency GT from ci-data-diffusion at per-platform commit (#32297)
|
2026-07-25 09:55:57 +08:00 |
|
 cctryandJialin Ouyang
|
a690e5e0b3
|
Add stream label to TTFT metrics (#32363)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
|
2026-07-24 17:44:04 -07:00 |
|
 cctryandYinghai Lu
|
ce705bb6dc
|
Report accelerator type in /v1/loads (#32348)
Co-authored-by: Yinghai Lu <yinghai@meta.com>
|
2026-07-24 17:31:28 -07:00 |
|
 DarkSharpnessandClaude Fable 5
|
9402012f0f
|
[Perf] Halve the non-finite sanitization overhead in per_token_group_quant (#32296)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-25 08:17:12 +08:00 |
|
 danielafrimiandDaniel Afrimi
|
7de8792758
|
Fix FP8 Triton dtype selection on A100 (#31340)
Co-authored-by: Daniel Afrimi <dafrimi@aws-dfw-cs-001-login-01.cm.cluster>
|
2026-07-24 17:03:19 -07:00 |
|
 
|
3079157175
|
Fix PyPI release: drop the git-only sgl-eval dep from packaged metadata (#32354)
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-24 20:00:57 -04:00 |
|
 Pranjal Shankhdharandpranjalssh
|
9c483cccfe
|
Support a same-size mixed q dtype in the fused RoPE kernels (#31834)
Co-authored-by: pranjalssh <pranjalssh@fb.com>
|
2026-07-24 16:19:08 -07:00 |
|
 
|
962c076934
|
Decode input_audio media containers with PyAV & Update memory profiler (#31832)
Signed-off-by: Shiyan Deng <dsy842974287@meta.com>
Signed-off-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
|
2026-07-24 16:18:45 -07:00 |
|
Zhihao Wang
|
14d6e1d3b1
|
[RL] Support FlashInfer TRT-LLM NVFP4 MoE in the RL weight checker (#31085)
Signed-off-by: zhihaow6 <zhihaow6@illinois.edu>
|
2026-07-24 16:01:54 -07:00 |
|
Zhihao Wang
|
0a212c6119
|
[RL] DSV4: add env to quantize SWA KV cache from bf16-rounded values (#31086)
Signed-off-by: zhihaow6 <zhihaow6@illinois.edu>
|
2026-07-24 15:52:00 -07:00 |
|
Zhihao Wang
|
f7986c8603
|
[RL] DSV4: dispatch indexer topk_transform_512 through DSATopKBackend (#31087)
Signed-off-by: zhihaow6 <zhihaow6@illinois.edu>
|
2026-07-24 14:45:44 -07:00 |
|
cctry
|
8727d105db
|
Add prefill and decode load counters to LoadSnapshot (#32245)
|
2026-07-24 14:04:54 -07:00 |
|
Baizhou Zhang
|
f15b43242b
|
Bump sgl-deep-gemm to 0.1.5 (#32345)
|
2026-07-24 14:03:38 -07:00 |
|
Qiaolin Yu
|
82fe0f041a
|
Fix stale flashinfer-MLA fallback poisoning spec verify capture (trtllm_mla + tc_piecewise) (#32288)
|
2026-07-24 13:24:26 -07:00 |
|
Zheng Wengang
|
be7c13af07
|
[BugFix] Fix DS/Kimi crash on non-first PP ranks when resolving input length (#31752)
|
2026-07-24 12:19:02 -07:00 |
|
 mosya415andmosya415
|
71015f3fea
|
fix(dsa): fail fast on fp8_e4m3 KV with tilelang DSA backend on CUDA (#31346)
Co-authored-by: mosya415 <263250241+mosya415@users.noreply.github.com>
|
2026-07-24 12:10:38 -07:00 |
|
Jinyan Chen
|
1e69765bae
|
Add FP4 Indexer for DeepSeek V4 on SM120 (#27059)
|
2026-07-24 11:37:23 -07:00 |
|
YAMY
|
2428f56145
|
[Bugfix] Fix Kimi-Linear state transfer across heterogeneous TP (#32262)
|
2026-07-24 10:31:17 -07:00 |
|
Zhiqiang Xie
|
5da0b6ec39
|
Write-back policy fix for unified tree (#31845)
|
2026-07-24 09:58:18 -07:00 |
|
Lu Fang
|
448662e85e
|
[mm] Accept per-item embedding lists from DataEmbeddingFunc (#31826)
|
2026-07-24 08:27:24 -07:00 |
|
khalilzhk
|
dfaf75b1a6
|
[NPU] bugfix for extra device memory on Ascend (#30112)
|
2026-07-24 21:13:21 +08:00 |
|
 
|
3d91a569ce
|
[MoE Backend] Add HPC-Ops FP8 MoE runner backend (#30541)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-24 19:46:11 +08:00 |
|
 Jun LiuandXiaoyu Zhang
|
4d5917e744
|
Add DeepSeek-reference 1e-20 epsilon to top-k renormalization to prevent 0/0 NaN (#31017)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-24 19:45:44 +08:00 |
|
 
|
8389d79e43
|
ci: add LongCat-Flash-Lite-FP8 8-GPU nightly test + fix NextN rope_theta (#32125)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-24 19:45:28 +08:00 |
|
 
|
841fa293b5
|
[Fix] Reject online weight updates while the HPC-Ops router GEMM split cache is active (#31943)
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-24 19:30:13 +08:00 |
|
Liangsheng Yin
|
a31542ebd9
|
[Feature] Add leveled invariant-check primitive for nan/inf/oob validity checks (#32308)
|
2026-07-24 04:27:20 -07:00 |
|
YAMY
|
de816e1eb5
|
[Disagg][StagingBuffer][1/2] Robustness and failure handling (#31217)
|
2026-07-24 17:22:55 +08:00 |
|
Liangsheng Yin
|
f4f15162bc
|
[Fix] Fail fast when a safetensors index references missing shard files (#32279)
|
2026-07-24 02:18:56 -07:00 |
|
Yuzhen Zhou
|
b954e9cf3d
|
[6/6][kimi-deterministic] Use deterministic seeded coins for EAGLE rejection sampling (#30822)
|
2026-07-24 02:11:21 -07:00 |
|
 Zheng Wengangandsiyu
|
364b5f23e6
|
[BugFix][EPD] Harden zmq_to_scheduler receiver failures; sync error info across TP (#31592)
Co-authored-by: siyu <liusy58@linux.alibaba.com>
|
2026-07-24 16:37:52 +08:00 |
|
 Jun LiuandXinyuan Tong
|
58f417049d
|
[PD] Fix multi-tokenizer disaggregation metrics labels (#30412)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-07-24 16:27:21 +08:00 |
|
Sam Shleifer
|
b8bb1b4e5a
|
[lora] Fix WAR race: never write MoE runner output into hidden_states in place (#31870)
|
2026-07-24 00:58:55 -07:00 |
|
Liangsheng Yin
|
d059b0f56e
|
[Fix] Clamp degenerate all-sentinel draft rows to token 0 in dspark _online_combine_kernel (#32277)
|
2026-07-24 00:53:56 -07:00 |
|
    
|
35e25f5356
|
[Feature] DCP: A2A + FlashInfer-MNNVL comm backends and q-replicate (Helix) (#21637)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-07-24 00:21:53 -07:00 |
|
Xinyuan Tong
|
39955d5314
|
[MoE] Make DeepEP auto serve flashinfer_cutedsl FP4 (coerce to low_latency) + guard (#29523)
|
2026-07-24 06:59:18 +00:00 |
|
Qiaolin Yu
|
15d73f1e03
|
Fix dynamo recompile limit in allreduce and bf16 gemm (#32239)
|
2026-07-23 22:24:01 -07:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
319055c191
|
[Intel GPU] Add XPU Platform support (#31949)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 12:49:29 +08:00 |
|
ANSHUMAN TRIPATHY
|
f0f78a6c93
|
Add deterministic inference for eagle parity test (#30026)
|
2026-07-24 12:48:21 +08:00 |
|
Liangsheng Yin
|
99b29bf188
|
[Fix] Support ENCODER_ONLY target-verify in the trtllm_mha backend (#32178)
|
2026-07-23 20:33:39 -07:00 |
|
Cheng Wan
|
eac7c7d7cd
|
fix(attention): read per-runner kv cache dtype off model_runner (#32251)
|
2026-07-23 20:08:57 -07:00 |
|
  
|
1e10ec93b3
|
[XPU] Add XPU device support for LMCache radix cache integration (#23534)
Co-authored-by: Christopher Manteuffel <christopher.manteuffel@intel.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 08:41:01 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
2f823a2eee
|
[Intel GPU] calculate free memory based on allocated memory for XPU (#32044)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 08:38:52 +08:00 |
|
Yihao Wang
|
433429b16a
|
diffusion: skip _save_gt_output for 3d/mesh (#32117)
|
2026-07-24 08:14:05 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
d4a0dfbc31
|
[Fix] Two root causes of the H100 deepep TBO CI break: scale-tensor use-after-free + missing non-finite quant sanitization (#32188)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-24 07:29:28 +08:00 |
|