 satyamk7054andSatyam Kumar
|
38dd4fbb66
|
Add overlap scheduling for embeddings code path (#14032)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2025-12-24 18:24:18 -08:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
92ddc46824
|
[Auto Sync] Update server_args.py (20251223) (#15700)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-12-24 14:55:14 -08:00 |
|
Xinyuan Tong
|
ae434f7821
|
Move limit-mm-data-per-request to make code clean (#15775)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-24 10:59:39 -08:00 |
|
fzyzcjy
|
e245cac0c5
|
Support JSON format request logging for easier parsing (#15743)
|
2025-12-24 19:52:43 +08:00 |
|
Lianmin Zheng
|
ff903a7eea
|
Simplify server args (#15704)
|
2025-12-23 22:46:12 -08:00 |
|
 Teng MaandXuchun Shang
|
d7301c89ba
|
[Feature] support fastsafetensors (#15091)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
|
2025-12-23 22:33:56 +08:00 |
|
Tianyu Guo
|
fa2966983a
|
Support PP for zmq_to_scheduler (#15312)
|
2025-12-23 17:07:55 +08:00 |
|
 Baidu-AIAKandsunhailiang
|
bc3ca30023
|
[PD] Support fake decode for PD disaggregation without prefill node (#14628)
Co-authored-by: sunhailiang <sunhailiang@baidu.com>
|
2025-12-23 12:43:33 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Liangsheng Yinandgemini-code-assist[bot]
|
3c882db3ad
|
Adjust wrong mtp meaning introduce by mimo (#15632)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-23 02:06:46 +08:00 |
|
fzyzcjy
|
e62e17442e
|
Tiny enable soft watchdog in CI for stuck without logs (#15616)
|
2025-12-22 17:01:04 +08:00 |
|
Hexq0210
|
cb30d056e3
|
add decode round robin policy (#15164)
|
2025-12-22 14:48:52 +08:00 |
|
Jincong Chen
|
350fbbf4dc
|
fix ds3.2 nsa backend prefill TBO (#14901)
|
2025-12-21 13:16:46 -08:00 |
|
 
|
bed301a5ac
|
[Feature] Enable return routed experts (#12162)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-21 15:16:43 +08:00 |
|
Xinyuan Tong
|
0a346d3bd9
|
feat: Add limit-mm-data-per-request argument to server arguments (#15418)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-20 10:44:47 -08:00 |
|
 sunxxunsandThomas Wang
|
f2d64e6782
|
[amd] Add deterministic all-reduce kernel for AMD (ROCm) (#15340)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
|
2025-12-18 23:36:03 -08:00 |
|
Qiaolin Yu
|
173940927f
|
Add customized sampler registration (#15423)
|
2025-12-18 23:10:23 -08:00 |
|
+6        
|
160a06cab2
|
[Feature] Xiaomi MiMo-V2-Flash day0 support (#15207)
Co-authored-by: 谢学扬 <xiexueyang@xiaomi.com>
Co-authored-by: tz <tangzhen3@xiaomi.com>
Co-authored-by: 李家乐 <lijiale10@xiaomi.com>
Co-authored-by: 张晨 <zhangchen50@xiaomi.com>
Co-authored-by: Shaohui Liu <liushaohui3@xiaomi.com>
Co-authored-by: 王晨 <wangchen77@xiaomi.com>
Co-authored-by: jiangzihan <jiangzihan@xiaomi.com>
Co-authored-by: xiexueyang <xyxie_wangyi@163.com>
Co-authored-by: Linghao Zhang <zhanglinghao@xiaomi.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: JoyFuture <35593546+JoyFuture@users.noreply.github.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: root <root@bj9-ml-g8h20e-k8s-slave106-20251106.alicn.idc.xiaomi.com>
|
2025-12-19 11:40:07 +08:00 |
|
Charles Chen
|
9e0ef04e5b
|
Support using different attention backend for draft decoding. (#14843)
|
2025-12-19 08:49:11 +08:00 |
|
Lianmin Zheng
|
d1f0063262
|
Clean up __init__ function of the scheduler and event loop for PD (#15298)
|
2025-12-18 01:35:14 -08:00 |
|
elvischenv
|
9970ee34e8
|
Mistral Large 3 NVFP4 TRTLLM MoE support (#15049)
|
2025-12-18 11:11:42 +08:00 |
|
elvischenv
|
435d1c83c1
|
[Perf] Enable Flashinfer autotune by default (#14357)
|
2025-12-16 23:01:39 -08:00 |
|
Lianmin Zheng
|
9d64a7b24f
|
Minor style fixes to the scheduler.py (#15218)
|
2025-12-16 17:09:44 -08:00 |
|
Even Zhou
|
71cb90378b
|
[NPU] fix for NPU memory settings logic (#15258)
|
2025-12-16 17:04:22 -08:00 |
|
amysaq2023
|
ccc8f3b266
|
support non disturbing remote instance weight loader v2 (#14997)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-12-16 14:39:56 -08:00 |
|
 Shangming Caiandybyang
|
36fcf71fff
|
[Qwen3-next] Add PD disaggregation support for mamba with extra_buffer (#15180)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
|
2025-12-16 14:36:00 +08:00 |
|
Yibo Cai
|
22587bc0b4
|
[BugFix] Fix CPU inference failure (#15231)
|
2025-12-15 21:32:41 -08:00 |
|
Liwansi
|
30da2f0598
|
[NPU][eagle3] support qwen eagle3 on NPU (#14820)
|
2025-12-16 02:25:13 +08:00 |
|
ratish
|
3d484be547
|
fix(attention): Prevent trtllm_mha auto-selection with eagle3 speculative decoding (#15127)
|
2025-12-15 16:42:26 +00:00 |
|
roikoren755
|
9003a4369d
|
Add missing assertion in NemotronH path (#15193)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2025-12-15 06:17:21 -08:00 |
|
 b8zhongandBrayden Zhong
|
1ab9b8e0a3
|
Enable TRT AllReduce Fusion by default (#14764)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2025-12-14 20:01:49 -08:00 |
|
Hanming Lu
|
e61dabf5e4
|
[Qwen3-next] support mamba radix cache for overlap scheduler (#14792)
|
2025-12-14 18:54:16 -08:00 |
|
roikoren755
|
3f0482174a
|
Fix Mamba2-based models' default attention backend (#15117)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2025-12-14 11:08:48 -08:00 |
|
    
|
9acb21ae27
|
feat: support EPD disaggregation (#12263)
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: Nicholas <45984215+liusy58@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2025-12-14 22:30:08 +08:00 |
|
fzyzcjy
|
168a31eb00
|
Support prefill max requests limitation (#14993)
|
2025-12-14 10:52:10 +08:00 |
|
Sam
|
4eda4194f2
|
[Fix] Disable trtllm moe backend for draft model for a qucik fix (#15002)
|
2025-12-12 23:58:47 -08:00 |
|
Lianmin Zheng
|
267170bf1d
|
Clean up server args and engine startup processes (#15015)
|
2025-12-12 18:46:07 -08:00 |
|
fzyzcjy
|
313f59ad80
|
Add soft watchdogs to debug soft hangs (#15023)
|
2025-12-13 10:41:35 +08:00 |
|
Ho-Ren (Jack) Chuang
|
171b442ad3
|
Add KV4-capable backend flashmla and update server args (#14989)
Signed-off-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
|
2025-12-12 11:50:27 -08:00 |
|
 Yineng Zhangandfzyzcjy
|
4b7b5af36a
|
Revert several PRs (#14958)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
|
2025-12-12 11:25:12 -08:00 |
|
Zaili Wang
|
4dabfbc827
|
[Fix] suppress remote weight loading engine w/o mooncake installed (#14937)
|
2025-12-11 23:12:26 -08:00 |
|
Sam
|
d7ed8a8c24
|
[NVIDIA] Enable TRTLLM BF16 MoE on Blackwell GPUs (#13798)
|
2025-12-11 22:56:13 -08:00 |
|
 b8zhongandBrayden Zhong
|
fe6d38d2fa
|
fix: trtllm mha attention auto-selection on sm120 (#14842)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2025-12-11 22:35:25 -08:00 |
|
    
|
c01b2ee094
|
[PP] Refactor PP to async mode (#11852)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: bluecoffee8 <jasperli2002@gmail.com>
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
|
2025-12-12 12:54:16 +08:00 |
|
Ho-Ren (Jack) Chuang
|
10146af099
|
Check KV4 compatibility with attention backends and add KV4 support to the attention_backend doc (#14467)
Signed-off-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
|
2025-12-11 19:00:53 -08:00 |
|
amysaq2023
|
70758d457e
|
support non-disturbing remote-instance-weight-loader (#13125)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-12-11 16:45:32 -08:00 |
|
Yinghai Lu
|
b05b346a13
|
[loader] enable private loader (#14620)
|
2025-12-11 11:49:46 -08:00 |
|
 Vladimir221andronnie_zheng
|
27032cecd9
|
[Ascend]Support of piecewise graph compilation for prefill on NPU (#12287)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-12-11 21:10:07 +08:00 |
|
b8zhong
|
da9b801eb7
|
fix lora target all + csgmv backend (#14796)
|
2025-12-10 11:42:51 -08:00 |
|
    
|
21028b5507
|
[RL] support weight reload for low-bit rollout (#9650)
Co-authored-by: Hecate0821 <hec4te0821@gmail.com>
Co-authored-by: eternally-z <zzywzj@gmail.com>
Co-authored-by: Wilboludriver <wilbolu@outlook.com>
Co-authored-by: Wilbolu <81792854+Wilboludriver@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2025-12-10 15:44:01 +08:00 |
|
TomerBN-Nvidia
|
b1cbfce612
|
fix server args bug (#14725)
|
2025-12-09 21:05:05 -08:00 |
|