fzyzcjy
|
f3e9336dcb
|
Support weight update for blackwell DeepGEMM (#13324)
|
2025-11-17 15:55:26 +08:00 |
|
fzyzcjy
|
d971f22898
|
Super tiny expose transform_scale_ue8m0 API for RL frameworks (#13323)
|
2025-11-15 17:31:04 +08:00 |
|
fzyzcjy
|
33f08a98b0
|
Tiny refactor condition to requant scale ue8m0 (#13286)
|
2025-11-15 16:36:00 +08:00 |
|
fzyzcjy
|
8e6083bfcf
|
Support inverse transform ue8m0 scale (#13285)
|
2025-11-15 16:34:32 +08:00 |
|
fzyzcjy
|
2fbc78a083
|
Support fast gemm when in batch invariant DeepGEMM fallback (#13259)
|
2025-11-15 16:34:15 +08:00 |
|
fzyzcjy
|
af9f71f9c5
|
Add script to create a model with fewer layers for debugging (#13284)
|
2025-11-14 22:13:27 +08:00 |
|
fzyzcjy
|
15264232ee
|
Super tiny fix CI (#13283)
|
2025-11-14 21:45:10 +08:00 |
|
fzyzcjy
|
3701f34dab
|
Tiny add utility to parse server logs (#12605)
|
2025-11-14 17:34:22 +08:00 |
|
fzyzcjy
|
821fb060c3
|
Enhance dumper comparator with tensor unifier and location finder (#12623)
|
2025-11-14 17:34:08 +08:00 |
|
fzyzcjy
|
ace27c0c01
|
Tiny enhance dumper with ctx and enable flags (#12622)
|
2025-11-14 17:33:49 +08:00 |
|
fzyzcjy
|
ed1d18d472
|
Tiny fix update version logic location (#12620)
|
2025-11-14 17:33:36 +08:00 |
|
fzyzcjy
|
1f134f850a
|
fix outdated router doc (#13255)
|
2025-11-14 14:00:42 +09:00 |
|
fzyzcjy
|
86255f27b4
|
Revert "fallback to triton mm_persistent kernel when deepGemm fail" (#13178)
|
2025-11-12 22:03:35 -08:00 |
|
fzyzcjy
|
b0ee99dd03
|
Super tiny fix typo (#13001)
|
2025-11-11 00:47:45 +08:00 |
|
fzyzcjy
|
2b6c4257a0
|
Fix sending all requests to the first rank in DP attention (#12832)
|
2025-11-09 00:18:53 +08:00 |
|
fzyzcjy
|
b7d7041190
|
Add sanity checks when a test file is not added to CI (reland) (#12594)
|
2025-11-04 18:04:26 +08:00 |
|
fzyzcjy
|
ff0b64e1e6
|
Ensure GPU work is finished when release memory occupation call is finished (#12592)
|
2025-11-04 18:01:27 +08:00 |
|
fzyzcjy
|
d84790db39
|
Support aggregating engine metrics in sgl-router (#11456)
|
2025-11-04 01:59:50 -08:00 |
|
fzyzcjy
|
60b0754cc9
|
Tiny fix ExpertDistributionReq error (#11760)
|
2025-11-04 13:39:25 +08:00 |
|
fzyzcjy
|
193fbb0bce
|
Super tiny add UT for copy_to_gpu_no_ce (#12270)
|
2025-11-04 09:40:51 +08:00 |
|
fzyzcjy
|
8834260739
|
Super tiny dump server info such as args in bench for post analysis (#12550)
|
2025-11-03 14:24:08 -08:00 |
|
fzyzcjy
|
fd7a72d62d
|
Super tiny allow profile activities in bench_serving (#12549)
|
2025-11-03 14:23:18 -08:00 |
|
fzyzcjy
|
385599cb04
|
Fix error when calling quantization (#12548)
|
2025-11-03 10:17:43 -08:00 |
|
fzyzcjy
|
c9db79117f
|
Super tiny fix naming in bench serving scripts (#12515)
|
2025-11-02 12:43:10 -08:00 |
|
fzyzcjy
|
30ad107028
|
Try to allow NCCL cumem for multi node nvlink case (#11987)
|
2025-10-31 12:48:25 -07:00 |
|
fzyzcjy
|
25257d8e00
|
Tiny assert no running requests when releasing memory to avoid IMA (#12341)
|
2025-11-01 01:28:53 +08:00 |
|
fzyzcjy
|
df5192cffa
|
Enable fast silu-and-mul-and-quant fused kernel (#11806)
|
2025-10-30 18:15:39 +08:00 |
|
fzyzcjy
|
fb52d35f63
|
Super tiny fix AMD ci (#12378)
|
2025-10-29 23:25:18 -07:00 |
|
fzyzcjy
|
25c5049870
|
Super tiny add tag for benchmark scripts (#12340)
|
2025-10-30 11:19:14 +08:00 |
|
fzyzcjy
|
29195aaa6e
|
Super tiny fix expert distribution dump error (#12271)
|
2025-10-28 15:20:55 -07:00 |
|
fzyzcjy
|
2a3763c335
|
Tiny fix sgl-kernel related CI installing the wrong binary (#12283)
|
2025-10-28 10:29:06 -07:00 |
|
 
|
691c8534cf
|
Support releasing CUDA graph memory when paused (#7873)
Co-authored-by: ryang-max <y1cunhui.yang@gmail.com>
Co-authored-by: ryang <38470282+ryang-max@users.noreply.github.com>
|
2025-10-28 14:40:50 +08:00 |
|
fzyzcjy
|
326c84c493
|
Compiling rope while preserving true on policy (#12161)
|
2025-10-28 08:02:17 +08:00 |
|
fzyzcjy
|
0103f374ba
|
Support DeepGEMM for deterministic inference (#12142)
|
2025-10-26 22:36:17 +08:00 |
|
fzyzcjy
|
c001deba37
|
Make bmm batch invariant injection optional (#12118)
|
2025-10-26 10:18:35 +08:00 |
|
fzyzcjy
|
20bd2271e2
|
Support true on-policy (#12058)
|
2025-10-25 10:23:42 +08:00 |
|
fzyzcjy
|
d7056c5236
|
Enhance tests in deterministic kernels (#12070)
|
2025-10-25 08:53:22 +08:00 |
|
fzyzcjy
|
e04340bf48
|
Fix multi processing serializer bug (#11958)
|
2025-10-24 22:53:45 +08:00 |
|
fzyzcjy
|
2342605ef0
|
Tiny cleanup send_single (#12056)
|
2025-10-23 23:53:42 -07:00 |
|
fzyzcjy
|
0f0c430e93
|
Install numactl in Dockerfile for GH200/GB200/GB300 (#11853)
|
2025-10-23 21:39:10 -07:00 |
|
fzyzcjy
|
8612811d85
|
Bump grace blackwell DeepEP version (#11990)
|
2025-10-22 21:08:12 -07:00 |
|
fzyzcjy
|
0917c5da8c
|
Support mixing cutedsl and deepgemm backend (#11807)
|
2025-10-21 07:38:35 +08:00 |
|
fzyzcjy
|
9e3be1fa2a
|
Tiny bump DeepEP version in ARM blackwell (#11810)
|
2025-10-20 08:15:14 +08:00 |
|
fzyzcjy
|
a8ba32798e
|
Fix triton_kernels import error on some hardwares (#11831)
|
2025-10-20 08:14:47 +08:00 |
|
fzyzcjy
|
12eb02e982
|
Change bf16 to fp8 for some gemms in attention for DeepSeek ckpt v2 (#11805)
|
2025-10-19 16:15:13 +08:00 |
|
 fzyzcjyandYineng Zhang
|
002d037359
|
Avoid generation gets hanging when user specifies multiple event loops (#5162)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-10-19 16:12:49 +08:00 |
|
fzyzcjy
|
a27825ae01
|
Support not officially supported high sgl-kernel version with low srt version (#11786)
|
2025-10-19 16:11:59 +08:00 |
|
fzyzcjy
|
ce399e154c
|
Make single-batch overlap compatible with NextN (#11804)
|
2025-10-19 16:10:44 +08:00 |
|
fzyzcjy
|
ea6275dfbc
|
Tiny add hints when users send requests to wrong place (#11808)
|
2025-10-19 16:10:20 +08:00 |
|
fzyzcjy
|
a7043c6f0d
|
Bump torch_memory_saver to avoid installing pre-release versions (#11797)
|
2025-10-18 01:20:42 -07:00 |
|
fzyzcjy
|
dbbd4e1891
|
Try add back no-commit-to-branch (#11799)
|
2025-10-18 12:05:12 +08:00 |
|
fzyzcjy
|
6c7c92eb02
|
Enable lint on main (#11794)
|
2025-10-17 19:08:50 -07:00 |
|
fzyzcjy
|
33e9bbec35
|
Make single-batch overlap compatible with offloading (#11614)
|
2025-10-18 08:45:54 +08:00 |
|
fzyzcjy
|
dcb8f090ad
|
Super tiny fix CI (#11788)
|
2025-10-17 17:41:58 -07:00 |
|
fzyzcjy
|
8af8491298
|
Support casting bf16 NextN moe to fp8 (#11613)
|
2025-10-18 08:02:15 +08:00 |
|
fzyzcjy
|
505329cab0
|
Support shared experts overlap in cutlass moe (#11611)
|
2025-10-18 07:59:40 +08:00 |
|
fzyzcjy
|
8a382fd399
|
Super tiny fix missing input throughput (#11607)
|
2025-10-18 07:58:48 +08:00 |
|
fzyzcjy
|
32803fb279
|
Super tiny improve FA3 import error message (#11590)
|
2025-10-14 22:06:31 -07:00 |
|
fzyzcjy
|
cb8ed2c09a
|
Make DeepEP combine recv do not overlap (#11535)
|
2025-10-13 18:40:42 -07:00 |
|
fzyzcjy
|
065ce81574
|
Tiny cleanup fp4 gemm calls (#11537)
|
2025-10-13 14:48:22 -07:00 |
|
fzyzcjy
|
bf3e7149be
|
Fix enable_v2 in int8 quant (#11470)
|
2025-10-11 21:56:30 +08:00 |
|
fzyzcjy
|
d957177a22
|
Super tiny delete unused openai router in sgl-router (#11448)
|
2025-10-11 15:59:30 +08:00 |
|
 fzyzcjyandYineng Zhang
|
21337b22b9
|
Reland [1/2] Optimizations and refactors about quant kernel (#10312)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-10-11 15:59:03 +08:00 |
|
+2        
|
efbc687c28
|
Support DeepSeek V3.2 Exp (#11061)
Co-authored-by: Stefan He <11166516+hebiao064@users.noreply.github.com>
Co-authored-by: Liangsheng Yin <95566987+hnyls2002@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <56809903+fridge003@users.noreply.github.com>
Co-authored-by: DarkSharpness <76582120+darksharpness@users.noreply.github.com>
Co-authored-by: ZhengdQin <46387172+zhengdqin@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Zhengda Qin <zhengdqin@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-10-06 00:24:15 -07:00 |
|
fzyzcjy
|
2f80bd9f0e
|
Bump torch_memory_saver 0.0.9rc2 (#11252)
|
2025-10-05 20:26:20 -07:00 |
|
fzyzcjy
|
fdc4e1e570
|
Tiny move files to utils folder (#11166)
|
2025-10-03 22:40:06 +08:00 |
|
fzyzcjy
|
6794d21051
|
Tiny add PD disaggregation + DP attention test (#11167)
|
2025-10-03 14:15:46 +08:00 |
|
fzyzcjy
|
afcd3e1089
|
Tiny remove duplicated code (#11164)
|
2025-10-02 21:56:31 +08:00 |
|
fzyzcjy
|
12d6818380
|
Tiny fix ep_gather behavior different in CI (#11130)
|
2025-10-02 21:55:53 +08:00 |
|
fzyzcjy
|
b65db0287b
|
Tiny cleanup deepseek_v2.py (#11163)
|
2025-10-02 21:54:52 +08:00 |
|
fzyzcjy
|
5e786cca3a
|
Support single batch overlap (#10422)
|
2025-10-02 18:04:36 +08:00 |
|
 fzyzcjyandKaixi Hou
|
0b9dfba787
|
Support dispatch low latency (#10263)
Co-authored-by: Kaixi Hou <4001424+kaixih@users.noreply.github.com>
|
2025-10-02 18:02:19 +08:00 |
|
fzyzcjy
|
2ac453b07f
|
Tiny detect slow ranks (#10508)
|
2025-10-02 18:00:33 +08:00 |
|
fzyzcjy
|
f35def8652
|
Fuse quantize and rope in trtllm_mla MTP (#10779)
|
2025-10-02 17:59:37 +08:00 |
|
fzyzcjy
|
d61615fe93
|
Tiny fix missing alt stream in nextn layer (#10768)
|
2025-10-02 17:58:23 +08:00 |
|
fzyzcjy
|
b1ccaf01cd
|
Tiny improve dumper (#11132)
|
2025-10-02 17:55:01 +08:00 |
|
fzyzcjy
|
44b1fbe258
|
Fix DeepSeek chunked prefill memory issue (#11149)
|
2025-10-01 23:56:59 -07:00 |
|
fzyzcjy
|
063c3791fe
|
Fix trtllm_mla slow concat kernel in MTP (#10777)
|
2025-09-22 22:47:49 -07:00 |
|
fzyzcjy
|
720c1c8ca3
|
Super tiny fix extra logs (#10697)
|
2025-09-20 21:30:54 -07:00 |
|
fzyzcjy
|
311de47bb7
|
[2/2] Speed up trtllm_mla attention backend (#10474)
|
2025-09-16 15:49:22 -07:00 |
|
fzyzcjy
|
8df7353af3
|
Support sgl-router parallel_batch in bench_one_batch_server (#10506)
|
2025-09-16 02:52:57 -07:00 |
|
fzyzcjy
|
ae4be601c2
|
Fix CI when sgl-kernel is changed but srt is not changed (#10515)
|
2025-09-16 02:49:54 -07:00 |
|
fzyzcjy
|
3b25dc127a
|
[1/2] Speed up trtllm_mla attention backend (>10% e2e) (#10473)
|
2025-09-15 11:53:21 -07:00 |
|
fzyzcjy
|
059c13de5c
|
Fix trtllm_moe wrong correction bias (#10440)
|
2025-09-15 01:02:05 -07:00 |
|
fzyzcjy
|
010181388c
|
Tiny fix wrong naming (#10437)
|
2025-09-14 19:24:41 -07:00 |
|
fzyzcjy
|
ca63f075b7
|
Revert "Fix FA4 import cause moe_fused_gate output be illegal memory" (#10432)
|
2025-09-14 19:03:27 -07:00 |
|
fzyzcjy
|
258d02c86d
|
Fix correction bias undefined behavior for nvfp4 models (#10426)
|
2025-09-14 18:41:09 -07:00 |
|
fzyzcjy
|
e3cf812f7d
|
Fix sgl-kernel + srt CI (#10419)
|
2025-09-14 01:44:47 -07:00 |
|
fzyzcjy
|
4da5533682
|
Support profile args in Engine API (#6539)
|
2025-09-14 01:21:10 -07:00 |
|
fzyzcjy
|
ac964d2e58
|
Support global scale in addition to per expert scale for cutedsl moe (#10270)
|
2025-09-14 01:17:00 -07:00 |
|
fzyzcjy
|
fa46e2bd40
|
Support offloading in fp8 (#9948)
|
2025-09-14 01:14:28 -07:00 |
|
fzyzcjy
|
b047b553c2
|
[2/2] Speed up prefill mla attention concat (#10157)
|
2025-09-14 01:12:04 -07:00 |
|
fzyzcjy
|
a0f844ed5a
|
Let sgl-kernel changes be tested on srt (#10313)
|
2025-09-14 01:09:17 -07:00 |
|
fzyzcjy
|
2df532ef20
|
Fix the global scale fix does not support EPLB and improve enabling condition (#10369)
|
2025-09-14 01:07:47 -07:00 |
|
fzyzcjy
|
abea9250da
|
Auto determine sgl kernel version in blackwell CI (#10318)
|
2025-09-14 01:06:30 -07:00 |
|
fzyzcjy
|
72dfa96aeb
|
Fix cutlass moe accuracy drop caused by attention UB from DP padding mode (#10414)
|
2025-09-13 22:29:09 -07:00 |
|
fzyzcjy
|
efedbe6ca9
|
Fix global input scale incompatible with CuTe DSL moe (#10370)
|
2025-09-12 03:22:49 -07:00 |
|
fzyzcjy
|
3a77c80b26
|
Fix FA4 import cause moe_fused_gate output be illegal memory (#10368)
|
2025-09-12 03:21:26 -07:00 |
|
fzyzcjy
|
0096798ed6
|
[1/2] Speed up prefill mla attention (#10156)
|
2025-09-08 09:00:33 -07:00 |
|
fzyzcjy
|
bc5fc332f7
|
Fix slow fused add RMSNorm (#10141)
|
2025-09-07 20:20:39 -07:00 |
|