Lianmin Zheng
|
dcc0a45618
|
Fix amd ci (#6360)
|
2025-05-16 15:33:10 -07:00 |
|
Lianmin Zheng
|
c2b7ddca49
|
[Minor] cleanup unused imports (#6358)
|
2025-05-16 14:52:38 -07:00 |
|
Lianmin Zheng
|
abebd9399c
|
Update CODEOWNERS (#6359)
|
2025-05-16 14:51:36 -07:00 |
|
Lianmin Zheng
|
e07a6977e7
|
Minor improvements of TokenizerManager / health check (#6327)
|
2025-05-15 15:29:25 -07:00 |
|
Lianmin Zheng
|
ac2324c177
|
Skip the flaky test_stateful_custom_logit_processor (#6251)
|
2025-05-12 18:29:41 -07:00 |
|
 Lianmin ZhengandSangBin Cho
|
d18c6b3358
|
Support incremental streaming of logprob/token_ids between scheduler and detokenizer (#6225)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
|
2025-05-12 14:33:38 -07:00 |
|
Lianmin Zheng
|
e8e18dcdcc
|
Revert "fix some typos" (#6244)
|
2025-05-12 12:53:26 -07:00 |
|
 Lianmin ZhengandSangBin Cho
|
fba8eccd7e
|
Log if cuda graph is used & extend cuda graph capture to cuda-graph-max-bs (#6201)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
|
2025-05-12 00:17:33 -07:00 |
|
Lianmin Zheng
|
6ea05950b1
|
Fix release-docs.yml to not use python 3.9 (#6204)
|
2025-05-11 16:04:55 -07:00 |
|
Lianmin Zheng
|
e7dd906c5c
|
Update README.md (#6202)
|
2025-05-11 14:34:12 -07:00 |
|
Lianmin Zheng
|
03227c5fa6
|
[CI] Reorganize the 8 gpu tests (#6192)
|
2025-05-11 10:55:06 -07:00 |
|
Lianmin Zheng
|
01bdbf7f80
|
Improve structured outputs: fix race condition, server crash, metrics and style (#6188)
|
2025-05-11 08:36:16 -07:00 |
|
Lianmin Zheng
|
17c36c5511
|
[CI] Disabled deepep tests temporarily because it takes too much time. (#6186)
|
2025-05-10 23:40:50 -07:00 |
|
Lianmin Zheng
|
de167cf5fa
|
Fix request abortion (#6184)
|
2025-05-10 21:54:46 -07:00 |
|
Lianmin Zheng
|
4319978c73
|
Fix data parallel perf regression (#6183)
|
2025-05-10 19:18:35 -07:00 |
|
Lianmin Zheng
|
38053c3372
|
Fix the timeout for 8 gpu tests (#6084)
|
2025-05-07 03:13:12 -07:00 |
|
Lianmin Zheng
|
26fc32d168
|
[CI] tune the test order to warmup the server (#5860)
|
2025-04-28 19:27:37 -07:00 |
|
Lianmin Zheng
|
849c83a0c0
|
[CI] test chunked prefill more (#5798)
|
2025-04-28 10:57:17 -07:00 |
|
Lianmin Zheng
|
693723d1f7
|
Revert "Tiny refactor DefaultModelLoader.Source" (#5825)
|
2025-04-28 01:18:57 -07:00 |
|
Lianmin Zheng
|
3029889cb4
|
Turn on overlap scheduler for multimodal models (#5771)
|
2025-04-27 23:45:09 -07:00 |
|
Lianmin Zheng
|
daed453e84
|
[CI] Improve github summary & enable fa3 for more models (#5796)
|
2025-04-27 15:29:46 -07:00 |
|
Lianmin Zheng
|
ded04b2e0a
|
Update nightly-test.yml (#5797)
|
2025-04-27 15:27:24 -07:00 |
|
Lianmin Zheng
|
a38f6932cc
|
[CI] Fix test case (#5790)
|
2025-04-27 08:55:35 -07:00 |
|
Lianmin Zheng
|
621e96bf9b
|
[CI] Fix ci tests (#5769)
|
2025-04-27 07:18:10 -07:00 |
|
Lianmin Zheng
|
35ca04d2fa
|
[CI] fix port conflicts (#5789)
|
2025-04-27 05:17:44 -07:00 |
|
Lianmin Zheng
|
3c4e0ee64d
|
[CI] Tune threshold (#5787)
|
2025-04-27 04:10:22 -07:00 |
|
Lianmin Zheng
|
9c088829ee
|
Revert "Use device_id in dist init to reduce NCCL communicator warmup & creation overhead" (#5786)
|
2025-04-27 04:03:02 -07:00 |
|
Lianmin Zheng
|
005aad32ad
|
Revert "[fix] fix bench_one_batch_server" (#5785)
|
2025-04-27 03:48:33 -07:00 |
|
Lianmin Zheng
|
4d23ba08f5
|
Simplify FA3 tests (#5779)
|
2025-04-27 01:30:17 -07:00 |
|
Lianmin Zheng
|
6e313c1b8b
|
Revert "Revert "fix: import vllm_rotary_embedding error when head_size not in 64, 128, 256, 512"" (#5777)
|
2025-04-27 01:04:15 -07:00 |
|
Lianmin Zheng
|
981a2619d5
|
Fix eagle test case (#5776)
|
2025-04-27 01:00:54 -07:00 |
|
Lianmin Zheng
|
8ba313304d
|
Revert "fix: import vllm_rotary_embedding error when head_size not in 64, 128, 256, 512" (#5772)
|
2025-04-26 23:26:08 -07:00 |
|
Lianmin Zheng
|
155890e4d1
|
[Minor] fix documentations (#5756)
|
2025-04-26 17:48:43 -07:00 |
|
Lianmin Zheng
|
21514ff5bd
|
Disable flaky eagle tests (#5753)
|
2025-04-25 15:54:39 -07:00 |
|
Lianmin Zheng
|
5641a09458
|
Revert "[Model] Support ArcticForCausalLM architecture (Snowflake/snowflake-arctic-instruct)" (#5754)
|
2025-04-25 15:50:28 -07:00 |
|
Lianmin Zheng
|
3dd3538c18
|
Pin torch audio to 2.6.0 (#5750)
|
2025-04-25 15:06:28 -07:00 |
|
Lianmin Zheng
|
de071366cd
|
tune the threshold of gemma-2-27b-it in test_nightly_gsm8k_eval.py (#5677)
|
2025-04-23 05:31:17 -07:00 |
|
Lianmin Zheng
|
1343200299
|
Clean up mem settings (#5610)
|
2025-04-21 17:19:00 -07:00 |
|
Lianmin Zheng
|
eef9433b46
|
Fix flush cache (#5590)
|
2025-04-20 22:56:40 -07:00 |
|
Lianmin Zheng
|
fbdc94ba59
|
Release v0.4.5.post2 (#5582)
|
2025-04-20 14:12:37 -07:00 |
|
Lianmin Zheng
|
177320a582
|
Clean up imports (#5467)
|
2025-04-16 15:26:49 -07:00 |
|
Lianmin Zheng
|
0769b14bf9
|
[Minor] Move torch.compile patch to a better place (#5397)
|
2025-04-15 18:37:07 -07:00 |
|
Lianmin Zheng
|
838fa0f218
|
[minor] cleanup cmakelists.txt (#5420)
|
2025-04-15 07:07:07 -07:00 |
|
Lianmin Zheng
|
dae7944440
|
minor clean up of sgl-kernel/CMakeLists.txt (#5393)
|
2025-04-14 18:38:44 -07:00 |
|
Lianmin Zheng
|
74885a848b
|
Revert "Replace enable_flashinfer_mla argument with attention_backend" (#5048)
|
2025-04-03 13:30:56 -07:00 |
|
Lianmin Zheng
|
9adf178cc2
|
Fix 2-gpu CI test and suppress some warnings (#4930)
|
2025-03-30 12:51:44 -07:00 |
|
Lianmin Zheng
|
f842853a40
|
Fix the timeout for unit-test-2-gpu in pr-test.yml (#4927)
|
2025-03-30 12:15:40 -07:00 |
|
Lianmin Zheng
|
4ede6770cd
|
Fix retract for page size > 1 (#4914)
|
2025-03-30 02:57:15 -07:00 |
|
Lianmin Zheng
|
b26bc86b36
|
Support page size > 1 + eagle (#4908)
|
2025-03-30 00:46:23 -07:00 |
|
Lianmin Zheng
|
74e0ac1dbd
|
Clean up import vllm in quantization/__init__.py (#4834)
|
2025-03-28 10:34:10 -07:00 |
|
Lianmin Zheng
|
47e6628aae
|
Fix CI tests (#4853)
|
2025-03-28 00:28:35 -07:00 |
|
Lianmin Zheng
|
2a882e8f3a
|
Fix the nightly eval by lowering the threshold of neuralmagic/gemma-2-2b-it-FP8 (#4830)
|
2025-03-27 16:09:49 -07:00 |
|
Lianmin Zheng
|
c38ca4fc8e
|
Update readme (#4517)
|
2025-03-17 08:22:42 -07:00 |
|
Lianmin Zheng
|
82dec1f70b
|
Remove redundant type conversion (#4513)
|
2025-03-17 05:57:35 -07:00 |
|
Lianmin Zheng
|
5493c3343e
|
Fix data parallel + tensor parallel (#4499)
|
2025-03-17 05:13:16 -07:00 |
|
Lianmin Zheng
|
754a0e8278
|
Update CODEOWNERS (#4484)
|
2025-03-16 17:10:15 -07:00 |
|
Lianmin Zheng
|
3db35c1af4
|
Release sgl-kernel v0.0.5.post2 (#4469)
|
2025-03-16 01:01:53 -07:00 |
|
Lianmin Zheng
|
06d12b39d3
|
Remove filter for pr-tests (#4468)
|
2025-03-16 00:57:26 -07:00 |
|
Lianmin Zheng
|
c30976fb41
|
Fix finish step for pr tests and notebook tests (#4467)
|
2025-03-16 00:52:06 -07:00 |
|
Lianmin Zheng
|
2c4f5ccac1
|
Fix minor style (#4460)
|
2025-03-15 21:51:12 -07:00 |
|
Lianmin Zheng
|
e73167ade3
|
Fix maximum recursion depth triggered on exception exit (#4438)
|
2025-03-14 15:12:26 -07:00 |
|
Lianmin Zheng
|
bb37855653
|
Update CODEOWNERS (#4403)
|
2025-03-13 17:54:40 -07:00 |
|
Lianmin Zheng
|
f0afaf5289
|
Add a dummy grok test case (#4399)
|
2025-03-13 15:29:48 -07:00 |
|
 
|
c6d7f8d370
|
Add some fused elementwise kernels for grok-1 (#4398)
Co-authored-by: dhou-xai <dhou@x.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
|
2025-03-13 13:39:10 -07:00 |
|
Lianmin Zheng
|
a5a892ffd3
|
Fix auto merge & add back get_flat_data_by_layer (#4393)
|
2025-03-13 08:46:25 -07:00 |
|
 
|
8e66fbecee
|
Improve DP attention (#4390)
Co-authored-by: dhou-xai <dhou@x.ai>
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
|
2025-03-13 08:23:56 -07:00 |
|
 
|
f141298a3c
|
Update ci_install_dependency.sh to use accelerate 1.4.0 (#4392)
Co-authored-by: wangyu <wangyu.steph@bytedance.com>
Co-authored-by: wangyu <yuwangauto@foxmail.com>
|
2025-03-13 07:16:11 -07:00 |
|
Lianmin Zheng
|
4fea040ca1
|
Fix a regression introduced by overlapping KV cache writing (#4375)
|
2025-03-13 03:49:05 -07:00 |
|
Lianmin Zheng
|
45de89719c
|
Revert "[XPU][CPU] Enable the native path of DeepSeek" (#4367)
|
2025-03-12 23:45:52 -07:00 |
|
Lianmin Zheng
|
c76040e31b
|
Support page size > 1 (#4356)
|
2025-03-12 22:22:39 -07:00 |
|
Lianmin Zheng
|
e35a93fa8a
|
Move output processing logic from scheduler.py into a separate file (#4354)
|
2025-03-12 16:21:49 -07:00 |
|
Lianmin Zheng
|
d40ee62b5d
|
Update nightly tests (#4352)
|
2025-03-12 15:36:13 -07:00 |
|
Lianmin Zheng
|
5524e7d057
|
Fix nightly eval for neuralmagic/Mixtral-8x7B-Instruct-v0.1-FP8 (#4279)
|
2025-03-10 16:50:28 -07:00 |
|
Lianmin Zheng
|
5a6400eec5
|
Test no vllm custom allreduce (#4256)
|
2025-03-10 10:08:25 -07:00 |
|
Lianmin Zheng
|
cf0ccd406e
|
Optimize rope in sgl kernel (#4267)
|
2025-03-10 10:07:45 -07:00 |
|
Lianmin Zheng
|
3d56585a97
|
increase the timeout of nightly-test.yml (#4262)
|
2025-03-10 05:07:03 -07:00 |
|
Lianmin Zheng
|
00d25a7f5e
|
Fix quantization and nightly tests (#4258)
|
2025-03-10 03:06:21 -07:00 |
|
Lianmin Zheng
|
1a5023e05d
|
Release sgl-kernel v0.0.4.post1 (#4255)
|
2025-03-10 02:39:50 -07:00 |
|
Lianmin Zheng
|
aa957102a9
|
Simplify tests & Fix trtllm custom allreduce registration (#4252)
|
2025-03-10 01:24:22 -07:00 |
|
Lianmin Zheng
|
7c0541b385
|
Move activation.cu to sgl-kernel/elementwise (#4250)
|
2025-03-09 22:41:13 -07:00 |
|
Lianmin Zheng
|
e8a69e4d0c
|
Clean up fp8 support (#4230)
|
2025-03-09 21:46:35 -07:00 |
|
Lianmin Zheng
|
fbd560028a
|
Auto balance CI tests (#4238)
|
2025-03-09 21:05:55 -07:00 |
|
Lianmin Zheng
|
730d084f2a
|
Minor style fix for sgl-kernel (#4243)
|
2025-03-09 20:15:13 -07:00 |
|
Lianmin Zheng
|
4a05bdfa86
|
Revert "Check eagle server args" (#4242)
|
2025-03-09 18:53:33 -07:00 |
|
Lianmin Zheng
|
eb06dbcbf8
|
Move rope and bmm into sgl-kernel (#4241)
|
2025-03-09 18:38:15 -07:00 |
|
Lianmin Zheng
|
1361ab9e03
|
Lazily import lora backends (#4225)
|
2025-03-08 23:39:26 -08:00 |
|
 Lianmin Zhengandzhyncs
|
8abf74e3c9
|
Rename files in sgl kernel to avoid nested folder structure (#4213)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-03-08 22:54:51 -08:00 |
|
Lianmin Zheng
|
48473684cc
|
Split test_mla.py into two files (#4216)
|
2025-03-08 15:40:49 -08:00 |
|
Lianmin Zheng
|
2cadd51d11
|
Test no vllm custom allreduce (#4210)
|
2025-03-08 05:23:06 -08:00 |
|
Lianmin Zheng
|
8d323e95e4
|
Use clang format 18 in pr-test-sgl-kernel.yml (#4203)
|
2025-03-08 01:28:10 -08:00 |
|
Lianmin Zheng
|
08c4d764a5
|
lazy import attn backends (#4200)
|
2025-03-08 00:41:35 -08:00 |
|
 
|
d4017a6b63
|
[EAGLE] many fixes for eagle (#4195)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <sehoon@x.ai>
|
2025-03-07 22:12:13 -08:00 |
|
Lianmin Zheng
|
d052f4c8a9
|
New clang format for sgl kernel (#4194)
|
2025-03-07 20:21:08 -08:00 |
|
Lianmin Zheng
|
9c58e68b4c
|
Release v0.4.3.post4 (#4140)
|
2025-03-06 12:50:28 -08:00 |
|
  
|
bc1534ff32
|
Fix a draft model accuracy bug in eagle; support step=1; return logprob in eagle (#4134)
Co-authored-by: Sehoon Kim <kssteven418@gmail.com>
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <sehoon@x.ai>
|
2025-03-06 06:13:59 -08:00 |
|
Lianmin Zheng
|
800bf018fb
|
Update CODEOWNER (#4138)
|
2025-03-06 03:42:10 -08:00 |
|
Lianmin Zheng
|
98c73d71cb
|
[Minor] make the __init__ function of model_runner.py shorter (#4132)
|
2025-03-06 01:51:12 -08:00 |
|
Lianmin Zheng
|
fcc2e37f69
|
Split the __init__ of scheduler as smaller functions. Improve the eagle tests (#4128)
|
2025-03-06 00:13:20 -08:00 |
|
Lianmin Zheng
|
286e6540a6
|
Remove prefill-only-one-req (#4117)
|
2025-03-05 20:58:48 -08:00 |
|
Lianmin Zheng
|
e074d84e5b
|
[Minor] more code cleanup (#4077)
|
2025-03-04 21:23:47 -08:00 |
|