Commit Graph
100 Commits
Author SHA1 Message Date
Lianmin Zheng dcc0a45618 Fix amd ci (#6360) 2025-05-16 15:33:10 -07:00
Lianmin Zheng c2b7ddca49 [Minor] cleanup unused imports (#6358) 2025-05-16 14:52:38 -07:00
Lianmin Zheng abebd9399c Update CODEOWNERS (#6359) 2025-05-16 14:51:36 -07:00
Lianmin Zheng e07a6977e7 Minor improvements of TokenizerManager / health check (#6327) 2025-05-15 15:29:25 -07:00
Lianmin Zheng ac2324c177 Skip the flaky test_stateful_custom_logit_processor (#6251) 2025-05-12 18:29:41 -07:00
Lianmin ZhengandSangBin Cho d18c6b3358 Support incremental streaming of logprob/token_ids between scheduler and detokenizer (#6225)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
2025-05-12 14:33:38 -07:00
Lianmin Zheng e8e18dcdcc Revert "fix some typos" (#6244) 2025-05-12 12:53:26 -07:00
Lianmin ZhengandSangBin Cho fba8eccd7e Log if cuda graph is used & extend cuda graph capture to cuda-graph-max-bs (#6201)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
2025-05-12 00:17:33 -07:00
Lianmin Zheng 6ea05950b1 Fix release-docs.yml to not use python 3.9 (#6204) 2025-05-11 16:04:55 -07:00
Lianmin Zheng e7dd906c5c Update README.md (#6202) 2025-05-11 14:34:12 -07:00
Lianmin Zheng 03227c5fa6 [CI] Reorganize the 8 gpu tests (#6192) 2025-05-11 10:55:06 -07:00
Lianmin Zheng 01bdbf7f80 Improve structured outputs: fix race condition, server crash, metrics and style (#6188) 2025-05-11 08:36:16 -07:00
Lianmin Zheng 17c36c5511 [CI] Disabled deepep tests temporarily because it takes too much time. (#6186) 2025-05-10 23:40:50 -07:00
Lianmin Zheng de167cf5fa Fix request abortion (#6184) 2025-05-10 21:54:46 -07:00
Lianmin Zheng 4319978c73 Fix data parallel perf regression (#6183) 2025-05-10 19:18:35 -07:00
Lianmin Zheng 38053c3372 Fix the timeout for 8 gpu tests (#6084) 2025-05-07 03:13:12 -07:00
Lianmin Zheng 26fc32d168 [CI] tune the test order to warmup the server (#5860) 2025-04-28 19:27:37 -07:00
Lianmin Zheng 849c83a0c0 [CI] test chunked prefill more (#5798) 2025-04-28 10:57:17 -07:00
Lianmin Zheng 693723d1f7 Revert "Tiny refactor DefaultModelLoader.Source" (#5825) 2025-04-28 01:18:57 -07:00
Lianmin Zheng 3029889cb4 Turn on overlap scheduler for multimodal models (#5771) 2025-04-27 23:45:09 -07:00
Lianmin Zheng daed453e84 [CI] Improve github summary & enable fa3 for more models (#5796) 2025-04-27 15:29:46 -07:00
Lianmin Zheng ded04b2e0a Update nightly-test.yml (#5797) 2025-04-27 15:27:24 -07:00
Lianmin Zheng a38f6932cc [CI] Fix test case (#5790) 2025-04-27 08:55:35 -07:00
Lianmin Zheng 621e96bf9b [CI] Fix ci tests (#5769) 2025-04-27 07:18:10 -07:00
Lianmin Zheng 35ca04d2fa [CI] fix port conflicts (#5789) 2025-04-27 05:17:44 -07:00
Lianmin Zheng 3c4e0ee64d [CI] Tune threshold (#5787) 2025-04-27 04:10:22 -07:00
Lianmin Zheng 9c088829ee Revert "Use device_id in dist init to reduce NCCL communicator warmup & creation overhead" (#5786) 2025-04-27 04:03:02 -07:00
Lianmin Zheng 005aad32ad Revert "[fix] fix bench_one_batch_server" (#5785) 2025-04-27 03:48:33 -07:00
Lianmin Zheng 4d23ba08f5 Simplify FA3 tests (#5779) 2025-04-27 01:30:17 -07:00
Lianmin Zheng 6e313c1b8b Revert "Revert "fix: import vllm_rotary_embedding error when head_size not in 64, 128, 256, 512"" (#5777) 2025-04-27 01:04:15 -07:00
Lianmin Zheng 981a2619d5 Fix eagle test case (#5776) 2025-04-27 01:00:54 -07:00
Lianmin Zheng 8ba313304d Revert "fix: import vllm_rotary_embedding error when head_size not in 64, 128, 256, 512" (#5772) 2025-04-26 23:26:08 -07:00
Lianmin Zheng 155890e4d1 [Minor] fix documentations (#5756) 2025-04-26 17:48:43 -07:00
Lianmin Zheng 21514ff5bd Disable flaky eagle tests (#5753) 2025-04-25 15:54:39 -07:00
Lianmin Zheng 5641a09458 Revert "[Model] Support ArcticForCausalLM architecture (Snowflake/snowflake-arctic-instruct)" (#5754) 2025-04-25 15:50:28 -07:00
Lianmin Zheng 3dd3538c18 Pin torch audio to 2.6.0 (#5750) 2025-04-25 15:06:28 -07:00
Lianmin Zheng de071366cd tune the threshold of gemma-2-27b-it in test_nightly_gsm8k_eval.py (#5677) 2025-04-23 05:31:17 -07:00
Lianmin Zheng 1343200299 Clean up mem settings (#5610) 2025-04-21 17:19:00 -07:00
Lianmin Zheng eef9433b46 Fix flush cache (#5590) 2025-04-20 22:56:40 -07:00
Lianmin Zheng fbdc94ba59 Release v0.4.5.post2 (#5582) 2025-04-20 14:12:37 -07:00
Lianmin Zheng 177320a582 Clean up imports (#5467) 2025-04-16 15:26:49 -07:00
Lianmin Zheng 0769b14bf9 [Minor] Move torch.compile patch to a better place (#5397) 2025-04-15 18:37:07 -07:00
Lianmin Zheng 838fa0f218 [minor] cleanup cmakelists.txt (#5420) 2025-04-15 07:07:07 -07:00
Lianmin Zheng dae7944440 minor clean up of sgl-kernel/CMakeLists.txt (#5393) 2025-04-14 18:38:44 -07:00
Lianmin Zheng 74885a848b Revert "Replace enable_flashinfer_mla argument with attention_backend" (#5048) 2025-04-03 13:30:56 -07:00
Lianmin Zheng 9adf178cc2 Fix 2-gpu CI test and suppress some warnings (#4930) 2025-03-30 12:51:44 -07:00
Lianmin Zheng f842853a40 Fix the timeout for unit-test-2-gpu in pr-test.yml (#4927) 2025-03-30 12:15:40 -07:00
Lianmin Zheng 4ede6770cd Fix retract for page size > 1 (#4914) 2025-03-30 02:57:15 -07:00
Lianmin Zheng b26bc86b36 Support page size > 1 + eagle (#4908) 2025-03-30 00:46:23 -07:00
Lianmin Zheng 74e0ac1dbd Clean up import vllm in quantization/__init__.py (#4834) 2025-03-28 10:34:10 -07:00
Lianmin Zheng 47e6628aae Fix CI tests (#4853) 2025-03-28 00:28:35 -07:00
Lianmin Zheng 2a882e8f3a Fix the nightly eval by lowering the threshold of neuralmagic/gemma-2-2b-it-FP8 (#4830) 2025-03-27 16:09:49 -07:00
Lianmin Zheng c38ca4fc8e Update readme (#4517) 2025-03-17 08:22:42 -07:00
Lianmin Zheng 82dec1f70b Remove redundant type conversion (#4513) 2025-03-17 05:57:35 -07:00
Lianmin Zheng 5493c3343e Fix data parallel + tensor parallel (#4499) 2025-03-17 05:13:16 -07:00
Lianmin Zheng 754a0e8278 Update CODEOWNERS (#4484) 2025-03-16 17:10:15 -07:00
Lianmin Zheng 3db35c1af4 Release sgl-kernel v0.0.5.post2 (#4469) 2025-03-16 01:01:53 -07:00
Lianmin Zheng 06d12b39d3 Remove filter for pr-tests (#4468) 2025-03-16 00:57:26 -07:00
Lianmin Zheng c30976fb41 Fix finish step for pr tests and notebook tests (#4467) 2025-03-16 00:52:06 -07:00
Lianmin Zheng 2c4f5ccac1 Fix minor style (#4460) 2025-03-15 21:51:12 -07:00
Lianmin Zheng e73167ade3 Fix maximum recursion depth triggered on exception exit (#4438) 2025-03-14 15:12:26 -07:00
Lianmin Zheng bb37855653 Update CODEOWNERS (#4403) 2025-03-13 17:54:40 -07:00
Lianmin Zheng f0afaf5289 Add a dummy grok test case (#4399) 2025-03-13 15:29:48 -07:00
c6d7f8d370 Add some fused elementwise kernels for grok-1 (#4398)
Co-authored-by: dhou-xai <dhou@x.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
2025-03-13 13:39:10 -07:00
Lianmin Zheng a5a892ffd3 Fix auto merge & add back get_flat_data_by_layer (#4393) 2025-03-13 08:46:25 -07:00
8e66fbecee Improve DP attention (#4390)
Co-authored-by: dhou-xai <dhou@x.ai>
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
2025-03-13 08:23:56 -07:00
f141298a3c Update ci_install_dependency.sh to use accelerate 1.4.0 (#4392)
Co-authored-by: wangyu <wangyu.steph@bytedance.com>
Co-authored-by: wangyu <yuwangauto@foxmail.com>
2025-03-13 07:16:11 -07:00
Lianmin Zheng 4fea040ca1 Fix a regression introduced by overlapping KV cache writing (#4375) 2025-03-13 03:49:05 -07:00
Lianmin Zheng 45de89719c Revert "[XPU][CPU] Enable the native path of DeepSeek" (#4367) 2025-03-12 23:45:52 -07:00
Lianmin Zheng c76040e31b Support page size > 1 (#4356) 2025-03-12 22:22:39 -07:00
Lianmin Zheng e35a93fa8a Move output processing logic from scheduler.py into a separate file (#4354) 2025-03-12 16:21:49 -07:00
Lianmin Zheng d40ee62b5d Update nightly tests (#4352) 2025-03-12 15:36:13 -07:00
Lianmin Zheng 5524e7d057 Fix nightly eval for neuralmagic/Mixtral-8x7B-Instruct-v0.1-FP8 (#4279) 2025-03-10 16:50:28 -07:00
Lianmin Zheng 5a6400eec5 Test no vllm custom allreduce (#4256) 2025-03-10 10:08:25 -07:00
Lianmin Zheng cf0ccd406e Optimize rope in sgl kernel (#4267) 2025-03-10 10:07:45 -07:00
Lianmin Zheng 3d56585a97 increase the timeout of nightly-test.yml (#4262) 2025-03-10 05:07:03 -07:00
Lianmin Zheng 00d25a7f5e Fix quantization and nightly tests (#4258) 2025-03-10 03:06:21 -07:00
Lianmin Zheng 1a5023e05d Release sgl-kernel v0.0.4.post1 (#4255) 2025-03-10 02:39:50 -07:00
Lianmin Zheng aa957102a9 Simplify tests & Fix trtllm custom allreduce registration (#4252) 2025-03-10 01:24:22 -07:00
Lianmin Zheng 7c0541b385 Move activation.cu to sgl-kernel/elementwise (#4250) 2025-03-09 22:41:13 -07:00
Lianmin Zheng e8a69e4d0c Clean up fp8 support (#4230) 2025-03-09 21:46:35 -07:00
Lianmin Zheng fbd560028a Auto balance CI tests (#4238) 2025-03-09 21:05:55 -07:00
Lianmin Zheng 730d084f2a Minor style fix for sgl-kernel (#4243) 2025-03-09 20:15:13 -07:00
Lianmin Zheng 4a05bdfa86 Revert "Check eagle server args" (#4242) 2025-03-09 18:53:33 -07:00
Lianmin Zheng eb06dbcbf8 Move rope and bmm into sgl-kernel (#4241) 2025-03-09 18:38:15 -07:00
Lianmin Zheng 1361ab9e03 Lazily import lora backends (#4225) 2025-03-08 23:39:26 -08:00
Lianmin Zhengandzhyncs 8abf74e3c9 Rename files in sgl kernel to avoid nested folder structure (#4213)
Co-authored-by: zhyncs <me@zhyncs.com>
2025-03-08 22:54:51 -08:00
Lianmin Zheng 48473684cc Split test_mla.py into two files (#4216) 2025-03-08 15:40:49 -08:00
Lianmin Zheng 2cadd51d11 Test no vllm custom allreduce (#4210) 2025-03-08 05:23:06 -08:00
Lianmin Zheng 8d323e95e4 Use clang format 18 in pr-test-sgl-kernel.yml (#4203) 2025-03-08 01:28:10 -08:00
Lianmin Zheng 08c4d764a5 lazy import attn backends (#4200) 2025-03-08 00:41:35 -08:00
d4017a6b63 [EAGLE] many fixes for eagle (#4195)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <sehoon@x.ai>
2025-03-07 22:12:13 -08:00
Lianmin Zheng d052f4c8a9 New clang format for sgl kernel (#4194) 2025-03-07 20:21:08 -08:00
Lianmin Zheng 9c58e68b4c Release v0.4.3.post4 (#4140) 2025-03-06 12:50:28 -08:00
bc1534ff32 Fix a draft model accuracy bug in eagle; support step=1; return logprob in eagle (#4134)
Co-authored-by: Sehoon Kim <kssteven418@gmail.com>
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <sehoon@x.ai>
2025-03-06 06:13:59 -08:00
Lianmin Zheng 800bf018fb Update CODEOWNER (#4138) 2025-03-06 03:42:10 -08:00
Lianmin Zheng 98c73d71cb [Minor] make the __init__ function of model_runner.py shorter (#4132) 2025-03-06 01:51:12 -08:00
Lianmin Zheng fcc2e37f69 Split the __init__ of scheduler as smaller functions. Improve the eagle tests (#4128) 2025-03-06 00:13:20 -08:00
Lianmin Zheng 286e6540a6 Remove prefill-only-one-req (#4117) 2025-03-05 20:58:48 -08:00
Lianmin Zheng e074d84e5b [Minor] more code cleanup (#4077) 2025-03-04 21:23:47 -08:00