Glen Liu
|
76b9c8de6f
|
[Feature] add LoRADrainer to address high P99 TTFT (#17913)
|
2026-05-02 16:13:43 -07:00 |
|
 Glen LiuandEthan Su
|
e0474fdd9b
|
throw ValueError for DoRA adapters (#22125)
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
|
2026-05-02 14:54:19 +00:00 |
|
Glen Liu
|
1eed219f47
|
add mixed chunk unit test and make small refactors (#18776)
|
2026-03-08 19:56:09 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Glen Liuandgemini-code-assist[bot]
|
cc860a2198
|
[TestFix] change LoRA tests to use NVIDIA adapter instead of Nutanix (#19642)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-02 12:55:41 -08:00 |
|
Glen Liu
|
8a0b7575b0
|
[Test] add unit test for skipping already preempted request (#18912)
|
2026-03-01 15:44:15 -08:00 |
|
Glen Liu
|
a06425184f
|
[Fix] remove redundant +1 when getting tail_str for Req (#18584)
|
2026-02-24 10:35:26 -08:00 |
|
Glen Liu
|
3f32a5831d
|
throw error if got adapter with added_tokens (#18046)
|
2026-02-05 23:55:43 +08:00 |
|
Glen Liu
|
fe57a887b1
|
[TestFix] use unit tests for LoRA overlap loading tests (#18140)
|
2026-02-02 22:06:50 -08:00 |
|
Glen Liu
|
99dad105fd
|
[TestFix] rewrite LoRA overlap loading tests (#18047)
|
2026-02-01 14:52:08 -08:00 |
|
Glen Liu
|
a6280b2a23
|
add documentation example for LoRA overlap loading and cleanup unused function (#17464)
|
2026-01-24 15:33:16 +08:00 |
|
Glen Liu
|
ad1b4e4728
|
[Feature] overlap LoRA weight loading with compute (#15512)
|
2026-01-19 10:43:17 +08:00 |
|
Glen Liu
|
6b065298b5
|
[Docs] add routing-key to schedule-policy in docs (#17101)
|
2026-01-14 22:22:07 -05:00 |
|
Glen Liu
|
6327dff242
|
enhance LoRA tests and fix base model LoRA eviction in Scheduler (#16333)
|
2026-01-10 16:49:00 +08:00 |
|
Glen Liu
|
eb1d885400
|
add LoRA warning if loading a preexisting LoRA adapter with a different name (#13822)
|
2025-11-24 15:16:41 -08:00 |
|
Glen Liu
|
53620a1b1a
|
fix test_lora_update.py starvation message check (#13702)
|
2025-11-21 19:33:04 -08:00 |
|
Glen Liu
|
750084ae08
|
remove unnecessary starvation check (#13619)
|
2025-11-20 19:10:51 -08:00 |
|
Glen Liu
|
ada8ce1fd0
|
allow loras to be implicitly evicted and loaded based on max_loaded_loras (#11526)
|
2025-11-20 13:34:32 -08:00 |
|
Glen Liu
|
d79e12941c
|
Small cleanups related to LoRA weight loading (#13474)
|
2025-11-18 09:14:34 -08:00 |
|
Glen Liu
|
cbf23dbbfa
|
[Feature] add --lora-request-distribution arg to bench_serving.py and support skewed and distinct workloads (#12175)
|
2025-11-04 21:41:40 -08:00 |
|
Glen Liu
|
fc86b18b3e
|
adjust dynamic vs static outputs comparison in test_lora_update.py (#11884)
|
2025-10-24 10:35:34 -07:00 |
|
Glen Liu
|
47c606d3dc
|
[Feature] support regex strings as a stopping condition (#10635)
|
2025-10-12 10:53:15 +08:00 |
|
Glen Liu
|
9a7e7a6576
|
[Bug Fix] prevent lora adapter from being loaded into LoRAManager if it is already loaded (#11365)
|
2025-10-09 18:43:03 -07:00 |
|
Glen Liu
|
ebd0e1c18b
|
[doc] add walkthrough for implementing and hosting a simple llama wrapper m… (#10093)
|
2025-09-10 12:05:06 +08:00 |
|