Qubitium-modelcloud
|
bbec01c9aa
|
Fix tp worker only checking req[0] for stream (#546)
|
2024-06-14 22:56:10 -07:00 |
|
 QubitiumandZX
|
a8c787d2b3
|
Add ChatGLM Model Support (#516)
Co-authored-by: ZX <zx@lbx.dev>
|
2024-06-11 16:39:52 -07:00 |
|
 QubitiumandZX
|
f70f72586a
|
Fix rid state map leak + Refractor .finished (#505)
Co-authored-by: ZX <zx@lbx.dev>
|
2024-06-07 13:20:40 -07:00 |
|
 
|
33b242df30
|
Compat with latest VLLM 0.4.2 main + fork.number rename + Flashinfer 0.0.4 (#380)
Co-authored-by: ZX <zx@lbx.dev>
Co-authored-by: ZhouXingg <165115237+ZhouXingg@users.noreply.github.com>
|
2024-05-11 16:37:49 -07:00 |
|
 Qubitiumandhnyls2002
|
c9de3e169c
|
Eliminate 2 gpu ops during sampling when logit_bias is zero (#338)
Co-authored-by: hnyls2002 <hnyls2002@gmail.com>
|
2024-04-03 13:56:06 +08:00 |
|
Qubitium
|
eddaa2b599
|
Add support for new autogptq quant_config.checkpoint_format (#332)
|
2024-03-28 19:24:16 -07:00 |
|
Qubitium
|
ce216c80dc
|
Cleanup codebase: removed unnecessary code/logic (#298)
|
2024-03-23 10:15:16 -07:00 |
|
Qubitium
|
92e2d74fd0
|
Fix env (docker) compat due to __file__ usage (#288)
|
2024-03-13 13:02:48 +08:00 |
|
Qubitium
|
ad1dd74673
|
Fix flashinfer >= 0.0.3 compat (#282)
|
2024-03-12 21:45:58 +08:00 |
|
Qubitium
|
b2eb080501
|
Fix Runtime missing some ServerArgs options (#281)
|
2024-03-11 22:32:15 +08:00 |
|