Direct model loading from object storage with Runai Model Streamer (#17948)
Signed-off-by: Noa Neria <noa@run.ai>
This commit is contained in:
@@ -0,0 +1,108 @@
|
||||
# Loading Models from Object Storage
|
||||
|
||||
SGLang supports direct loading of models from object storage (S3 and Google Cloud Storage) without requiring a full local download. This feature uses the `runai_streamer` load format to stream model weights directly from cloud storage, significantly reducing startup time and local storage requirements.
|
||||
|
||||
## Overview
|
||||
|
||||
When loading models from object storage, SGLang uses a two-phase approach:
|
||||
|
||||
1. **Metadata Download** (once, before process launch): Configuration files and tokenizer files are downloaded to a local cache
|
||||
2. **Weight Streaming** (lazy, during model loading): Model weights are streamed directly from object storage as needed
|
||||
|
||||
## Supported Storage Backends
|
||||
|
||||
1. **Amazon S3**: `s3://bucket-name/path/to/model/`
|
||||
2. **Google Cloud Storage**: `gs://bucket-name/path/to/model/`
|
||||
3. **Azure Blob**: `az://some-azure-container/path/`
|
||||
4. **S3 compatible**: `s3://bucket-name/path/to/model/`
|
||||
|
||||
## Quick Start
|
||||
|
||||
### Basic Usage
|
||||
|
||||
Simply provide an object storage URI as the model path:
|
||||
|
||||
```bash
|
||||
# S3
|
||||
python -m sglang.launch_server \
|
||||
--model-path s3://my-bucket/models/llama-3-8b/ \
|
||||
--load-format runai_streamer
|
||||
|
||||
# Google Cloud Storage
|
||||
python -m sglang.launch_server \
|
||||
--model-path gs://my-bucket/models/llama-3-8b/ \
|
||||
--load-format runai_streamer
|
||||
```
|
||||
|
||||
**Note**: The `--load-format runai_streamer` is automatically detected when using object storage URIs, so you can omit it:
|
||||
|
||||
```bash
|
||||
python -m sglang.launch_server \
|
||||
--model-path s3://my-bucket/models/llama-3-8b/
|
||||
```
|
||||
|
||||
### With Tensor Parallelism
|
||||
|
||||
```bash
|
||||
python -m sglang.launch_server \
|
||||
--model-path gs://my-bucket/models/llama-70b/ \
|
||||
--tp 4 \
|
||||
--model-loader-extra-config '{"distributed": true}'
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
### Load Format
|
||||
|
||||
The `runai_streamer` load format is specifically designed for object storage, ssd and shared file systems
|
||||
|
||||
```bash
|
||||
python -m sglang.launch_server \
|
||||
--model-path s3://bucket/model/ \
|
||||
--load-format runai_streamer
|
||||
```
|
||||
|
||||
### Extended Configuration Parameters
|
||||
|
||||
Use `--model-loader-extra-config` to pass additional configuration as a JSON string:
|
||||
|
||||
```bash
|
||||
python -m sglang.launch_server \
|
||||
--model-path s3://bucket/model/ \
|
||||
--model-loader-extra-config '{
|
||||
"distributed": true,
|
||||
"concurrency": 8,
|
||||
"memory_limit": 2147483648
|
||||
}'
|
||||
```
|
||||
|
||||
#### Available Parameters
|
||||
|
||||
| Parameter | Type | Description | Default |
|
||||
|-----------|------|-------------|---------|
|
||||
| `distributed` | bool | Enable distributed streaming for multi-GPU setups. Automatically set to `true` for object storage paths and cuda alike devices. | Auto-detected |
|
||||
| `concurrency` | int | Number of concurrent download streams. Higher values can improve throughput for large models. | 4 |
|
||||
| `memory_limit` | int | Memory limit (in bytes) for the streaming buffer. | System-dependent |
|
||||
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### Distributed Streaming
|
||||
|
||||
For multi-GPU setups, enable distributed streaming to parallelize weight loading between the processes:
|
||||
|
||||
```bash
|
||||
python -m sglang.launch_server \
|
||||
--model-path s3://bucket/model/ \
|
||||
--tp 8 \
|
||||
--model-loader-extra-config '{"distributed": true}'
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- **Supported Formats**: Currently only supports `.safetensors` weight format (recommended format)
|
||||
- **Supported Device**: Distributed streaming is supported on cuda alike devices. Otherwise fallback to non distributed streaming
|
||||
|
||||
## See Also
|
||||
|
||||
- [Runai model streamer documentation](https://github.com/run-ai/runai-model-streamer)
|
||||
@@ -84,7 +84,7 @@ Please consult the documentation below and [server_args.py](https://github.com/s
|
||||
| `--tokenizer-mode` | Tokenizer mode. 'auto' will use the fast tokenizer if available, and 'slow' will always use the slow tokenizer. | `auto` | `auto`, `slow` |
|
||||
| `--tokenizer-worker-num` | The worker num of the tokenizer manager. | `1` | Type: int |
|
||||
| `--skip-tokenizer-init` | If set, skip init tokenizer and pass input_ids in generate request. | `False` | bool flag (set to enable) |
|
||||
| `--load-format` | The format of the model weights to load. "auto" will try to load the weights in the safetensors format and fall back to the pytorch bin format if safetensors format is not available. "pt" will load the weights in the pytorch bin format. "safetensors" will load the weights in the safetensors format. "npcache" will load the weights in pytorch format and store a numpy cache to speed up the loading. "dummy" will initialize the weights with random values, which is mainly for profiling."gguf" will load the weights in the gguf format. "bitsandbytes" will load the weights using bitsandbytes quantization."layered" loads weights layer by layer so that one can quantize a layer before loading another to make the peak memory envelope smaller. "flash_rl" will load the weights in flash_rl format. "fastsafetensors" and "private" are also supported. | `auto` | `auto`, `pt`, `safetensors`, `npcache`, `dummy`, `sharded_state`, `gguf`, `bitsandbytes`, `layered`, `flash_rl`, `remote`, `remote_instance`, `fastsafetensors`, `private` |
|
||||
| `--load-format` | The format of the model weights to load. "auto" will try to load the weights in the safetensors format and fall back to the pytorch bin format if safetensors format is not available. "pt" will load the weights in the pytorch bin format. "safetensors" will load the weights in the safetensors format. "npcache" will load the weights in pytorch format and store a numpy cache to speed up the loading. "dummy" will initialize the weights with random values, which is mainly for profiling."gguf" will load the weights in the gguf format. "bitsandbytes" will load the weights using bitsandbytes quantization."layered" loads weights layer by layer so that one can quantize a layer before loading another to make the peak memory envelope smaller. "flash_rl" will load the weights in flash_rl format. "fastsafetensors" and "private" are also supported. "runai_streamer" enables direct model loading from object storage and shared file systems.| `auto` | `auto`, `pt`, `safetensors`, `npcache`, `dummy`, `sharded_state`, `gguf`, `bitsandbytes`, `layered`, `flash_rl`, `remote`, `remote_instance`, `fastsafetensors`, `private`, `runai_streamer` |
|
||||
| `--model-loader-extra-config` | Extra config for model loader. This will be passed to the model loader corresponding to the chosen load_format. | `{}` | Type: str |
|
||||
| `--trust-remote-code` | Whether or not to allow for custom models defined on the Hub in their own modeling files. | `False` | bool flag (set to enable) |
|
||||
| `--context-length` | The model's maximum context length. Defaults to None (will use the value from the model's config.json instead). | `None` | Type: int |
|
||||
|
||||
Reference in New Issue
Block a user