Direct model loading from object storage with Runai Model Streamer (#17948)

Signed-off-by: Noa Neria <noa@run.ai>
This commit is contained in:
Noa Neria
2026-04-01 18:41:22 -07:00
committed by GitHub
parent ae3b207dfd
commit 8d9145d97e
14 changed files with 659 additions and 17 deletions
+108
View File
@@ -0,0 +1,108 @@
# Loading Models from Object Storage
SGLang supports direct loading of models from object storage (S3 and Google Cloud Storage) without requiring a full local download. This feature uses the `runai_streamer` load format to stream model weights directly from cloud storage, significantly reducing startup time and local storage requirements.
## Overview
When loading models from object storage, SGLang uses a two-phase approach:
1. **Metadata Download** (once, before process launch): Configuration files and tokenizer files are downloaded to a local cache
2. **Weight Streaming** (lazy, during model loading): Model weights are streamed directly from object storage as needed
## Supported Storage Backends
1. **Amazon S3**: `s3://bucket-name/path/to/model/`
2. **Google Cloud Storage**: `gs://bucket-name/path/to/model/`
3. **Azure Blob**: `az://some-azure-container/path/`
4. **S3 compatible**: `s3://bucket-name/path/to/model/`
## Quick Start
### Basic Usage
Simply provide an object storage URI as the model path:
```bash
# S3
python -m sglang.launch_server \
--model-path s3://my-bucket/models/llama-3-8b/ \
--load-format runai_streamer
# Google Cloud Storage
python -m sglang.launch_server \
--model-path gs://my-bucket/models/llama-3-8b/ \
--load-format runai_streamer
```
**Note**: The `--load-format runai_streamer` is automatically detected when using object storage URIs, so you can omit it:
```bash
python -m sglang.launch_server \
--model-path s3://my-bucket/models/llama-3-8b/
```
### With Tensor Parallelism
```bash
python -m sglang.launch_server \
--model-path gs://my-bucket/models/llama-70b/ \
--tp 4 \
--model-loader-extra-config '{"distributed": true}'
```
## Configuration
### Load Format
The `runai_streamer` load format is specifically designed for object storage, ssd and shared file systems
```bash
python -m sglang.launch_server \
--model-path s3://bucket/model/ \
--load-format runai_streamer
```
### Extended Configuration Parameters
Use `--model-loader-extra-config` to pass additional configuration as a JSON string:
```bash
python -m sglang.launch_server \
--model-path s3://bucket/model/ \
--model-loader-extra-config '{
"distributed": true,
"concurrency": 8,
"memory_limit": 2147483648
}'
```
#### Available Parameters
| Parameter | Type | Description | Default |
|-----------|------|-------------|---------|
| `distributed` | bool | Enable distributed streaming for multi-GPU setups. Automatically set to `true` for object storage paths and cuda alike devices. | Auto-detected |
| `concurrency` | int | Number of concurrent download streams. Higher values can improve throughput for large models. | 4 |
| `memory_limit` | int | Memory limit (in bytes) for the streaming buffer. | System-dependent |
## Performance Considerations
### Distributed Streaming
For multi-GPU setups, enable distributed streaming to parallelize weight loading between the processes:
```bash
python -m sglang.launch_server \
--model-path s3://bucket/model/ \
--tp 8 \
--model-loader-extra-config '{"distributed": true}'
```
## Limitations
- **Supported Formats**: Currently only supports `.safetensors` weight format (recommended format)
- **Supported Device**: Distributed streaming is supported on cuda alike devices. Otherwise fallback to non distributed streaming
## See Also
- [Runai model streamer documentation](https://github.com/run-ai/runai-model-streamer)
+1 -1
View File
@@ -84,7 +84,7 @@ Please consult the documentation below and [server_args.py](https://github.com/s
| `--tokenizer-mode` | Tokenizer mode. 'auto' will use the fast tokenizer if available, and 'slow' will always use the slow tokenizer. | `auto` | `auto`, `slow` |
| `--tokenizer-worker-num` | The worker num of the tokenizer manager. | `1` | Type: int |
| `--skip-tokenizer-init` | If set, skip init tokenizer and pass input_ids in generate request. | `False` | bool flag (set to enable) |
| `--load-format` | The format of the model weights to load. "auto" will try to load the weights in the safetensors format and fall back to the pytorch bin format if safetensors format is not available. "pt" will load the weights in the pytorch bin format. "safetensors" will load the weights in the safetensors format. "npcache" will load the weights in pytorch format and store a numpy cache to speed up the loading. "dummy" will initialize the weights with random values, which is mainly for profiling."gguf" will load the weights in the gguf format. "bitsandbytes" will load the weights using bitsandbytes quantization."layered" loads weights layer by layer so that one can quantize a layer before loading another to make the peak memory envelope smaller. "flash_rl" will load the weights in flash_rl format. "fastsafetensors" and "private" are also supported. | `auto` | `auto`, `pt`, `safetensors`, `npcache`, `dummy`, `sharded_state`, `gguf`, `bitsandbytes`, `layered`, `flash_rl`, `remote`, `remote_instance`, `fastsafetensors`, `private` |
| `--load-format` | The format of the model weights to load. "auto" will try to load the weights in the safetensors format and fall back to the pytorch bin format if safetensors format is not available. "pt" will load the weights in the pytorch bin format. "safetensors" will load the weights in the safetensors format. "npcache" will load the weights in pytorch format and store a numpy cache to speed up the loading. "dummy" will initialize the weights with random values, which is mainly for profiling."gguf" will load the weights in the gguf format. "bitsandbytes" will load the weights using bitsandbytes quantization."layered" loads weights layer by layer so that one can quantize a layer before loading another to make the peak memory envelope smaller. "flash_rl" will load the weights in flash_rl format. "fastsafetensors" and "private" are also supported. "runai_streamer" enables direct model loading from object storage and shared file systems.| `auto` | `auto`, `pt`, `safetensors`, `npcache`, `dummy`, `sharded_state`, `gguf`, `bitsandbytes`, `layered`, `flash_rl`, `remote`, `remote_instance`, `fastsafetensors`, `private`, `runai_streamer` |
| `--model-loader-extra-config` | Extra config for model loader. This will be passed to the model loader corresponding to the chosen load_format. | `{}` | Type: str |
| `--trust-remote-code` | Whether or not to allow for custom models defined on the Hub in their own modeling files. | `False` | bool flag (set to enable) |
| `--context-length` | The model's maximum context length. Defaults to None (will use the value from the model's config.json instead). | `None` | Type: int |
+1
View File
@@ -41,6 +41,7 @@ Its core features include:
:caption: Advanced Features
advanced_features/server_arguments.md
advanced_features/object_storage.md
advanced_features/hyperparameter_tuning.md
advanced_features/attention_backend.md
advanced_features/speculative_decoding.ipynb