vlm: refactor engine vlm params and support processor output as input (#14091)

Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: BenYao21 <cyao22@asu.edu>
Co-authored-by: minleminzui <minleminzui@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
This commit is contained in:
mlmz
2025-12-20 18:31:24 +08:00
committed by GitHub
co-authored by Mick zhaochenyang20 Xinyuan Tong BenYao21 minleminzui gemini-code-assist[bot] 赵晨阳
parent 165f5c04cb
commit 1f1f05a85e
16 changed files with 783 additions and 305 deletions
+225 -161
View File
@@ -5,7 +5,13 @@
"id": "0",
"metadata": {},
"source": [
"# Query Vision Language Model"
"# Query VLM with Offline Engine\n",
"\n",
"This tutorial demonstrates how to use SGLang's **offline Engine API** to query VLMs. We will demonstrate usage with Qwen2.5-VL and Llama 4. This section demonstrates three different calling approaches:\n",
"\n",
"1. **Basic Call**: Directly pass images and text.\n",
"2. **Processor Output**: Use HuggingFace processor for data preprocessing.\n",
"3. **Precomputed Embeddings**: Pre-calculate image features to improve inference efficiency."
]
},
{
@@ -13,22 +19,38 @@
"id": "1",
"metadata": {},
"source": [
"## Querying Qwen-VL"
"## Understanding the Three Input Formats\n",
"\n",
"SGLang supports three ways to pass visual data, each optimized for different scenarios:\n",
"\n",
"### 1. **Raw Images** - Simplest approach\n",
"- Pass PIL Images, file paths, URLs, or base64 strings directly\n",
"- SGLang handles all preprocessing automatically\n",
"- Best for: Quick prototyping, simple applications\n",
"\n",
"### 2. **Processor Output** - For custom preprocessing\n",
"- Pre-process images with HuggingFace processor\n",
"- Pass the complete processor output dict with `format: \"processor_output\"`\n",
"- Best for: Custom image transformations, integration with existing pipelines\n",
"- Requirement: Must use `input_ids` instead of text prompt\n",
"\n",
"### 3. **Precomputed Embeddings** - For maximum performance\n",
"- Pre-calculate visual embeddings using the vision encoder\n",
"- Pass embeddings with `format: \"precomputed_embedding\"`\n",
"- Best for: Repeated queries on same images, caching, high-throughput serving\n",
"- Performance gain: Avoids redundant vision encoder computation (30-50% speedup)\n",
"\n",
"**Key Rule**: Within a single request, use only one format for all images. Don't mix formats.\n",
"\n",
"The examples below demonstrate all three approaches with both Qwen2.5-VL and Llama 4 models."
]
},
{
"cell_type": "code",
"execution_count": null,
"cell_type": "markdown",
"id": "2",
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply() # Run this first.\n",
"\n",
"model_path = \"Qwen/Qwen2.5-VL-3B-Instruct\"\n",
"chat_template = \"qwen2-vl\""
"## Querying Qwen2.5-VL Model"
]
},
{
@@ -38,8 +60,21 @@
"metadata": {},
"outputs": [],
"source": [
"# Lets create a prompt.\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"model_path = \"Qwen/Qwen2.5-VL-3B-Instruct\"\n",
"chat_template = \"qwen2-vl\""
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "4",
"metadata": {},
"outputs": [],
"source": [
"from io import BytesIO\n",
"import requests\n",
"from PIL import Image\n",
@@ -59,30 +94,18 @@
"conv.append_message(conv.roles[1], \"\")\n",
"conv.image_data = [image]\n",
"\n",
"print(\"Generated prompt text:\")\n",
"print(conv.get_prompt())\n",
"print(f\"\\nImage size: {image.size}\")\n",
"image"
]
},
{
"cell_type": "markdown",
"id": "4",
"metadata": {},
"source": [
"### Query via the offline Engine API"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "5",
"metadata": {},
"outputs": [],
"source": [
"from sglang import Engine\n",
"\n",
"llm = Engine(\n",
" model_path=model_path, chat_template=chat_template, mem_fraction_static=0.8\n",
")"
"### Basic Offline Engine API Call"
]
},
{
@@ -92,27 +115,73 @@
"metadata": {},
"outputs": [],
"source": [
"out = llm.generate(prompt=conv.get_prompt(), image_data=[image])\n",
"print(out[\"text\"])"
]
},
{
"cell_type": "markdown",
"id": "7",
"metadata": {},
"source": [
"### Query via the offline Engine API, but send precomputed embeddings"
"from sglang import Engine\n",
"\n",
"\n",
"llm = Engine(model_path=model_path, chat_template=chat_template, log_level=\"warning\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "8",
"id": "7",
"metadata": {},
"outputs": [],
"source": [
"# Compute the image embeddings using Huggingface.\n",
"out = llm.generate(prompt=conv.get_prompt(), image_data=[image])\n",
"print(\"Model response:\")\n",
"print(out[\"text\"])"
]
},
{
"cell_type": "markdown",
"id": "8",
"metadata": {},
"source": [
"### Call with Processor Output\n",
"\n",
"Using a HuggingFace processor to preprocess text and images, and passing the `processor_output` directly into `Engine.generate`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "9",
"metadata": {},
"outputs": [],
"source": [
"from transformers import AutoProcessor\n",
"\n",
"processor = AutoProcessor.from_pretrained(model_path, use_fast=True)\n",
"processor_output = processor(\n",
" images=[image], text=conv.get_prompt(), return_tensors=\"pt\"\n",
")\n",
"\n",
"out = llm.generate(\n",
" input_ids=processor_output[\"input_ids\"][0].detach().cpu().tolist(),\n",
" image_data=[dict(processor_output, format=\"processor_output\")],\n",
")\n",
"print(\"Response using processor output:\")\n",
"print(out[\"text\"])"
]
},
{
"cell_type": "markdown",
"id": "10",
"metadata": {},
"source": [
"### Call with Precomputed Embeddings\n",
"\n",
"You can pre-calculate image features to avoid repeated visual encoding processes."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "11",
"metadata": {},
"outputs": [],
"source": [
"from transformers import AutoProcessor\n",
"from transformers import Qwen2_5_VLForConditionalGeneration\n",
"\n",
@@ -122,53 +191,6 @@
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "9",
"metadata": {},
"outputs": [],
"source": [
"processed_prompt = processor(\n",
" images=[image], text=conv.get_prompt(), return_tensors=\"pt\"\n",
")\n",
"input_ids = processed_prompt[\"input_ids\"][0].detach().cpu().tolist()\n",
"precomputed_embeddings = vision(\n",
" processed_prompt[\"pixel_values\"].cuda(), processed_prompt[\"image_grid_thw\"].cuda()\n",
")\n",
"\n",
"mm_item = dict(\n",
" modality=\"IMAGE\",\n",
" image_grid_thw=processed_prompt[\"image_grid_thw\"],\n",
" precomputed_embeddings=precomputed_embeddings,\n",
")\n",
"out = llm.generate(input_ids=input_ids, image_data=[mm_item])\n",
"print(out[\"text\"])"
]
},
{
"cell_type": "markdown",
"id": "10",
"metadata": {},
"source": [
"## Querying Llama 4 (Vision)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "11",
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply() # Run this first.\n",
"\n",
"model_path = \"meta-llama/Llama-4-Scout-17B-16E-Instruct\"\n",
"chat_template = \"llama-4\""
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -176,7 +198,39 @@
"metadata": {},
"outputs": [],
"source": [
"# Lets create a prompt.\n",
"processor_output = processor(\n",
" images=[image], text=conv.get_prompt(), return_tensors=\"pt\"\n",
")\n",
"\n",
"input_ids = processor_output[\"input_ids\"][0].detach().cpu().tolist()\n",
"\n",
"precomputed_embeddings = vision(\n",
" processor_output[\"pixel_values\"].cuda(), processor_output[\"image_grid_thw\"].cuda()\n",
")\n",
"\n",
"multi_modal_item = dict(\n",
" processor_output,\n",
" format=\"precomputed_embedding\",\n",
" feature=precomputed_embeddings,\n",
")\n",
"\n",
"out = llm.generate(input_ids=input_ids, image_data=[multi_modal_item])\n",
"print(\"Response using precomputed embeddings:\")\n",
"print(out[\"text\"])\n",
"\n",
"llm.shutdown()"
]
},
{
"cell_type": "markdown",
"id": "13",
"metadata": {},
"source": [
"## Querying Llama 4 Vision Model\n",
"\n",
"```python\n",
"model_path = \"meta-llama/Llama-4-Scout-17B-16E-Instruct\"\n",
"chat_template = \"llama-4\"\n",
"\n",
"from io import BytesIO\n",
"import requests\n",
@@ -184,6 +238,7 @@
"\n",
"from sglang.srt.parser.conversation import chat_templates\n",
"\n",
"# Download the same example image\n",
"image = Image.open(\n",
" BytesIO(\n",
" requests.get(\n",
@@ -197,53 +252,62 @@
"conv.append_message(conv.roles[1], \"\")\n",
"conv.image_data = [image]\n",
"\n",
"print(\"Llama 4 generated prompt text:\")\n",
"print(conv.get_prompt())\n",
"print(f\"Image size: {image.size}\")\n",
"\n",
"image"
"image\n",
"```"
]
},
{
"cell_type": "markdown",
"id": "13",
"metadata": {},
"source": [
"### Query via the offline Engine API"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "14",
"metadata": {},
"outputs": [],
"source": [
"from sglang.test.test_utils import is_in_ci\n",
"### Llama 4 Basic Call\n",
"\n",
"if not is_in_ci():\n",
" from sglang import Engine\n",
"Llama 4 requires more computational resources, so it's configured with multi-GPU parallelism (tp_size=4) and larger context length.\n",
"\n",
" llm = Engine(\n",
" model_path=model_path,\n",
" trust_remote_code=True,\n",
" enable_multimodal=True,\n",
" mem_fraction_static=0.8,\n",
" tp_size=4,\n",
" attention_backend=\"fa3\",\n",
" context_length=65536,\n",
" )"
"```python\n",
"llm = Engine(\n",
" model_path=model_path,\n",
" enable_multimodal=True,\n",
" attention_backend=\"fa3\",\n",
" tp_size=4,\n",
" context_length=65536,\n",
")\n",
"\n",
"out = llm.generate(prompt=conv.get_prompt(), image_data=[image])\n",
"print(\"Llama 4 response:\")\n",
"print(out[\"text\"])\n",
"```"
]
},
{
"cell_type": "code",
"execution_count": null,
"cell_type": "markdown",
"id": "15",
"metadata": {},
"outputs": [],
"source": [
"if not is_in_ci():\n",
" out = llm.generate(prompt=conv.get_prompt(), image_data=[image])\n",
" print(out[\"text\"])"
"### Call with Processor Output\n",
"\n",
"Using HuggingFace processor to preprocess data can reduce computational overhead during inference.\n",
"\n",
"```python\n",
"from transformers import AutoProcessor\n",
"\n",
"processor = AutoProcessor.from_pretrained(model_path, use_fast=True)\n",
"processor_output = processor(\n",
" images=[image], text=conv.get_prompt(), return_tensors=\"pt\"\n",
")\n",
"\n",
"out = llm.generate(\n",
" input_ids=processor_output[\"input_ids\"][0].detach().cpu().tolist(),\n",
" image_data=[dict(processor_output, format=\"processor_output\")],\n",
")\n",
"print(\"Response using processor output:\")\n",
"print(out)\n",
"```"
]
},
{
@@ -251,54 +315,48 @@
"id": "16",
"metadata": {},
"source": [
"### Query via the offline Engine API, but send precomputed embeddings"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "17",
"metadata": {},
"outputs": [],
"source": [
"if not is_in_ci():\n",
" # Compute the image embeddings using Huggingface.\n",
"### Call with Precomputed Embeddings\n",
"\n",
" from transformers import AutoProcessor\n",
" from transformers import Llama4ForConditionalGeneration\n",
"```python\n",
"from transformers import AutoProcessor\n",
"from transformers import Llama4ForConditionalGeneration\n",
"\n",
" processor = AutoProcessor.from_pretrained(model_path, use_fast=True)\n",
" model = Llama4ForConditionalGeneration.from_pretrained(\n",
" model_path, torch_dtype=\"auto\"\n",
" ).eval()\n",
" vision = model.vision_model.cuda()\n",
" multi_modal_projector = model.multi_modal_projector.cuda()"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "18",
"metadata": {},
"outputs": [],
"source": [
"if not is_in_ci():\n",
" processed_prompt = processor(\n",
" images=[image], text=conv.get_prompt(), return_tensors=\"pt\"\n",
" )\n",
" print(f'{processed_prompt[\"pixel_values\"].shape=}')\n",
" input_ids = processed_prompt[\"input_ids\"][0].detach().cpu().tolist()\n",
"processor = AutoProcessor.from_pretrained(model_path, use_fast=True)\n",
"model = Llama4ForConditionalGeneration.from_pretrained(\n",
" model_path, torch_dtype=\"auto\"\n",
").eval()\n",
"\n",
" image_outputs = vision(\n",
" processed_prompt[\"pixel_values\"].to(\"cuda\"), output_hidden_states=False\n",
" )\n",
" image_features = image_outputs.last_hidden_state\n",
" vision_flat = image_features.view(-1, image_features.size(-1))\n",
" precomputed_embeddings = multi_modal_projector(vision_flat)\n",
"vision = model.vision_model.cuda()\n",
"multi_modal_projector = model.multi_modal_projector.cuda()\n",
"\n",
" mm_item = dict(modality=\"IMAGE\", precomputed_embeddings=precomputed_embeddings)\n",
" out = llm.generate(input_ids=input_ids, image_data=[mm_item])\n",
" print(out[\"text\"])"
"print(f'Image pixel values shape: {processor_output[\"pixel_values\"].shape}')\n",
"input_ids = processor_output[\"input_ids\"][0].detach().cpu().tolist()\n",
"\n",
"# Process image through vision encoder\n",
"image_outputs = vision(\n",
" processor_output[\"pixel_values\"].to(\"cuda\"), \n",
" aspect_ratio_ids=processor_output[\"aspect_ratio_ids\"].to(\"cuda\"),\n",
" aspect_ratio_mask=processor_output[\"aspect_ratio_mask\"].to(\"cuda\"),\n",
" output_hidden_states=False\n",
")\n",
"image_features = image_outputs.last_hidden_state\n",
"\n",
"# Flatten image features and pass through multimodal projector\n",
"vision_flat = image_features.view(-1, image_features.size(-1))\n",
"precomputed_embeddings = multi_modal_projector(vision_flat)\n",
"\n",
"# Build precomputed embedding data item\n",
"mm_item = dict(\n",
" processor_output, \n",
" format=\"precomputed_embedding\", \n",
" feature=precomputed_embeddings\n",
")\n",
"\n",
"# Use precomputed embeddings for efficient inference\n",
"out = llm.generate(input_ids=input_ids, image_data=[mm_item])\n",
"print(\"Llama 4 precomputed embedding response:\")\n",
"print(out[\"text\"])\n",
"```"
]
}
],
@@ -306,7 +364,13 @@
"jupytext": {
"cell_metadata_filter": "-all",
"custom_cell_magics": "kql",
"encoding": "# -*- coding: utf-8 -*-"
"encoding": "# -*- coding: utf-8 -*-",
"text_representation": {
"extension": ".py",
"format_name": "light",
"format_version": "1.5",
"jupytext_version": "1.16.1"
}
},
"language_info": {
"codemirror_mode": {
+1 -1
View File
@@ -12,7 +12,7 @@ The `/generate` endpoint accepts the following parameters in JSON format. For de
| text | `Optional[Union[List[str], str]] = None` | The input prompt. Can be a single prompt or a batch of prompts. |
| input_ids | `Optional[Union[List[List[int]], List[int]]] = None` | The token IDs for text; one can specify either text or input_ids. |
| input_embeds | `Optional[Union[List[List[List[float]]], List[List[float]]]] = None` | The embeddings for input_ids; one can specify either text, input_ids, or input_embeds. |
| image_data | `Optional[Union[List[List[ImageDataItem]], List[ImageDataItem], ImageDataItem]] = None` | The image input. Can be an image instance, file name, URL, or base64 encoded string. Can be a single image, list of images, or list of lists of images. |
| image_data | `Optional[Union[List[List[ImageDataItem]], List[ImageDataItem], ImageDataItem]] = None` | The image input. Supports three formats: (1) **Raw images**: PIL Image, file path, URL, or base64 string; (2) **Processor output**: Dict with `format: "processor_output"` containing HuggingFace processor outputs; (3) **Precomputed embeddings**: Dict with `format: "precomputed_embedding"` and `feature` containing pre-calculated visual embeddings. Can be a single image, list of images, or list of lists of images. See [Multimodal Input Formats](#multimodal-input-formats) for details. |
| audio_data | `Optional[Union[List[AudioDataItem], AudioDataItem]] = None` | The audio input. Can be a file name, URL, or base64 encoded string. |
| sampling_params | `Optional[Union[List[Dict], Dict]] = None` | The sampling parameters as described in the sections below. |
| rid | `Optional[Union[List[str], str]] = None` | The request ID. |