Last of four; stacked on #38048. The record is the operator's input; the bags are what is in effect. A reader that takes the record and reads a field off it gets the input, which is the wrong one of the two whenever resolution decided something -- and the mistake is silent, because for most fields and most launches the two agree. Several of these files already read both ways, sometimes in the same expression: ```python get_tokenizer( get_serving().tokenizer_path, tokenizer_mode=server_args.tokenizer_mode, # the input, not the decision ... ) ``` Sixty-odd files convert. Record field reads in runtime code go from 199 to 11. Nine parameters that the conversion emptied are dropped along with the argument at every call site -- the dead-parameter ratchet is what names them. ### "Runs after its process publishes" is a per-entry-point claim Most converted reads sit in the serving and model-executor layers, which only exist after publication, or in the two subprocess entry points, which publish first thing. Three places are not like that, and they keep reading the record they were handed: - **`HttpServerEngineAdapter`** launches the server as a *child*. The parent resolves the record and never publishes, so the adapter's own reads -- the launch banner, the API key in its readiness loop, the TP width in `update_weights_from_tensor` -- are of `self.server_args`. A bag read here fails closed in a bare process, or answers for an unrelated engine in one that happens to have published. - **`serve_grpc`** reads its sidecar port before the integrated servicer builds the `Engine` that publishes. The comment above that line already said so and already bound `cfg = resolving_view(server_args)` for it; the sidecar port and the port it derives from read `cfg`. - **`initialize_dp_attention`** runs from callers whose publish is not guaranteed, so its one predicate stays on the resolution view. `ROLE_NAMESPACE_SETS["dp_controller"]` gains `observability` and `serving`, because the controller's metrics gate, tracing setup and worker-port broadcast now read those namespaces. Under `SGLANG_ROLE_NAMESPACES=enforce` that set is what the process may read, so a conversion that reaches a new namespace has to widen it in the same change. ## Three things worth a reviewer's attention **Eleven reads were `getattr(record, "field", default)`.** An AST scan for attribute access does not see those, so the census that said "43 readers" was counting the shape it could match rather than the thing it was after. `incremental_streaming_output` was read that way twice, and the transcription tests were the only reason it surfaced. **Not every record read is a bag read waiting to happen.** A multimodal processor's `base_gpu_id` is the instance's, not the process's: two engines in one process keep different ones, and `test_publishing_another_config_does_not_move_the_device` exists to say so. It stays on the record while `rl_on_policy_target` beside it moves. `RequestMetricsExporter` is the same shape -- it is handed the directory it writes to, and a test builds several with different ones. `configure_logger` is a third: 17 call sites, one of which passes an `argparse.Namespace`, so it is not a global-context reader at all. Those eleven remaining reads are the ones with a reason. **The fixtures move with the code.** Tests that hung config off a mock manager now publish a record, which is what the serving layer reads; where a test states a value it says so with `override_server_args` instead of assigning through the mock. `test_hisparse_unit` is the last of them: it stubbed a `server_args` onto a fake scheduler to say the decode radix cache was off, and the value it was standing in for is the published default, so the stub goes and the class publishes. ## Two things CI caught that a local sweep could not **`unittest.TestCase.enterContext` is Python 3.11+.** The converted fixtures used it at 18 sites; `requires-python` is `>=3.10` and CI runs 3.10, so every one of them raised `AttributeError` there while passing on a newer local interpreter. They call `enter_override(self, ...)` now -- a four-line helper in `sglang/test/test_utils.py` over the override's own `install()` / `restore()`. **A batched sweep cannot see a missing publish.** Three fixtures needed a published config and did not have one; each *passed* inside a shard where some other file had published, and failed when run alone. The affected cases are `test_serving_completions` (which set `incremental_streaming_output` on the mock manager's record, where nothing reads it now), `test_qwen3_vl_feature_materialization` (same shape for `mm_enable_dp_encoder`), and the two Qwen Rust tests -- whose fixture already carried the comment `# Non-auto: get_resolved_model_impl would choke on a SimpleNamespace` next to the `model_impl` it sets, which is exactly what happened once `get_mm_processor_cls` started reading that value from the bag. Its `publish` mirrors `model_impl` now, like the four fields it already mirrored. ## Verification A full registered-unit sweep (648 files) against this stack's merge-base: 19 failures on both sides, the same 19, none of them config. That sweep is what caught 23 failures the file-scoped runs missed -- and, later, that the narrower 139-file list did not even contain the files this change reaches. It is also what caught the `test_hisparse_unit` fixture above: the file passes inside a shard where something else published, and fails when it is run on its own, which is why every failing file is re-run alone before it is counted.
291 lines
11 KiB
Python
291 lines
11 KiB
Python
from __future__ import annotations
|
|
|
|
import json
|
|
import logging
|
|
import uuid
|
|
from abc import ABC, abstractmethod
|
|
from typing import TYPE_CHECKING, Any, List, Optional, Tuple, Union
|
|
|
|
import orjson
|
|
from fastapi import HTTPException, Request
|
|
from fastapi.responses import ORJSONResponse, StreamingResponse
|
|
|
|
from sglang.srt.entrypoints.openai.encoding_dsv32 import DS32EncodingError
|
|
from sglang.srt.entrypoints.openai.protocol import ErrorResponse, OpenAIServingRequest
|
|
from sglang.srt.managers.io_struct import EmbeddingReqInput, GenerateReqInput
|
|
from sglang.srt.observability.req_time_stats import monotonic_time
|
|
from sglang.srt.runtime_context import get_observability
|
|
from sglang.srt.server_args import ServerArgs
|
|
|
|
if TYPE_CHECKING:
|
|
from sglang.srt.managers.tokenizer_manager import TokenizerManager
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
# Base class for specific endpoint handlers
|
|
class OpenAIServingBase(ABC):
|
|
"""Abstract base class for OpenAI endpoint handlers"""
|
|
|
|
def __init__(self, tokenizer_manager: TokenizerManager):
|
|
self.tokenizer_manager = tokenizer_manager
|
|
self.allowed_custom_labels = (
|
|
set(get_observability().tokenizer_metrics_allowed_custom_labels)
|
|
if isinstance(self.tokenizer_manager.server_args, ServerArgs)
|
|
and get_observability().tokenizer_metrics_allowed_custom_labels
|
|
else None
|
|
)
|
|
|
|
def _parse_model_parameter(self, model: str) -> Tuple[str, Optional[str]]:
|
|
"""Parse 'base-model:adapter-name' syntax to extract LoRA adapter.
|
|
|
|
Returns (base_model, adapter_name) or (model, None) if no colon present.
|
|
"""
|
|
if ":" not in model:
|
|
return model, None
|
|
|
|
# Split on first colon only to handle model paths with multiple colons
|
|
parts = model.split(":", 1)
|
|
base_model = parts[0].strip()
|
|
adapter_name = parts[1].strip() or None
|
|
|
|
return base_model, adapter_name
|
|
|
|
def _resolve_lora_path(
|
|
self,
|
|
request_model: str,
|
|
explicit_lora_path: Optional[Union[str, List[Optional[str]]]],
|
|
) -> Optional[Union[str, List[Optional[str]]]]:
|
|
"""Resolve LoRA adapter with priority: model parameter > explicit lora_path.
|
|
|
|
Returns adapter name or None. Supports both single values and lists (batches).
|
|
"""
|
|
_, adapter_from_model = self._parse_model_parameter(request_model)
|
|
|
|
# Model parameter adapter takes precedence
|
|
if adapter_from_model is not None:
|
|
return adapter_from_model
|
|
|
|
# Fall back to explicit lora_path
|
|
return explicit_lora_path
|
|
|
|
async def handle_request(
|
|
self, request: OpenAIServingRequest, raw_request: Request
|
|
) -> Union[Any, StreamingResponse, ErrorResponse]:
|
|
"""Handle the specific request type with common pattern
|
|
If you want to override this method, you should be careful to record the validation time.
|
|
"""
|
|
received_time = monotonic_time()
|
|
|
|
try:
|
|
# Validate request
|
|
error_msg = self._validate_request(request)
|
|
if error_msg:
|
|
return self.create_error_response(error_msg)
|
|
|
|
# Log the raw OpenAI request payload before conversion to tokenized form.
|
|
request_logger = self.tokenizer_manager.request_logger
|
|
if request_logger.log_requests and request_logger.log_requests_level >= 2:
|
|
request_logger.log_openai_received_request(request, request=raw_request)
|
|
|
|
# Convert to internal format
|
|
adapted_request, processed_request = self._convert_to_internal_request(
|
|
request, raw_request
|
|
)
|
|
|
|
if isinstance(adapted_request, (GenerateReqInput, EmbeddingReqInput)):
|
|
# Only set timing fields if adapted_request supports them
|
|
adapted_request.received_time = received_time
|
|
|
|
# Note(Xinyuan): raw_request below is only used for detecting the connection of the client
|
|
if hasattr(request, "stream") and request.stream:
|
|
return await self._handle_streaming_request(
|
|
adapted_request, processed_request, raw_request
|
|
)
|
|
else:
|
|
return await self._handle_non_streaming_request(
|
|
adapted_request, processed_request, raw_request
|
|
)
|
|
except HTTPException as e:
|
|
return self.create_error_response(
|
|
message=e.detail, err_type=str(e.status_code), status_code=e.status_code
|
|
)
|
|
except ValueError as e:
|
|
return self.create_error_response(
|
|
message=str(e),
|
|
err_type="BadRequest",
|
|
status_code=400,
|
|
)
|
|
except DS32EncodingError as e:
|
|
logger.info(f"DS32EncodingError: {e}")
|
|
return self.create_error_response(
|
|
message=str(e),
|
|
err_type="BadRequest",
|
|
status_code=400,
|
|
)
|
|
except Exception as e:
|
|
logger.exception(f"Error in request: {e}")
|
|
return self.create_error_response(
|
|
message=f"Internal server error: {str(e)}",
|
|
err_type="InternalServerError",
|
|
status_code=500,
|
|
)
|
|
|
|
@abstractmethod
|
|
def _request_id_prefix(self) -> str:
|
|
"""Generate request ID based on request type"""
|
|
pass
|
|
|
|
def _generate_request_id_base(self, request: OpenAIServingRequest) -> Optional[str]:
|
|
"""Generate request ID based on request type"""
|
|
return None
|
|
|
|
# TODO(chang): the rid is used in io_strcut check and often violates `The rid should be a list` AssertionError
|
|
# Temporarily return None in this function until the rid logic is clear.
|
|
if rid := getattr(request, "rid", None):
|
|
return rid
|
|
|
|
return f"{self._request_id_prefix()}{uuid.uuid4().hex}"
|
|
|
|
@abstractmethod
|
|
def _convert_to_internal_request(
|
|
self,
|
|
request: OpenAIServingRequest,
|
|
raw_request: Request = None,
|
|
) -> tuple[GenerateReqInput, OpenAIServingRequest]:
|
|
"""Convert OpenAI request to internal format"""
|
|
pass
|
|
|
|
async def _handle_streaming_request(
|
|
self,
|
|
adapted_request: GenerateReqInput,
|
|
request: OpenAIServingRequest,
|
|
raw_request: Request,
|
|
) -> Union[StreamingResponse, ErrorResponse, ORJSONResponse]:
|
|
"""Handle streaming request
|
|
|
|
Override this method in child classes that support streaming requests.
|
|
"""
|
|
return self.create_error_response(
|
|
message=f"{self.__class__.__name__} does not support streaming requests",
|
|
err_type="NotImplementedError",
|
|
status_code=501,
|
|
)
|
|
|
|
async def _handle_non_streaming_request(
|
|
self,
|
|
adapted_request: GenerateReqInput,
|
|
request: OpenAIServingRequest,
|
|
raw_request: Request,
|
|
) -> Union[Any, ErrorResponse, ORJSONResponse]:
|
|
"""Handle non-streaming request
|
|
|
|
Override this method in child classes that support non-streaming requests.
|
|
"""
|
|
return self.create_error_response(
|
|
message=f"{self.__class__.__name__} does not support non-streaming requests",
|
|
err_type="NotImplementedError",
|
|
status_code=501,
|
|
)
|
|
|
|
def _validate_request(self, _: OpenAIServingRequest) -> Optional[str]:
|
|
"""Validate request"""
|
|
pass
|
|
|
|
def create_error_response(
|
|
self,
|
|
message: str,
|
|
err_type: str = "BadRequestError",
|
|
status_code: int = 400,
|
|
param: Optional[str] = None,
|
|
) -> ORJSONResponse:
|
|
"""Create an error response"""
|
|
# TODO: remove fastapi dependency in openai and move response handling to the entrypoint
|
|
error = ErrorResponse(
|
|
object="error",
|
|
message=message,
|
|
type=err_type,
|
|
param=param,
|
|
code=status_code,
|
|
)
|
|
return ORJSONResponse(content=error.model_dump(), status_code=status_code)
|
|
|
|
def create_streaming_error_response(
|
|
self,
|
|
message: str,
|
|
err_type: str = "BadRequestError",
|
|
status_code: int = 400,
|
|
) -> str:
|
|
"""Create a streaming error response"""
|
|
error = ErrorResponse(
|
|
object="error",
|
|
message=message,
|
|
type=err_type,
|
|
param=None,
|
|
code=status_code,
|
|
)
|
|
return json.dumps({"error": error.model_dump()})
|
|
|
|
def extract_custom_labels(self, raw_request):
|
|
if (
|
|
not self.allowed_custom_labels
|
|
or not get_observability().tokenizer_metrics_custom_labels_header
|
|
):
|
|
return None
|
|
|
|
custom_labels = None
|
|
header = get_observability().tokenizer_metrics_custom_labels_header
|
|
try:
|
|
raw_labels = (
|
|
orjson.loads(raw_request.headers.get(header))
|
|
if raw_request and raw_request.headers.get(header)
|
|
else None
|
|
)
|
|
except json.JSONDecodeError as e:
|
|
logger.exception(f"Error in request: {e}")
|
|
raw_labels = None
|
|
|
|
if isinstance(raw_labels, dict):
|
|
custom_labels = {
|
|
label: value
|
|
for label, value in raw_labels.items()
|
|
if label in self.allowed_custom_labels
|
|
}
|
|
return custom_labels
|
|
|
|
def extract_routing_key(self, raw_request):
|
|
if raw_request is None:
|
|
return None
|
|
return raw_request.headers.get("x-smg-routing-key")
|
|
|
|
def extract_routed_dp_rank_from_header(
|
|
self, raw_request: Request, body_routed_dp_rank: Optional[int] = None
|
|
) -> Optional[int]:
|
|
"""Extract routed_dp_rank from HTTP header, with higher priority than routed_dp_rank in body.
|
|
|
|
Header name: X-Data-Parallel-Rank (case-insensitive in HTTP/1.1/2)
|
|
"""
|
|
if raw_request is None:
|
|
return body_routed_dp_rank
|
|
|
|
header_value = raw_request.headers.get("x-data-parallel-rank")
|
|
if header_value is not None:
|
|
try:
|
|
header_dp_rank = int(header_value)
|
|
if (
|
|
body_routed_dp_rank is not None
|
|
and header_dp_rank != body_routed_dp_rank
|
|
):
|
|
logger.debug(
|
|
f"X-Data-Parallel-Rank header ({header_dp_rank}) overrides "
|
|
f"body routed_dp_rank ({body_routed_dp_rank})"
|
|
)
|
|
return header_dp_rank
|
|
except ValueError:
|
|
raise HTTPException(
|
|
status_code=400,
|
|
detail=f"Invalid X-Data-Parallel-Rank header: must be an integer, got '{header_value}'",
|
|
)
|
|
|
|
return body_routed_dp_rank
|