[Config] Round 6.4: the runtime reads the bags, not the record (#38049)
Last of four; stacked on #38048. The record is the operator's input; the bags are what is in effect. A reader that takes the record and reads a field off it gets the input, which is the wrong one of the two whenever resolution decided something -- and the mistake is silent, because for most fields and most launches the two agree. Several of these files already read both ways, sometimes in the same expression: ```python get_tokenizer( get_serving().tokenizer_path, tokenizer_mode=server_args.tokenizer_mode, # the input, not the decision ... ) ``` Sixty-odd files convert. Record field reads in runtime code go from 199 to 11. Nine parameters that the conversion emptied are dropped along with the argument at every call site -- the dead-parameter ratchet is what names them. ### "Runs after its process publishes" is a per-entry-point claim Most converted reads sit in the serving and model-executor layers, which only exist after publication, or in the two subprocess entry points, which publish first thing. Three places are not like that, and they keep reading the record they were handed: - **`HttpServerEngineAdapter`** launches the server as a *child*. The parent resolves the record and never publishes, so the adapter's own reads -- the launch banner, the API key in its readiness loop, the TP width in `update_weights_from_tensor` -- are of `self.server_args`. A bag read here fails closed in a bare process, or answers for an unrelated engine in one that happens to have published. - **`serve_grpc`** reads its sidecar port before the integrated servicer builds the `Engine` that publishes. The comment above that line already said so and already bound `cfg = resolving_view(server_args)` for it; the sidecar port and the port it derives from read `cfg`. - **`initialize_dp_attention`** runs from callers whose publish is not guaranteed, so its one predicate stays on the resolution view. `ROLE_NAMESPACE_SETS["dp_controller"]` gains `observability` and `serving`, because the controller's metrics gate, tracing setup and worker-port broadcast now read those namespaces. Under `SGLANG_ROLE_NAMESPACES=enforce` that set is what the process may read, so a conversion that reaches a new namespace has to widen it in the same change. ## Three things worth a reviewer's attention **Eleven reads were `getattr(record, "field", default)`.** An AST scan for attribute access does not see those, so the census that said "43 readers" was counting the shape it could match rather than the thing it was after. `incremental_streaming_output` was read that way twice, and the transcription tests were the only reason it surfaced. **Not every record read is a bag read waiting to happen.** A multimodal processor's `base_gpu_id` is the instance's, not the process's: two engines in one process keep different ones, and `test_publishing_another_config_does_not_move_the_device` exists to say so. It stays on the record while `rl_on_policy_target` beside it moves. `RequestMetricsExporter` is the same shape -- it is handed the directory it writes to, and a test builds several with different ones. `configure_logger` is a third: 17 call sites, one of which passes an `argparse.Namespace`, so it is not a global-context reader at all. Those eleven remaining reads are the ones with a reason. **The fixtures move with the code.** Tests that hung config off a mock manager now publish a record, which is what the serving layer reads; where a test states a value it says so with `override_server_args` instead of assigning through the mock. `test_hisparse_unit` is the last of them: it stubbed a `server_args` onto a fake scheduler to say the decode radix cache was off, and the value it was standing in for is the published default, so the stub goes and the class publishes. ## Two things CI caught that a local sweep could not **`unittest.TestCase.enterContext` is Python 3.11+.** The converted fixtures used it at 18 sites; `requires-python` is `>=3.10` and CI runs 3.10, so every one of them raised `AttributeError` there while passing on a newer local interpreter. They call `enter_override(self, ...)` now -- a four-line helper in `sglang/test/test_utils.py` over the override's own `install()` / `restore()`. **A batched sweep cannot see a missing publish.** Three fixtures needed a published config and did not have one; each *passed* inside a shard where some other file had published, and failed when run alone. The affected cases are `test_serving_completions` (which set `incremental_streaming_output` on the mock manager's record, where nothing reads it now), `test_qwen3_vl_feature_materialization` (same shape for `mm_enable_dp_encoder`), and the two Qwen Rust tests -- whose fixture already carried the comment `# Non-auto: get_resolved_model_impl would choke on a SimpleNamespace` next to the `model_impl` it sets, which is exactly what happened once `get_mm_processor_cls` started reading that value from the bag. Its `publish` mirrors `model_impl` now, like the four fields it already mirrored. ## Verification A full registered-unit sweep (648 files) against this stack's merge-base: 19 failures on both sides, the same 19, none of them config. That sweep is what caught 23 failures the file-scoped runs missed -- and, later, that the narrower 139-file list did not even contain the files this change reaches. It is also what caught the `test_hisparse_unit` fixture above: the file passes inside a shard where something else published, and fails when it is run on its own, which is why every failing file is re-run alone before it is counted.
This commit is contained in:
@@ -9,6 +9,7 @@ import unittest
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
from sglang.srt.entrypoints.openai.serving_base import OpenAIServingBase
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_amd_ci, register_cpu_ci
|
||||
|
||||
@@ -16,13 +17,22 @@ register_amd_ci(est_time=30, suite="nightly-amd-1-gpu", nightly=True)
|
||||
register_cpu_ci(est_time=7, suite="stage-b-test-cpu-intel")
|
||||
|
||||
|
||||
def publish_config(case):
|
||||
"""`OpenAIServingBase.__init__` reads `get_observability()`, so a case that
|
||||
builds one needs a published config. The mock manager cannot stand in for
|
||||
it: `MagicMock(spec=ServerArgs)` passes the `isinstance` guard, so the read
|
||||
happens and there is no bag to answer from."""
|
||||
reset_context()
|
||||
case.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
|
||||
class MockTokenizerManager:
|
||||
"""Mock TokenizerManager for testing."""
|
||||
|
||||
def __init__(self, enable_lora=False):
|
||||
self.server_args = MagicMock(spec=ServerArgs)
|
||||
self.server_args.enable_lora = enable_lora
|
||||
self.server_args.tokenizer_metrics_allowed_custom_labels = None
|
||||
|
||||
|
||||
class ConcreteServingBase(OpenAIServingBase):
|
||||
@@ -42,6 +52,7 @@ class TestParseModelParameter(unittest.TestCase):
|
||||
"""Test _parse_model_parameter method."""
|
||||
|
||||
def setUp(self):
|
||||
publish_config(self)
|
||||
self.tokenizer_manager = MockTokenizerManager(enable_lora=True)
|
||||
self.serving = ConcreteServingBase(self.tokenizer_manager)
|
||||
|
||||
@@ -98,6 +109,7 @@ class TestResolveLoraPath(unittest.TestCase):
|
||||
"""Test _resolve_lora_path method."""
|
||||
|
||||
def setUp(self):
|
||||
publish_config(self)
|
||||
self.tokenizer_manager = MockTokenizerManager(enable_lora=True)
|
||||
self.serving = ConcreteServingBase(self.tokenizer_manager)
|
||||
|
||||
@@ -146,6 +158,7 @@ class TestIntegrationScenarios(unittest.TestCase):
|
||||
"""Integration tests for common usage scenarios."""
|
||||
|
||||
def setUp(self):
|
||||
publish_config(self)
|
||||
self.tokenizer_manager = MockTokenizerManager(enable_lora=True)
|
||||
self.serving = ConcreteServingBase(self.tokenizer_manager)
|
||||
|
||||
@@ -197,6 +210,7 @@ class TestEdgeCases(unittest.TestCase):
|
||||
"""Test edge cases and error conditions."""
|
||||
|
||||
def setUp(self):
|
||||
publish_config(self)
|
||||
self.tokenizer_manager = MockTokenizerManager(enable_lora=True)
|
||||
self.serving = ConcreteServingBase(self.tokenizer_manager)
|
||||
|
||||
|
||||
@@ -14,7 +14,8 @@ from sglang.srt.disaggregation.utils import DisaggregationMode
|
||||
from sglang.srt.distributed.parallel_state_wrapper import ParallelState
|
||||
from sglang.srt.managers.schedule_batch import FINISH_ABORT
|
||||
from sglang.srt.managers.scheduler import Scheduler
|
||||
from sglang.srt.runtime_context import get_context
|
||||
from sglang.srt.runtime_context import get_context, publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -34,6 +35,12 @@ class FakeReceiver:
|
||||
|
||||
|
||||
class TestDecodeQueueCleanup(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_paged_swa_retraction_resume_uses_physical_page_budget(self):
|
||||
# resume_retracted_reqs reads the retraction backend off the disagg
|
||||
# bag, so the case publishes a config instead of injecting one.
|
||||
|
||||
@@ -43,9 +43,11 @@ from sglang.srt.parser.jinja_template_utils import (
|
||||
jinja_template_may_reorder_tool_results,
|
||||
)
|
||||
from sglang.srt.parser.template_detection import ReasoningToggleConfig
|
||||
from sglang.srt.runtime_context import get_context, publish, reset_context
|
||||
from sglang.srt.sampling.sampling_params import (
|
||||
REQUEST_REASONING_END_TOKEN_IDS_KEY,
|
||||
)
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.srt.utils import get_or_create_event_loop
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
|
||||
@@ -196,6 +198,22 @@ class _MockTemplateManager:
|
||||
class ServingChatTestCase(unittest.TestCase):
|
||||
# ------------- common fixtures -------------
|
||||
def setUp(self):
|
||||
# The serving layer reads its config from the bags, so the fixture has
|
||||
# to publish one rather than hang the values off a mock manager.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(
|
||||
ServerArgs(
|
||||
model_path="dummy",
|
||||
revision=None,
|
||||
enable_cache_report=False,
|
||||
tool_call_parser="hermes",
|
||||
reasoning_parser=None,
|
||||
stream_response_default_include_usage=False,
|
||||
default_chat_template_kwargs=None,
|
||||
),
|
||||
role="tokenizer",
|
||||
)
|
||||
self.tm = _MockTokenizerManager()
|
||||
self.template_manager = _MockTemplateManager()
|
||||
self.chat = OpenAIServingChat(self.tm, self.template_manager)
|
||||
@@ -3643,14 +3661,14 @@ class ServingChatTestCase(unittest.TestCase):
|
||||
|
||||
def test_continuous_usage_reports_cached_tokens(self):
|
||||
"""continuous_usage_stats chunks include cached tokens when cache reporting is on."""
|
||||
self.tm.server_args.enable_cache_report = True
|
||||
self.enterContext(get_context().override_server_args(enable_cache_report=True))
|
||||
usages = self._collect_continuous_usage(cached_tokens=6)
|
||||
self.assertTrue(usages, "continuous_usage_stats attached no usage")
|
||||
self.assertEqual(usages[0]["prompt_tokens_details"]["cached_tokens"], 6)
|
||||
|
||||
def test_continuous_usage_omits_cached_tokens_when_report_disabled(self):
|
||||
"""With cache reporting off, continuous_usage_stats must not leak cached tokens."""
|
||||
self.tm.server_args.enable_cache_report = False
|
||||
self.enterContext(get_context().override_server_args(enable_cache_report=False))
|
||||
usages = self._collect_continuous_usage(cached_tokens=6)
|
||||
self.assertTrue(usages, "continuous_usage_stats attached no usage")
|
||||
self.assertIsNone(usages[0].get("prompt_tokens_details"))
|
||||
@@ -3666,7 +3684,9 @@ class ServingChatTestCase(unittest.TestCase):
|
||||
Regression test for https://github.com/sgl-project/sglang/issues/22510.
|
||||
"""
|
||||
# Enable incremental_streaming_output on the mock
|
||||
self.tm.server_args.incremental_streaming_output = True
|
||||
self.enterContext(
|
||||
get_context().override_server_args(incremental_streaming_output=True)
|
||||
)
|
||||
|
||||
# Simulate incremental streaming: each yield has ONLY the new text (delta),
|
||||
# NOT the full accumulated text.
|
||||
@@ -4146,6 +4166,9 @@ class TestProcessToolCallsWithRequiredToolChoice(unittest.TestCase):
|
||||
"""Test _process_tool_calls with tool_choice='required' uses model-specific parser."""
|
||||
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
tm = _MockTokenizerManager()
|
||||
tm.server_args.tool_call_parser = "kimi_k2"
|
||||
self.chat = OpenAIServingChat(tm, _MockTemplateManager())
|
||||
|
||||
@@ -19,6 +19,8 @@ from fastapi import Request
|
||||
from sglang.srt.entrypoints.openai.protocol import CompletionRequest
|
||||
from sglang.srt.entrypoints.openai.serving_completions import OpenAIServingCompletion
|
||||
from sglang.srt.managers.tokenizer_manager import TokenizerManager
|
||||
from sglang.srt.runtime_context import get_context, publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.srt.utils import get_or_create_event_loop
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
|
||||
@@ -66,6 +68,9 @@ class ServingCompletionTestCase(unittest.TestCase):
|
||||
|
||||
# ---------- shared test fixtures ----------
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
# build the mock TokenizerManager once for every test
|
||||
tm = Mock(spec=TokenizerManager)
|
||||
|
||||
@@ -322,15 +327,18 @@ class ServingCompletionTestCase(unittest.TestCase):
|
||||
return_token_ids=True,
|
||||
)
|
||||
adapted_request, _ = self.sc._convert_to_internal_request(req)
|
||||
self.sc.tokenizer_manager.server_args.stream_response_default_include_usage = (
|
||||
False
|
||||
)
|
||||
|
||||
for incremental in (False, True):
|
||||
with self.subTest(incremental_streaming_output=incremental):
|
||||
self.sc.tokenizer_manager.server_args.incremental_streaming_output = (
|
||||
incremental
|
||||
)
|
||||
# Both of these are read through `get_serving()` now, so assigning
|
||||
# them on the mock manager's record has no effect on what the code
|
||||
# under test sees. State them where the code reads them.
|
||||
with (
|
||||
self.subTest(incremental_streaming_output=incremental),
|
||||
get_context().override_server_args(
|
||||
stream_response_default_include_usage=False,
|
||||
incremental_streaming_output=incremental,
|
||||
),
|
||||
):
|
||||
texts = ("a", "b", "c") if incremental else ("a", "ab", "abc")
|
||||
output_ids = (
|
||||
([5], [6], [7]) if incremental else ([5], [5, 6], [5, 6, 7])
|
||||
|
||||
@@ -23,9 +23,11 @@ from sglang.srt.entrypoints.openai.serving_responses import (
|
||||
)
|
||||
from sglang.srt.function_call.core_types import ToolCallItem
|
||||
from sglang.srt.parser.template_detection import ReasoningToggleConfig
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.sampling.sampling_params import (
|
||||
REQUEST_REASONING_END_TOKEN_IDS_KEY,
|
||||
)
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -625,6 +627,9 @@ class OutputItemsTestCase(CustomTestCase):
|
||||
def setUp(self):
|
||||
# qwen3_coder is the default for this class; the one no-native-parser
|
||||
# case overrides it.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
self.serving = make_serving()
|
||||
self.serving.tool_call_parser = "qwen3_coder"
|
||||
|
||||
|
||||
@@ -11,6 +11,8 @@ from utils import (
|
||||
)
|
||||
|
||||
from sglang.srt.entrypoints.openai.protocol import ResponsesRequest
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -246,6 +248,10 @@ class MultiToolCallStreamingOrderTestCase(CustomTestCase):
|
||||
def setUp(self):
|
||||
from sglang.srt.function_call.qwen3_coder_detector import Qwen3CoderDetector
|
||||
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
self.serving = make_serving()
|
||||
self.serving.tool_call_parser = "qwen3_coder"
|
||||
self.serving.reasoning_parser = None
|
||||
|
||||
@@ -10,7 +10,7 @@ The tests mock ``TokenizerManager.generate_request`` to yield synthetic
|
||||
``text`` chunks for each of the happy, abort, and boundary cases.
|
||||
"""
|
||||
|
||||
from sglang.test.test_utils import maybe_stub_sgl_kernel
|
||||
from sglang.test.test_utils import enter_override, maybe_stub_sgl_kernel
|
||||
|
||||
maybe_stub_sgl_kernel() # must precede any import that pulls in sgl_kernel
|
||||
|
||||
@@ -32,6 +32,8 @@ from sglang.srt.entrypoints.openai.serving_transcription import (
|
||||
OpenAIServingTranscription,
|
||||
)
|
||||
from sglang.srt.managers.io_struct import GenerateReqInput
|
||||
from sglang.srt.runtime_context import get_context, publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.srt.utils import get_or_create_event_loop
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
@@ -102,6 +104,16 @@ def _deltas_from_sse(sse_lines: List[str]) -> List[str]:
|
||||
class TestStreamingFusedAutodetect(CustomTestCase):
|
||||
"""_generate_transcription_stream with _fused_autodetect=True."""
|
||||
|
||||
def setUp(self):
|
||||
# The transcription serving layer reads its config from the bags, so
|
||||
# the fixture publishes one instead of hanging values off a mock.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(
|
||||
ServerArgs(model_path="dummy", asr_max_concurrent_sessions=32),
|
||||
role="tokenizer",
|
||||
)
|
||||
|
||||
def _run_stream(
|
||||
self, chunks: List[dict], fused: bool = True, ts_variant: bool = False
|
||||
):
|
||||
@@ -318,6 +330,16 @@ class TestLongAudioChunkedNonStreaming(CustomTestCase):
|
||||
requests and the transcripts stitched in order — without chunking the
|
||||
feature extractor silently truncates everything past 30 s."""
|
||||
|
||||
def setUp(self):
|
||||
# The transcription serving layer reads its config from the bags, so
|
||||
# the fixture publishes one instead of hanging values off a mock.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(
|
||||
ServerArgs(model_path="dummy", asr_max_concurrent_sessions=32),
|
||||
role="tokenizer",
|
||||
)
|
||||
|
||||
def _create_transcription(self, tm, audio_bytes, language="en", **kwargs):
|
||||
serving = OpenAIServingTranscription(tm)
|
||||
loop = get_or_create_event_loop()
|
||||
@@ -510,6 +532,16 @@ class TestLongAudioChunkedStreaming(CustomTestCase):
|
||||
"""_generate_long_audio_stream: chunks transcribed sequentially, deltas
|
||||
emitted in audio order, exactly one finish frame."""
|
||||
|
||||
def setUp(self):
|
||||
# The transcription serving layer reads its config from the bags, so
|
||||
# the fixture publishes one instead of hanging values off a mock.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(
|
||||
ServerArgs(model_path="dummy", asr_max_concurrent_sessions=32),
|
||||
role="tokenizer",
|
||||
)
|
||||
|
||||
def _run_stream(self, results_per_request, fused=False, n_chunks=2):
|
||||
tm = _MockChunkTokenizerManager(results_per_request)
|
||||
serving = OpenAIServingTranscription(tm)
|
||||
@@ -679,6 +711,16 @@ class TestStreamingIncrementalOutputMode(CustomTestCase):
|
||||
server already sent as a delta.
|
||||
"""
|
||||
|
||||
def setUp(self):
|
||||
# The transcription serving layer reads its config from the bags, so
|
||||
# the fixture publishes one instead of hanging values off a mock.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(
|
||||
ServerArgs(model_path="dummy", asr_max_concurrent_sessions=32),
|
||||
role="tokenizer",
|
||||
)
|
||||
|
||||
def _run_incremental_stream(self, chunk_deltas, fused=False):
|
||||
"""Server in incremental mode: yield per-chunk delta, not cumulative."""
|
||||
chunks = [
|
||||
@@ -686,9 +728,8 @@ class TestStreamingIncrementalOutputMode(CustomTestCase):
|
||||
for i, d in enumerate(chunk_deltas)
|
||||
]
|
||||
tm = _MockTokenizerManager(chunks)
|
||||
tm.server_args = Mock(
|
||||
incremental_streaming_output=True,
|
||||
asr_max_concurrent_sessions=32,
|
||||
enter_override(
|
||||
self, get_context().override_server_args(incremental_streaming_output=True)
|
||||
)
|
||||
serving = OpenAIServingTranscription(tm)
|
||||
|
||||
|
||||
@@ -27,6 +27,8 @@ from unittest.mock import Mock
|
||||
|
||||
from sglang.srt.entrypoints.openai.protocol import RequestResponseMetadata
|
||||
from sglang.srt.entrypoints.openai.serving_responses import OpenAIServingResponses
|
||||
from sglang.srt.runtime_context import get_context, publish
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
|
||||
register_cpu_ci(
|
||||
@@ -83,6 +85,11 @@ class MockTemplateManager:
|
||||
|
||||
|
||||
def make_serving(*, is_multimodal: bool = False) -> OpenAIServingResponses:
|
||||
"""The serving layer reads its config from the bags, so the fixture
|
||||
publishes one. Idempotent: a caller that already published keeps its own,
|
||||
which is how a test states a value the default record does not carry."""
|
||||
if not get_context().is_config_namespace_published("serving"):
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
return OpenAIServingResponses(
|
||||
MockTokenizerManager(is_multimodal=is_multimodal), MockTemplateManager()
|
||||
)
|
||||
|
||||
@@ -11,7 +11,13 @@ from types import SimpleNamespace
|
||||
from sglang.srt.arg_groups.layernorm_sp_hook import validate_layernorm_sp
|
||||
from sglang.srt.layers import layernorm_sp
|
||||
from sglang.srt.model_executor.forward_batch_info import ForwardMode
|
||||
from sglang.srt.runtime_context import get_flags, get_forward, reset_context
|
||||
from sglang.srt.runtime_context import (
|
||||
get_flags,
|
||||
get_forward,
|
||||
publish,
|
||||
reset_context,
|
||||
)
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -19,8 +25,10 @@ register_cpu_ci(est_time=9, suite="base-a-test-cpu")
|
||||
|
||||
|
||||
def _initialize(*, enable=True, arch="Qwen3ForCausalLM"):
|
||||
publish(
|
||||
ServerArgs(model_path="dummy", enable_layernorm_sp=enable), role="tokenizer"
|
||||
)
|
||||
layernorm_sp.initialize_layernorm_sp(
|
||||
server_args=SimpleNamespace(enable_layernorm_sp=enable),
|
||||
model_config=SimpleNamespace(
|
||||
hf_config=SimpleNamespace(architectures=[arch] if arch else [])
|
||||
),
|
||||
|
||||
@@ -21,6 +21,8 @@ from sglang.srt.managers.tokenizer_manager import TokenizerManager
|
||||
from sglang.srt.managers.tokenizer_manager_score_mixin import (
|
||||
TokenizerManagerScoreMixin,
|
||||
)
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -45,6 +47,12 @@ def _vec2d(val: float = 1.0) -> torch.Tensor:
|
||||
|
||||
|
||||
class TestPositionalEmbeds(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_from_list_of_1d_tensors(self):
|
||||
pe = PositionalEmbeds(embeds=[_vec(1), _vec(2)], positions=[0, 5])
|
||||
self.assertEqual(pe.embeds.shape, (2, HIDDEN_DIM))
|
||||
@@ -75,6 +83,12 @@ class TestPositionalEmbeds(CustomTestCase):
|
||||
|
||||
|
||||
class TestConvertEmbedsToTensors(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_none_returns_none(self):
|
||||
self.assertIsNone(convert_embeds_to_tensors(None))
|
||||
|
||||
@@ -110,6 +124,12 @@ class TestConvertEmbedsToTensors(CustomTestCase):
|
||||
|
||||
|
||||
class TestResolveEmbedOverrides(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_basic_resolution(self):
|
||||
embeds = [_vec(1), _vec(2)]
|
||||
pe = TokenizerManager._resolve_embed_overrides(
|
||||
@@ -144,6 +164,12 @@ class TestResolveEmbedOverrides(CustomTestCase):
|
||||
|
||||
|
||||
class TestGenerateReqInputEmbedOverride(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_single_override_in_getitem(self):
|
||||
"""Single PositionalEmbeds is shared across all items in __getitem__."""
|
||||
pe = PositionalEmbeds(embeds=[_vec()], positions=[0])
|
||||
@@ -176,6 +202,12 @@ class TestGenerateReqInputEmbedOverride(CustomTestCase):
|
||||
|
||||
|
||||
class TestEmbeddingReqInputEmbedOverride(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_override_fields_in_getitem(self):
|
||||
"""embed_override_token_id, embed_overrides, and positional_embed_overrides
|
||||
are correctly sliced in __getitem__."""
|
||||
@@ -220,6 +252,9 @@ class _FakeMixin(TokenizerManagerScoreMixin):
|
||||
|
||||
class TestResolveOverridesForSequence(CustomTestCase):
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
self.mixin = _FakeMixin()
|
||||
|
||||
def test_none_embeds_returns_empty(self):
|
||||
@@ -276,6 +311,9 @@ class TestResolveOverridesForSequence(CustomTestCase):
|
||||
|
||||
class TestResolveEmbedOverridesForRequest(CustomTestCase):
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
self.mixin = _FakeMixin()
|
||||
|
||||
def test_no_overrides_returns_none(self):
|
||||
@@ -339,6 +377,9 @@ DELIM_TOKEN = MIS_DELIMITER_TOKEN_ID
|
||||
|
||||
class TestBuildTokenIdInputs(CustomTestCase):
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
self.mixin = _FakeMixin(enable_mis=True)
|
||||
|
||||
# --- single-item mode, no embeds ---
|
||||
@@ -505,6 +546,9 @@ class TestScoreRequestValidation(CustomTestCase):
|
||||
"""Test validation guards in score_request without running full pipeline."""
|
||||
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
self.mixin = _FakeMixin()
|
||||
|
||||
def _call(self, **kwargs):
|
||||
|
||||
@@ -15,6 +15,8 @@ from types import SimpleNamespace
|
||||
import torch
|
||||
|
||||
from sglang.srt.managers.schedule_batch import ReqKvInfo
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.srt.utils import is_cuda, is_hip, is_npu, is_xpu
|
||||
from sglang.srt.utils.common import Range
|
||||
from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
@@ -165,7 +167,14 @@ class TestHiSparseUnit(unittest.TestCase):
|
||||
|
||||
Without this, a mid-test assertion failure skips cleanup and leaks
|
||||
resources, causing unrelated failures in later tests.
|
||||
|
||||
The code under test reads its configuration from the bags -- the PD
|
||||
decode prealloc path asks whether the decode radix cache is on -- so a
|
||||
case here needs a published config, the way a real process has one.
|
||||
"""
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="scheduler")
|
||||
self.allocator.clear()
|
||||
self.req_to_token_pool.clear()
|
||||
self.coordinator.mem_pool_host.clear()
|
||||
@@ -758,7 +767,6 @@ class TestHiSparseUnit(unittest.TestCase):
|
||||
queue.scheduler = SimpleNamespace(
|
||||
enable_hisparse=True,
|
||||
hisparse_coordinator=self.coordinator,
|
||||
server_args=SimpleNamespace(disaggregation_decode_enable_radix_cache=False),
|
||||
)
|
||||
|
||||
host_indices = queue._pre_alloc(req)
|
||||
|
||||
@@ -19,13 +19,20 @@ from sglang.srt.managers.schedule_batch import ( # noqa: E402
|
||||
ReqKvInfo,
|
||||
)
|
||||
from sglang.srt.managers.scheduler import Scheduler # noqa: E402
|
||||
from sglang.srt.runtime_context import get_context # noqa: E402
|
||||
from sglang.srt.runtime_context import get_context, publish, reset_context # noqa: E402
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
|
||||
register_cpu_ci(est_time=12, suite="base-a-test-cpu")
|
||||
|
||||
|
||||
class TestDisaggregationPriorityQueueing(unittest.TestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def _new_scheduler(self, disaggregation_mode: DisaggregationMode) -> Scheduler:
|
||||
scheduler = Scheduler.__new__(Scheduler)
|
||||
scheduler.disaggregation_mode = disaggregation_mode
|
||||
@@ -91,6 +98,12 @@ class TestDisaggregationPriorityQueueing(unittest.TestCase):
|
||||
|
||||
|
||||
class TestDecodePreallocQueuePriority(unittest.TestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def _new_decode_req(self, rid: str, priority: int, *, failed: bool = False):
|
||||
req = SimpleNamespace(
|
||||
rid=rid,
|
||||
@@ -227,6 +240,12 @@ class TestDecodePreallocQueueRebootstrapPayload(unittest.TestCase):
|
||||
dispatch itself now lives on the kv manager (see
|
||||
``TestCommonKVManagerPrefillRecompute``)."""
|
||||
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def _sampling_params(self):
|
||||
return SimpleNamespace(
|
||||
temperature=0.0,
|
||||
@@ -286,6 +305,12 @@ class TestCommonKVManagerPrefillRecompute(unittest.TestCase):
|
||||
``KVPoll.Failed`` so the scheduler's normal transfer-failure streaming runs.
|
||||
"""
|
||||
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def _new_manager(self):
|
||||
from sglang.srt.disaggregation.common.conn import CommonKVManager
|
||||
|
||||
@@ -423,6 +448,12 @@ class TestCommonKVManagerPrefillRecompute(unittest.TestCase):
|
||||
|
||||
|
||||
class TestDecodePrebuilt(unittest.TestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def _new_scheduler(self, *, enable_overlap: bool) -> Scheduler:
|
||||
scheduler = Scheduler.__new__(Scheduler)
|
||||
scheduler.grammar_manager = MagicMock()
|
||||
|
||||
@@ -10,7 +10,8 @@ from sglang.srt.managers.schedule_batch import ReqKvInfo
|
||||
from sglang.srt.mem_cache.allocator.hisparse import (
|
||||
DeepSeekV4HiSparseTokenToKVPoolAllocator,
|
||||
)
|
||||
from sglang.srt.runtime_context import get_context
|
||||
from sglang.srt.runtime_context import get_context, publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -18,6 +19,12 @@ register_cpu_ci(est_time=10, suite="base-a-test-cpu")
|
||||
|
||||
|
||||
class TestDeepSeekV4HiSparseAllocator(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The code under test reads its config from the bags.
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="tokenizer")
|
||||
|
||||
def test_forwards_swa_tail_allocation_to_logical_allocator(self):
|
||||
allocator = object.__new__(DeepSeekV4HiSparseTokenToKVPoolAllocator)
|
||||
logical_allocator = MagicMock(spec=["alloc_extend_swa_tail"])
|
||||
|
||||
@@ -12,6 +12,7 @@ from sglang.srt.multimodal.processors.qwen_vl import QwenVLImageProcessor
|
||||
from sglang.srt.multimodal.transport.cuda_ipc import (
|
||||
DEFER_CUDA_IPC_FEATURE_RECONSTRUCTION_KEY,
|
||||
)
|
||||
from sglang.srt.runtime_context import get_context
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
@@ -33,6 +34,15 @@ class _RecordingVisual:
|
||||
|
||||
|
||||
class TestQwen3VLFeatureMaterialization(CustomTestCase):
|
||||
def setUp(self):
|
||||
# The transport decision is read from the `mm` bag.
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
publish(ServerArgs(model_path="dummy"), role="test")
|
||||
|
||||
@staticmethod
|
||||
def _model(visual, *, use_data_parallel):
|
||||
model = Qwen3VLForConditionalGeneration.__new__(Qwen3VLForConditionalGeneration)
|
||||
@@ -43,10 +53,15 @@ class TestQwen3VLFeatureMaterialization(CustomTestCase):
|
||||
|
||||
def test_processor_defers_gpu_transport_for_encoder_dp(self):
|
||||
for transport in ("cuda_ipc", "cuda_vmm"):
|
||||
with self.subTest(transport=transport):
|
||||
# `mm_enable_dp_encoder` is read through `get_mm()` now, so stating
|
||||
# it on the processor's own `server_args` no longer reaches the
|
||||
# code under test.
|
||||
with (
|
||||
self.subTest(transport=transport),
|
||||
get_context().override_server_args(mm_enable_dp_encoder=True),
|
||||
):
|
||||
processor = QwenVLImageProcessor.__new__(QwenVLImageProcessor)
|
||||
processor.mm_feature_transport = transport
|
||||
processor.server_args = SimpleNamespace(mm_enable_dp_encoder=True)
|
||||
processor.model_type = "qwen3_vl"
|
||||
items = [
|
||||
MultimodalDataItem(modality=Modality.IMAGE),
|
||||
|
||||
@@ -94,6 +94,11 @@ def make_processor(case, config, image_processor_cls=None):
|
||||
publish(
|
||||
ServerArgs(
|
||||
model_path="dummy",
|
||||
# Mirrored for the same reason the stub sets it: `get_mm_processor_cls`
|
||||
# reads `model_impl` from this bag now, and "auto" would send it into
|
||||
# `get_resolved_model_impl`, which chokes on the SimpleNamespace
|
||||
# `model_config` these tests hand it.
|
||||
model_impl=server_args.model_impl,
|
||||
mm_feature_transport=server_args.mm_feature_transport,
|
||||
mm_process_config=server_args.mm_process_config,
|
||||
allowed_media_domains=server_args.allowed_media_domains,
|
||||
|
||||
@@ -20,7 +20,9 @@ from sglang.srt.managers.multimodal_processor import ( # noqa: E402
|
||||
get_mm_processor_cls,
|
||||
import_processors,
|
||||
)
|
||||
from sglang.srt.runtime_context import publish
|
||||
from sglang.srt.rust_server.multimodal import rust_mm_family_for # noqa: E402
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
|
||||
register_cpu_ci(est_time=14, suite="base-a-test-cpu")
|
||||
|
||||
@@ -28,7 +30,10 @@ register_cpu_ci(est_time=14, suite="base-a-test-cpu")
|
||||
def processor_cls_for(architecture, model_type):
|
||||
"""Through the production selection, as `resolve_spec` calls it."""
|
||||
hf_config = SimpleNamespace(architectures=[architecture], model_type=model_type)
|
||||
return get_mm_processor_cls(hf_config, SimpleNamespace(model_impl="sglang"))
|
||||
# `model_impl` is read from the bags now, so it has to be published rather
|
||||
# than handed over on a stand-in.
|
||||
publish(ServerArgs(model_path="dummy", model_impl="sglang"), role="tokenizer")
|
||||
return get_mm_processor_cls(hf_config)
|
||||
|
||||
|
||||
class TestRustMmGate(CustomTestCase):
|
||||
|
||||
@@ -12,6 +12,7 @@ from types import SimpleNamespace
|
||||
from unittest.mock import patch
|
||||
|
||||
from sglang.srt.multimodal.processors.base_processor import BaseMultimodalProcessor
|
||||
from sglang.srt.runtime_context import publish, reset_context
|
||||
from sglang.srt.server_args import ServerArgs
|
||||
from sglang.test.ci.ci_register import register_cpu_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
@@ -31,12 +32,25 @@ class _StubProcessor(BaseMultimodalProcessor):
|
||||
|
||||
|
||||
def _make(**fields):
|
||||
"""Both surfaces, because the device decision reads both.
|
||||
|
||||
`base_gpu_id` is the instance's own -- two engines in one process keep
|
||||
different ones, which `test_publishing_another_config_does_not_move_the_device`
|
||||
pins -- so it stays on the record the processor holds. `rl_on_policy_target`
|
||||
is the process's, so it is published.
|
||||
"""
|
||||
server_args = ServerArgs(model_path="dummy", **fields)
|
||||
publish(server_args, role="tokenizer")
|
||||
processor = _StubProcessor.__new__(_StubProcessor)
|
||||
processor.server_args = ServerArgs(model_path="dummy", **fields)
|
||||
processor.server_args = server_args
|
||||
return processor
|
||||
|
||||
|
||||
class TestFastImageProcessorDevice(CustomTestCase):
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
|
||||
def _device(self, processor, **platform):
|
||||
flags = {"_is_cpu": False, "_is_xpu": False, "_is_npu": False}
|
||||
flags.update(platform)
|
||||
@@ -80,6 +94,10 @@ class TestFastImageProcessorDevice(CustomTestCase):
|
||||
|
||||
|
||||
class TestFastImageProcessorMemoryPool(CustomTestCase):
|
||||
def setUp(self):
|
||||
reset_context()
|
||||
self.addCleanup(reset_context)
|
||||
|
||||
def _processor(self, *, transport="cpu", precompute_hash=False):
|
||||
processor = _make(base_gpu_id=0)
|
||||
processor.mm_feature_transport = transport
|
||||
|
||||
Reference in New Issue
Block a user