Cheng Wan
f1a512c51c
[Config] msgspec.Struct for the config tier ( #38753 )
2026-09-09 19:41:19 -07:00
Alex Nails and Alison Shao
28262c20df
[CI][RFC] Replace black-jupyter with ruff-format ( #37210 )
...
Co-authored-by: Alison Shao <a.shao@wustl.edu >
2026-09-02 19:46:08 -07:00
Khoa Pham
78d36f5f62
fix: kill_process_tree waits for the reap by default ( #36589 )
2026-08-27 02:34:05 -07:00
Cheng Wan
937af8538b
config: the runtime readers take the published bags ( #36254 )
2026-08-26 05:08:25 -07:00
Cheng Wan and Claude Opus 5
64aa859da2
config: constructing a config no longer resolves it ( #35907 )
...
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
2026-08-23 01:18:53 -07:00
Lianmin Zheng
198a7b2fc9
[Misc] Clean up python/sglang package structure ( #35062 )
2026-08-17 14:24:35 -07:00
Cheng Wan
99cfc90658
config: retire ServerArgs.override in favour of derive()
...
`ServerArgs.override(source, **fields)` was the last way to change a resolved
`ServerArgs` in place. Every remaining call-site was one of two things, and
neither wanted an in-place write:
- **A config for someone else.** A draft worker's context length, an encode
worker's device, the compile script's watchdog, the client's port pick, a test
fixture's backends. These already deepcopied first — the write was on the copy.
- **A launcher-stage resolution.** `resolve_auto_parsers` detected the chat
template's parsers and wrote them back, to be inherited by the schedulers it
spawns.
Both are "one config becomes another", so `derive(source, **fields)` returns the
variant and leaves the receiver — and any bags projected from it — untouched. It
deliberately is not `dataclasses.replace`: resolution does not re-run, because
the values being set are decided after it, from inputs it never had. Provenance
and the resolvable-field stash work as before, on the copy.
`resolve_auto_parsers` now computes the parsers and returns the config to launch
with; the detection helpers stop taking a config to mutate. `HiMambaRadixCache`
re-applied a HiCache layout normalization `__post_init__` already performs (the
same duplicate removed from `UnifiedRadixCache` in ebb1c88d23 ) and just goes.
With no in-place mutation left, `ServerArgs.__setattr__` raising after
resolution *is* the guarantee, so the textual writer ratchet retires and
`test_server_args_derive.py` pins the contract instead: the receiver survives
deriving, the published instance still refuses assignment, and deriving does not
publish. `SGLANG_STRICT_CONFIG_MUTATION` was already unused — the guard has been
unconditional since the mutation sweep — and goes with it.
The detection tests drop their `SimpleNamespace` stand-in for a real
`ServerArgs`; the test kit and the MLA chunk-metadata fixture publish a derived
variant instead of writing the runner's published config.
2026-08-05 19:30:53 -07:00
zijiexia and Claude Opus 4.8
b819d2fb5b
[Docs] Rename docs_new/ to docs/ ( #32123 )
...
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-08-03 16:51:00 -07:00
zijiexia
50ed4c011f
Remove legacy Sphinx docs/ and finish the Mintlify cutover ( #28964 )
2026-07-13 15:06:08 -07:00
Cheng Wan
b14f7b4f75
[refactor] Move model-capability adjustments into the resolution pipeline ( #30299 )
2026-07-07 21:26:55 -07:00
ishandhanani
aaa31eb0a1
feat: first-class session identity in SGLang ( #29436 )
2026-06-28 07:16:38 -07:00
shuwenn and Claude Opus 4.7
66ef97c00f
[Lint] Fix Optional[X] = (None,) typo defaults in two dataclasses ( #25252 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-15 00:32:40 -07:00
Emmanuel Acheampong and Claude Sonnet 4.6
b49d05fd0e
feat: add Crusoe managed inference backend ( #20475 )
...
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-05-12 16:23:23 -07:00
2813cb6d9a
[New Model] Gemma 4 ( #21952 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Pengyu Chen <pychen96@gmail.com >
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Andy Luo <andy.luo@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com >
2026-04-06 20:24:44 -07:00
Aurick Qiao
3178f3959f
Align incremental streaming logprobs with streamed output tokens ( #21583 )
2026-04-06 00:30:02 -07:00
David Cheung
ed427e1299
Migrate all callers from /get_server_info to /server_info ( #21463 )
2026-04-01 21:17:50 -07:00
blzheng
a98b456c70
[CPU] Add frontend support for Gemma ( #12590 )
2026-03-18 23:02:26 -07:00
Liangsheng Yin
f0458e0b49
[Utils] Move network/socket utilities from common.py to network.py ( #20646 )
2026-03-15 20:35:24 -07:00
Lianmin Zheng
5e1a495c65
Improve engine customization interface ( #15635 )
2025-12-22 14:24:16 -08:00
Baizhou Zhang
42fcf5438f
Revert "tiny remove deprecated endpoint call" ( #14533 )
2025-12-05 23:48:54 -08:00
b8zhong
ec7b2c16d9
tiny remove deprecated endpoint call ( #13607 )
2025-12-05 09:54:49 -08:00
Liangsheng Yin
e019f233f9
Remove unused code / testcases in lang ( #13335 )
2025-11-16 22:48:35 +08:00
Praneth Paruchuri
a53f2d6c12
Support orion ( #10665 )
2025-11-15 03:08:32 +08:00
Lianmin Zheng
ffc722a690
Revert "lang: support direct video inference" ( #12038 )
2025-10-23 19:21:31 -07:00
Mick and Lianmin Zheng
823b442945
lang: support direct video inference ( #9936 )
...
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
2025-10-23 18:12:39 -07:00
Glen Liu
47c606d3dc
[Feature] support regex strings as a stopping condition ( #10635 )
2025-10-12 10:53:15 +08:00
fzyzcjy
fdc4e1e570
Tiny move files to utils folder ( #11166 )
2025-10-03 22:40:06 +08:00
Lianmin Zheng
60e37f8028
Move parsers under a single folder ( #9912 )
2025-09-02 18:25:04 -07:00
Lianmin Zheng
b58ae7a2a0
Simplify frontend language ( #9029 )
2025-08-10 10:59:30 -07:00
f29aba8c6e
Support glm4.1v and glm4.5v ( #8798 )
...
Signed-off-by: Xinyuan Tong <justinning0323@outlook.com >
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Xinyuan Tong <justinning0323@outlook.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: zRzRzRzRzRzRzR <2448370773@qq.com >
Co-authored-by: Minglei Zhu <mingleizhu1122@gmail.com >
Co-authored-by: Chang Su <csu272@usc.edu >
2025-08-09 00:59:13 -07:00
b7094a5ef1
model: support intern-s1 ( #8350 )
...
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: zxy <zhou0493@e.ntu.edu.sg >
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Mick <mickjagger19@icloud.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
2025-07-26 13:48:51 -07:00
Yudi Xue
14c18d25df
Frontend language separate reasoning support ( #6031 )
2025-06-10 17:11:29 -07:00
Yueyang Pan
98c00a2df1
Fix torch profiler bugs for bench_offline_throughput.py ( #6557 )
2025-06-09 20:33:41 +08:00
Kiv Chen
5380cd7ea3
model(vlm): pixtral ( #5084 )
2025-05-13 00:16:10 -07:00
applesaucethebun and Brayden Zhong
2ce8793519
Add typo checker in pre-commit ( #6179 )
...
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca >
2025-05-11 12:55:00 +08:00
XinyuanTong
9d8ec2e67e
Fix and Clean up chat-template requirement for VLM ( #6114 )
...
Signed-off-by: Xinyuan Tong <justinning0323@outlook.com >
2025-05-11 00:14:09 +08:00
3409aaab32
Support InternVL3 ( #5350 )
...
Co-authored-by: Mick <mickjagger19@icloud.com >
Co-authored-by: Chayenne <zhaochen20@outlook.com >
2025-05-01 22:38:59 -07:00
Chuyue Sun and Shan Yu
08289eaa3e
Support o1 model on Azure ( #4980 )
...
Co-authored-by: Shan Yu <shanyu1@g.ucla.edu >
2025-04-21 00:46:09 -07:00
fzyzcjy
fba86b6b54
Tiny improve error message ( #5526 )
2025-04-20 16:00:15 -07:00
Lianmin Zheng
177320a582
Clean up imports ( #5467 )
2025-04-16 15:26:49 -07:00
f04c80dc42
Add Llama4 support ( #5092 )
...
Co-authored-by: Cheng Wan <cwan39@gatech.edu >
Co-authored-by: fzyzcjy <ch271828n@outlook.com >
Co-authored-by: ispobock <ispobaoke@163.com >
2025-04-07 00:29:36 -07:00
Mick
1e86457c90
model: Minicpmo ( #3023 )
2025-03-24 20:08:40 -07:00
fad86a6863
Support n in OpenAI API completions ( #3446 )
...
Co-authored-by: Shan Yu <shanyu1@g.ucla.edu >
Co-authored-by: Yineng Zhang <me@zhyncs.com >
Co-authored-by: chuyue sun <chuyue@lmsys.us-northcentral1-a.compute.internal >
2025-03-20 13:46:46 +08:00
Mick
9d02bb3e2a
Urgent model support: support gemma-3-it ( #4424 )
2025-03-16 17:37:32 -07:00
Mick
01090e8ac3
model: Support Janus-pro ( #3203 )
2025-03-12 11:02:11 -07:00
Qiaolin Yu
57a404fd55
Remove outdated test utils and fix links for the doc of sampling params ( #3999 )
2025-03-03 09:41:38 -08:00
Lianmin Zheng
66301e124f
Improve code styles ( #4021 )
2025-03-03 03:20:23 -08:00
ac2387279e
Support penalty in overlap mode; return logprob with chunked prefill; improve benchmark scripts ( #3988 )
...
Co-authored-by: SangBin Cho <rkooo567@gmail.com >
Co-authored-by: dhou-xai <dhou@x.ai >
Co-authored-by: Hanming Lu <hanming_lu@berkeley.edu >
2025-03-03 00:12:04 -08:00
Mick
45205d88a0
bench: Add MMMU benchmark for vLM ( #3562 )
2025-02-22 08:10:59 -08:00
Mick
bcc213df61
Model: Support Qwen 2.5 vl ( #3258 )
2025-02-16 00:58:53 -08:00