Rain Jiang
|
b83a59835d
|
sglang-server remove opaque type (#38095)
|
2026-09-08 00:41:31 -07:00 |
|
Rain Jiang
|
7d7ab4b5c6
|
Rainj me/rust server refactor2 (#35239)
|
2026-08-21 16:37:02 -07:00 |
|
Rain Jiang
|
4494fb96b2
|
refactor the tcp listener binding logic (#33420)
|
2026-08-03 22:10:52 -07:00 |
|
 Rain JiangandKan Wu
|
c844244da5
|
support dp attn with client lb (#33105)
Co-authored-by: Kan Wu <wukanustc@gmail.com>
|
2026-08-02 18:32:35 -07:00 |
|
Rain Jiang
|
1d640aaea2
|
bump dynamo-tokenizers to 1.7.0 (#32981)
|
2026-07-31 15:02:28 -07:00 |
|
 Rain JiangandAlex Nails
|
9dcaf6bfdf
|
rust server build release artifacts (#33096)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-31 12:36:25 -07:00 |
|
Rain Jiang
|
d7a4c830e5
|
sglang rust server tokenizer manager, ring and runtime (#32358)
|
2026-07-29 18:42:12 -07:00 |
|
Rain Jiang
|
6e48c13497
|
sglang rust server egress message (#32342)
|
2026-07-29 17:16:06 -07:00 |
|
Rain Jiang
|
d24de56995
|
sglang rust server sampling message (#32343)
|
2026-07-29 14:55:06 -07:00 |
|
Rain Jiang
|
2ca2ca753a
|
sglang rust server request message (#32242)
|
2026-07-29 10:54:56 -07:00 |
|
Rain Jiang
|
d943636a48
|
support regex that compatible with python re lib however apply more l… (#32676)
|
2026-07-28 14:25:31 -07:00 |
|
Rain Jiang
|
be7cc17307
|
sglang rust server environ fsm error id gen (#32240)
|
2026-07-24 20:55:10 +00:00 |
|
Rain Jiang
|
20eb37a2a1
|
init sglang rust server project (#32256)
|
2026-07-23 21:47:15 +00:00 |
|
Rain Jiang
|
7fe82dd02e
|
create rust workspace (#32014)
|
2026-07-23 12:02:41 -07:00 |
|
Rain Jiang
|
ae1f0c6d07
|
session_id dataclass field should not put in msgpack struct (#29977)
|
2026-07-02 14:10:12 -07:00 |
|
 Rain JiangandLianmin Zheng
|
be1930133a
|
Convert IPC dataclasses to msgspec.Struct with opt-in msgpack transport (#28688)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-26 12:04:03 -07:00 |
|
Rain Jiang
|
ef01618dfb
|
support MPServer and embedded server for granian to enable muti tokenizer worker (#28573)
|
2026-06-18 17:59:12 -07:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
Rain Jiang
|
cb1e63aba4
|
bump fa4 to official released fa4 pkg (#20303)
|
2026-03-17 17:22:56 -07:00 |
|
Rain Jiang
|
ab4b863546
|
fix ci by removing nvidia-cutlass-dsl-libs-base and force reinstall n… (#20380)
|
2026-03-11 13:37:33 -07:00 |
|
Rain Jiang
|
61b228239e
|
bump sgl-fa4 version to 4.0.5 to loose torch deps (#20378)
|
2026-03-11 13:08:09 -07:00 |
|
Rain Jiang
|
472eef4071
|
fa4 cleanup (#19727)
|
2026-03-05 17:54:25 +08:00 |
|
Rain Jiang
|
0ffd0a3995
|
Nsa trtllm mla sparse fp8 support with Deepseek v3.2 NVFP4 (#18389)
|
2026-02-16 09:29:54 +08:00 |
|
Rain Jiang
|
f0e948a0f1
|
fix the deepep 8 gpu unit test (#14601)
|
2025-12-09 01:40:09 -08:00 |
|
 Rain JiangandTrevor Morris
|
ea177372bd
|
support mtp with deepseek r1 nvfp4 model (#13115)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
|
2025-12-06 00:45:54 -08:00 |
|
Rain Jiang
|
0f76976c3c
|
remove the fa4 page_size hardcode to 128 restriction on mla model arch (#12801)
|
2025-11-07 13:30:30 -08:00 |
|
Rain Jiang
|
a119363f08
|
ignore the deepgemm check when the model weight with nvfp4 and moe ba… (#12782)
|
2025-11-06 15:27:28 -08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Rain Jiangandgemini-code-assist[bot]
|
8e797a47f0
|
fix: the hardcode hf repo name comparison for deepseek-ocr (#12031)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-10-23 21:37:56 -07:00 |
|
 
|
2286e85e77
|
pass a_scale from fp8 quant result instead of hard code to 1.0f (#10241)
Co-authored-by: Yichen Wang <yichen.wang@bytedance.com>
Co-authored-by: Jinwu Guo <641876696@qq.com>
|
2025-09-10 12:56:05 -07:00 |
|
Rain Jiang
|
df5407fb53
|
Revert "feat: add fused moe config for Qwen3-30B-A3B on B200" (#10185)
|
2025-09-08 18:11:15 -07:00 |
|
Rain Jiang
|
7a40e4f4a6
|
fix the cutlass moe tests (#10182)
|
2025-09-08 16:24:55 -07:00 |
|
Rain Jiang
|
6049ca209e
|
move compile threads to an option to avoid OOM on low memory host (#10123)
|
2025-09-07 21:36:14 -07:00 |
|
Rain Jiang
|
7802586cab
|
fix the fp8 topk_config.correction_bias is none bug (#10040)
|
2025-09-07 20:28:14 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Rain Jiangandgemini-code-assist[bot]
|
9db8025376
|
support fp8 kvcache for hybrid attn backend on GPT-OSS (#9783)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-09-01 19:17:12 +00:00 |
|
Rain Jiang
|
6b39f9cf8c
|
Support compile sgl-kernel on cuda 13.0 (#9721)
|
2025-08-28 10:18:03 -07:00 |
|
Rain Jiang
|
79e6a8a6ac
|
support cuda 13.0 and trtllm kernel by Aug 25 2025 (#9495)
|
2025-08-26 23:13:27 -07:00 |
|