Commit Graph
122 Commits
Author SHA1 Message Date
54eb2904a4 minor: docs include mac installation (#25178)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-05-18 15:48:59 +08:00
Xia WeiwenandMa Mingfei 8d5ed330cc [XPU] Enable qwen3.5 on XPU (#21668)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-18 14:59:19 +08:00
a080358cac [Refactor] Refactor DeepEP dispatcher (#22822)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2026-05-18 04:36:42 +03:00
Baizhou Zhang 6dcacb1159 [Doc] Fix several places for dpsk v4 cookbook (#25506) 2026-05-16 21:54:15 -07:00
Yuhao Yang 57eb5bdaf6 [Doc] DSV4 cookbook: clean up env vars, add MegaMoE toggle, unify docker image (#25412) 2026-05-16 11:28:05 -07:00
zijiexia 9f26697d6a [Docs] Update DeepSeek V4 cookbook to use the latest docker image (#25410) 2026-05-16 11:17:51 -07:00
Zheng Luo 435ea41cf0 Delegate ModelExpress loading to package (#24723)
Signed-off-by: Zheng Luo <zheluo@nvidia.com>
2026-05-16 11:16:44 -07:00
Xinyuan Tong 33f1d3915f [NEW MODEL] Add H200 validation for Ring-2.6-1T cookbook (#25370) 2026-05-15 11:47:15 -07:00
Baizhou Zhang 1a1d69507d [Doc] Update MegaMoE usage (#25378) 2026-05-15 02:17:50 -07:00
Zhangheng c7e879e43f Add hicache feature in dsv4 cookbook (#25369) 2026-05-15 00:06:44 -07:00
Xinyuan Tong c3daa77e9a [NEW MODEL] Add Ring-2.6-1T cookbook (#25360) 2026-05-14 23:32:22 -07:00
Haoguang CaiandClaude Sonnet 4.6 626fd61308 📝 docs: add canonical URL to fix Google indexing lmsysorg.mintlify.app instead of docs.sglang.io (#24935)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-14 16:46:40 -07:00
Liangsheng Yin 67096f48bf Revert "[MoE] Decouple Mega MoE from DeepEP backend" (#25317) 2026-05-14 16:00:41 -07:00
Yuhao Yang 37f030a0de [MoE] Decouple Mega MoE from DeepEP backend (#24884) 2026-05-15 02:01:44 +08:00
amote-i 373a22c225 [NPU] [DOC] fix issues in ascend npu docs (#25268) 2026-05-14 17:25:47 +08:00
zijiexiaandClaude Opus 4.7 1f119f6a44 [Docs] update dsv4 cookbook with H100 deployment commands (#25243)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 23:45:29 -07:00
Lewisand百麒 0680f1b3d1 Add IntraNode NVLink configration in PD disaggregation docs (#23329)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
2026-05-13 23:23:57 -07:00
jianzhao-xu edb1b3f8f5 [NPU] add Ascend NPU Accuracy Evaluation and Faq docs (#24777) 2026-05-14 11:28:16 +08:00
amote-i 65e9f81c7d [NPU] [DOC] add performance testing and optimization docs for npu (#25114) 2026-05-14 09:47:55 +08:00
Le Zhangandlezhang 6ac30192fa [MLX] Add on-the-fly --quantization mlx_q4 / mlx_q8 for Apple Silicon (#24907)
Co-authored-by: lezhang <lezhang@local>
2026-05-13 11:06:13 -07:00
zijiexiaandClaude Opus 4.7 7d515c6d1f docs: prepend SGLANG_JIT_DEEPGEMM_PRECOMPILE=0 for H200 FP8 Flash max-throughput (#25152)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 00:50:15 -07:00
Khoa Pham c665edec6e [env] Make max KV chunk capacity configurable via SGLANG_MAX_KV_CHUNK_CAPACITY (#25120) 2026-05-12 22:37:45 -07:00
zijiexiaandClaude Opus 4.7 b0018ad015 [Doc]: refactor Intern-S2-Preview cookbook with interactive command generator (#25134)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 22:26:18 -07:00
RunningLeon 622baa17bd [Doc]: add interns2preview in cookbook (#25115) 2026-05-13 12:05:59 +08:00
Jimmy Shongandgithub-actions[bot] fd3eb77d45 [Cookbook]: add Laguna-XS.2 (Poolside) (#24730)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-12 16:06:26 +01:00
Ma Mingfei 71285335f7 Revert "Migrate Intel CPU cases to the test/registered." (#25044) 2026-05-12 13:32:47 +08:00
jundu ecf5d844f5 Migrate Intel CPU cases to the test/registered. (#22670) 2026-05-12 13:27:51 +08:00
R0CKSTAR 0a37d24e62 [diffusion] hardware: support sage attention backend on MUSA (attn backend, 21/N) (#24752)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2026-05-11 19:50:52 -07:00
shuwenn 5495026a3b [HiCache] feat: default storage prefetch timeout (#23309) 2026-05-11 18:49:35 -07:00
R0CKSTAR 74d70af09a [Apple Silicon] Add Metal kernel support in sgl-kernel (#23449)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-05-11 17:54:27 -07:00
Yihao Wang 7407a62c1c [Docs] Update MiniCPM-V-4.6 documentation and deployment configuration (#24991) 2026-05-11 11:10:19 -07:00
Ke Bao d9cb38012e Update MiMo V2.5 cookbook image to nightly (#24983) 2026-05-11 22:04:31 +08:00
Ke Bao 62edbc37c4 [Doc] Add rerun-test slash command usage (#24979) 2026-05-11 21:30:01 +08:00
Ke Bao 9c03171f6c Fix gb envs in deployment guide (#24977) 2026-05-11 21:06:02 +08:00
Mick 6e5b4de01a [diffusion] fix: further align ltx2.3 accuracy with tp (#24660) 2026-05-11 13:42:08 +08:00
Junlin Wu a623ee4cb5 📝 docs(diffusion): add MXFP8 quantization docs for Wan2.2 on Ascend NPU (#24918) 2026-05-11 08:13:34 +03:00
Yihao Wang 9f066cb55b [Docs] Add MiniCPM-V 4.6 cookbook (#24876) 2026-05-10 21:32:15 -07:00
egvenediktov 2473659e76 [NPU]Documentation update for communications quantization feature (#24668) 2026-05-10 23:49:21 +03:00
d82e339ce2 [Session R3] Add routed_experts_start_len for absolute routing slice control (#24851)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: zyzshishui <zyzshishui@gmail.com>
Co-authored-by: Yuzhen Zhou <82826991+zyzshishui@users.noreply.github.com>
2026-05-10 10:04:43 -07:00
ef5e9f8aba [DSV4] Cherry pick missing commits from deepseek_v4 branch and enhance tests (#24793)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
2026-05-09 04:15:37 -07:00
Brayden Zhongandb8zhong 8f33bee31b Reland Cute-DSL FP4 dense GEMM (#23590)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-09 02:20:58 -07:00
Jimmy Shong 096ad02b06 [Model] Laguna-XS.2 Model Support (#24204) 2026-05-09 05:43:13 +08:00
Mick 17888fa92a [diffusion] doc: update ltx2 multi-gpu deployment guide (#24682) 2026-05-08 18:38:05 +08:00
amote-i d32e283947 [NPU] [DOC] refresh npu supported model list (#24676) 2026-05-08 17:08:15 +08:00
amote-i 47e9ec11ad [NPU] [DOC] fix ascend_npu_support_new_models TOC (#24658) 2026-05-08 14:07:00 +08:00
+6 35870d55ac Deepseek V4 (#23882)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Qiaolin Yu <90088090+qiaolin-yu@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <11704492+yushengsu-thu@users.noreply.github.com>
Co-authored-by: Mingyi <27337995+wisclmy0611@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+againstentropy@users.noreply.github.com>
2026-05-07 18:32:21 -07:00
Revanth Reddy Airre be088f8076 fix(router): configure HTTP client connection settings (#24330)
Signed-off-by: Revanth Reddy Airre <revanthreddy@hippocraticai.com>
2026-05-07 11:42:45 -07:00
Revanth Reddy Airre d363315de9 fix(router): make HTTP pool idle timeout configurable (#24329)
Signed-off-by: Revanth Reddy Airre <revanthreddy@hippocraticai.com>
2026-05-06 22:11:11 -07:00
Baizhou Zhang 7ec18f7e4e [Doc] Fix instruction on Cuda 13 environments (#24516) 2026-05-06 02:37:23 -07:00
zijiexia 83b48fd523 [codex] update Nemotron3 Nano Omni cookbook benchmarks (#23998) 2026-05-05 14:53:15 -07:00