Commit Graph
100 Commits
Author SHA1 Message Date
Yuhao YangandClaude Code a37ded1693 [Cookbook] DeepSeek-V4.1: add the HiCache L2 knob to the Playground (#38844)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-10 18:20:34 +08:00
Yuhao Yang 13469c16d3 Store mamba prefix-cache checkpoints at the configured SSM state dtype (#34820) 2026-09-09 15:25:57 +08:00
Yuhao Yang 8eaffdf382 docs: point the Qwen3.8-Flash-Next cookbook at model support PR #36497 (#36499) 2026-08-26 12:53:55 +00:00
c7b5e76fa9 Add Qwen3.8-Flash-Next cookbook (#36496)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 20:36:59 +08:00
8a1e6e4e46 Qwen3.8-27B Model Support (#34859)
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-08-19 16:31:43 +08:00
Yuhao Yang 861eca8e25 docs: add NVFP4 quantization option to Kimi-K3 deploy panel (#35168) 2026-08-17 11:01:47 -07:00
Yuhao YangandClaude 6ad3f2d8fd docs: link dots3.note checkpoints, add H100 cells (#34797)
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-14 02:42:27 +00:00
c1142677a8 [do not merge] add new cookbooks (#34658)
Co-authored-by: Jianfei Wang <jianfei.wangg@outlook.com>
Co-authored-by: Jianfei Wang <905787410@qq.com>
2026-08-13 13:35:15 +00:00
Yuhao Yang 971932d661 [Kimi-K3] Allow DSPARK verify on cutedsl_mla (fold_sq) (#33650) 2026-08-06 13:54:25 -07:00
Yuhao Yang 7cae831e41 Update mi35x ROCm image to k3-20260727 (#32559) 2026-07-27 10:52:41 -07:00
Yuhao Yang 37a830b667 [Feature] Add DWDP (Distributed Weight Data Parallelism) for MoE prefill (#29778) 2026-07-20 23:59:54 -07:00
Yuhao Yang a03ca46a28 Fix KDA prefix caching under mamba extra_buffer and enable it for kimi_linear (#31474) 2026-07-19 20:09:03 +08:00
Yuhao Yang 5abec3fbf8 Fix MiMo-V2 on Blackwell: FA3 fallback and TP-aware audio weight loading (#31343) 2026-07-15 13:26:27 -07:00
Yuhao Yang dd2e4cdc99 Add Inkling cookbook (#31360) 2026-07-16 02:24:31 +08:00
Yuhao Yang 4a8200565e Fused QK GemmaRMSNorm + RoPE + gate kernel for Qwen3.5 (#28320) 2026-06-25 15:58:54 +08:00
Yuhao Yang 0fc815aa2c [CI] Temporarily disable openbmb MiniCPM tests (#29095) 2026-06-24 01:39:44 -07:00
Yuhao Yang aea0e30853 [4/N] Qwen3.5Opt: Overlap mamba verify update with draft extend (#26924) 2026-06-13 20:29:20 +08:00
Yuhao YangandXinyuan Tong 50815d54a7 docs (#28061)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-06-12 21:19:40 +08:00
Yuhao Yang 46c58b5c70 bench: fix MMMU VLM eval max_tokens for CoT prompt (#27327) 2026-06-05 10:28:16 +08:00
Yuhao Yang a8cfae0b30 doc: update step-3.7-flash docker image tag (#26625) 2026-05-29 08:40:17 +08:00
3bdea78ad1 model: support Step-3.7-Flash (#26565)
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: luotingdan <luotingdan@stepfun.com>
2026-05-29 08:00:54 +08:00
Yuhao YangandKe Bao c9153da5dc Fix SWA double-free in disagg decode with MTP speculation (#25805)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-05-22 00:28:54 -07:00
Yuhao YangandYaochen Han fa6f4dfb35 improve: combine vit calls for images from different reqs from one batch (#25910)
Co-authored-by: Yaochen Han <814073252@qq.com>
2026-05-22 14:57:41 +08:00
Yuhao Yang 81d686d9fa Default MegaMoE to W4A8 for Max-Throughput recipe (#26004) 2026-05-21 11:54:16 -07:00
Yuhao Yang 79ea30d1f1 [Bug] Fix V4-Pro NaN on Blackwell by converting fp8_einsum input scale to ue8m0 (#25733) 2026-05-18 23:48:34 -07:00
Yuhao Yang d8e66e54e5 fix: use triton_attn as default vision attention on B300 (SM103) (#25570) 2026-05-19 11:00:07 +08:00
Yuhao Yang 57eb5bdaf6 [Doc] DSV4 cookbook: clean up env vars, add MegaMoE toggle, unify docker image (#25412) 2026-05-16 11:28:05 -07:00
Yuhao Yang b2c6db0cc4 [MoE] Decouple Mega MoE from DeepEP backend (#25406) 2026-05-16 00:18:43 -07:00
Yuhao Yang 18c16f8660 Port SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN from deepseek_v4_dev (#25419) 2026-05-15 17:39:13 -07:00
5ba69f50fb Add multi-detokenizer support (#24944)
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-05-15 17:26:26 -07:00
Yuhao Yang 88d3ed7df1 Enable SGLANG_OPT_FP8_WO_A_GEMM by default (#25181) 2026-05-15 02:09:13 +08:00
Yuhao Yang 37f030a0de [MoE] Decouple Mega MoE from DeepEP backend (#24884) 2026-05-15 02:01:44 +08:00
e2290b155a Port KV Compression V2 from deepseek_v4_dev (#24890)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: DarkSharpness <2040703891@qq.com>
2026-05-13 22:40:38 +08:00
d0913fca8d Port fused SiLU+clamp+FP8 quant from DSV4 dev branch (#24897)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: zcnrex <zcnrex@gmail.com>
2026-05-13 22:36:44 +08:00
d6d3d0f599 Optimize SWA memory preallocation for disaggregated decode (#24857)
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
2026-05-13 09:09:34 +08:00
2f06867128 Optimize MHC pipeline: DeepGemm, fused norm, fused hc_head (#24775)
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
2026-05-10 19:03:37 +08:00
Yuhao YangandCheng Wan bd0aa22309 Fix PD bootstrap failure handling (#24772)
Co-authored-by: Cheng Wan <chwan@rice.edu>
2026-05-10 19:02:47 +08:00
Yuhao Yang 09912fd89d Remove unnecessary bf16 assert in rotate_activation (#24686) 2026-05-09 05:00:52 +08:00
Yuhao YangandClaude Opus 4.6 6279aee716 [docs] Update B300 Pro cookbook with accuracy-verified serving configs (#24367)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-05 07:26:03 +01:00
Yuhao Yang b437f6be48 model: Nemotron-omni-v3-alias (#23857) 2026-04-29 11:08:23 +08:00
Yuhao Yang 4e1ef6b3cf [Docs] Add single-node H200 DeepSeek-V4-Pro low-latency recipe (#23943) 2026-04-28 23:03:26 +08:00
Yuhao Yang 049f1bf6fb docs(DeepSeek-V4): add GB200 platform to cookbook recipe (#23725) 2026-04-25 20:54:55 -07:00
Yuhao Yangandtrangdough 4a3fe2a091 model: support parakeet nemotron encoder (#23568)
Co-authored-by: trangdough <trangtdo22@gmail.com>
2026-04-25 11:00:23 +08:00
Yuhao Yang f41f1a74a4 [diffusion] chore: support custom output folder name in GT generation workflow (#23422) 2026-04-22 11:18:21 +08:00
Yuhao Yang 5595f6e988 Fix trtllm mla chunked-prefill zero-length bug (#22291) (#22688) 2026-04-20 22:10:13 -07:00
fe9b9b254b Fix segfault in cudaMemcpyBatchAsync on CUDA 13.0 (#23136)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-04-20 12:20:22 -07:00
Yuhao Yang a77abbe005 [VLM] Reduce GPU memory footprint of CUDA IPC MM feature transport (#22662) 2026-04-17 10:38:36 +08:00
Yuhao Yang 9da998a882 [diffusion] feat: disaggregated diffusion (#21701) 2026-04-16 23:51:32 +08:00
Yuhao Yang b8794baa6d [Step3p5] Optimize allreduce in MoE layers (#22773) 2026-04-16 09:33:12 +08:00
Yuhao Yang 8686f42acb [VLM] Enable per-image ViT cache and avoid TP CUDA context creation for Kimi-K2.5 (#22858) 2026-04-16 01:14:24 +08:00
Yuhao Yang 16f306fd85 [VLM] GPU Image Preprocessing for Kimi-K2.5 (#22368) 2026-04-11 11:13:30 +08:00
Yuhao Yang f5fd5ab622 add whisper test (#22302) 2026-04-10 15:34:53 +08:00
Yuhao YangandMick 2b119ba388 [diffusion] fix: fix accuracy for flux series (#22059)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-04-05 16:03:17 +08:00
Yuhao Yang 34d5765e2f [VLM] Chunk-aware ViT encoding with per-image cache and lazy device transfer (#22038) 2026-04-04 16:55:17 +08:00
Yuhao Yang 69e89a1fcc [VLM] Enable per-image MM splitting by default and remove MULTI_IMAGES modality (#21899) 2026-04-03 11:04:41 +08:00
Yuhao Yang 2ef12073f4 [VLM] Add VLM TP=4 per-commit CI test and improve MMMU eval prompt/parser (#21841) 2026-04-01 20:09:47 -07:00
Yuhao Yang 1aabe44b64 [VLM] remove AsyncMMDataProcessor wrapper (#21651) 2026-04-01 17:39:50 +08:00
Yuhao Yang 68a4573627 [diffusion] fix: fix Flux.2 with tp(#21664) 2026-03-31 14:14:59 +08:00
Yuhao Yang 4e69f14b95 fix bench_serving sglang backend to support image dataset (#21294) 2026-03-29 10:02:11 +08:00
Yuhao Yang 57cf4790ca [VLM] Optimize ShmPointerMMData for multi-pickle safety and deferred unwrap (#21465) 2026-03-28 23:11:12 +08:00
Yuhao Yang 5ef56682b8 reduce CPU peak memory in multimodal tensor hashing (#21123) 2026-03-28 11:09:16 +08:00
Yuhao Yang 32a85ef128 [diffusion] CI: auto-skip diffusion tests when required pipeline class is missing from diffusers (#21139) 2026-03-23 12:15:21 +08:00
Yuhao Yang c32e35a2a5 [diffusion] CI: fix picklingerror for diffusion models using diffusers backend (#20854) 2026-03-22 11:51:03 +08:00
Yuhao Yang 24a27d5320 vlm: support piecewise cuda graph for Kimi-K2.5 (#20747) 2026-03-18 00:32:07 +08:00
Yuhao Yang 2ccdb7373e [diffusion] CI: fix consistency test workflow (#20704) 2026-03-17 07:42:30 +08:00
Yuhao Yangandwili-65535 1c456a0af5 VLM: add Conv2dLayer/Conv3dLayer to fix PyTorch 2.9.1 CuDNN Conv3d (#20282)
Co-authored-by: wili-65535 <wili-65535@users.noreply.github.com>
2026-03-15 19:17:44 +08:00
Yuhao Yang a6ecf050be diffusion: fix helios accuracy issue (#20036) 2026-03-15 13:55:51 +08:00
Yuhao Yang a57a44739f [diffusion] deps: upgrade diffusers from 0.36.0 to 0.37.0 (#20318) 2026-03-12 19:17:28 +08:00
Yuhao Yang ecca8c553d [diffusion] fix: fix diffusers backend issues in diffusion ci gt workflow (#20173) 2026-03-10 00:51:48 +08:00
Yuhao Yang 1cb86f5171 [diffusion] CI: fix CI script path and missing server arg in perf baseline generator (#20138) 2026-03-09 10:35:21 +08:00
Yuhao Yang 57f28fda90 [diffusion] chore: add diffusion new model skill (#19605) 2026-03-09 09:45:23 +08:00
Yuhao Yang 115f879958 Helios: Real Real-Time Long Video Generation Model (#19782) 2026-03-04 14:58:04 +08:00
Yuhao Yang ca44aa25af Fix dp_attention crash when dp_size < tp_size in warmup dummy run (#19760) 2026-03-03 19:43:13 -08:00
Yuhao YangandProzac614 b01b07aa16 [diffusion] CI: GT generation flow for diffusion CI (#19236)
Co-authored-by: Prozac614 <dwt614707404@163.com>
2026-02-28 14:07:45 +08:00
Yuhao Yangandyizhang2077 c7c4a1cbbd refactor linear attention backend (#18622)
Co-authored-by: yizhang2077 <1109276519@qq.com>
2026-02-25 23:02:44 +08:00
Yuhao Yang 5a7ae059e3 Add DP ViT support for Kimi K2.5 (#18689) 2026-02-18 23:03:07 +08:00
Yuhao Yangandltd0924 980d2936cd model: support Step-3.5-Flash (#18084)
Co-authored-by: ltd0924 <ltd0924@sina.com>
2026-02-03 00:40:07 +08:00
Yuhao Yang d11ccc0a0a fix: avoid double reduce in VLM dp attention (#17991) 2026-02-02 09:44:32 +08:00
Yuhao Yang 3c2f4c7bbe [diffusion] model: sync with upstream z-Image (#17822) 2026-01-29 21:10:11 +08:00
Yuhao YangandMick 479ab7a4e7 model: support Kimi-K2.5 (#17789)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-27 10:57:00 +08:00
Yuhao Yang f7a0bcda1e model: step3-vl-10b (#17513) 2026-01-22 23:15:08 +08:00
Yuhao Yangandjianyingzhu a0b4ba9032 [diffusion] model: GLM-Image (#16894)
Co-authored-by: jianyingzhu <53300651@qq.com>
2026-01-14 02:02:03 +08:00
Yuhao YangandMick e14f5ec8a8 [diffusion] refactor: eliminate redundant parameters in req (#16505)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-08 11:14:03 +08:00
Yuhao Yang 10174e1114 Revert "[grpc] update api to scheduler in grpc request manager" (#16387) 2026-01-04 22:05:39 -08:00
Yuhao Yang 2138ff48c6 Revert "[FEAT] optimize tensor zmq transfer for multimodal inputs" (#16386) 2026-01-04 22:05:26 -08:00
Yuhao Yang 6c8587b5db [diffusion] fix: align negative prompt with official readme for new model (#16222) 2026-01-02 15:20:56 +08:00
Yuhao YangandMick 4280a18a13 [diffusion] CI: add test for cache-dit (#16204)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-31 19:58:10 +08:00
Yuhao YangandMick 39ca57cd28 [diffusion] chore: tiny fix model config (#16159)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-30 22:11:38 +08:00
Yuhao Yang 0cd2b719a5 [diffusion] chore: remove useless params (#15925) 2025-12-28 01:01:08 +08:00
Yuhao Yang 29ce7b3612 [diffusion] chore: remove stepvideo code (#15918) 2025-12-27 13:25:05 +08:00
Yuhao Yang ba41080892 [diffusion] model: support qwen-image-edit-2511 (#15458) 2025-12-19 20:06:00 +08:00
Yuhao Yang 3d42b7e7b0 unified management of environment variables for vlm cuda ipc transport (#14501) 2025-12-18 12:28:06 +08:00
Yuhao Yang 01b955ac3d [diffusion] model: support mutli-image input and qwen-image-edit-2509 (#15005) 2025-12-15 16:17:10 +08:00
Yuhao Yang a81cc1b8b3 add transformers version validation for glm-4.6v moe models (#14998) 2025-12-13 10:54:08 -08:00
Yuhao Yang 06b58c5dc5 fix flaky image access in ci by switching to raw content url (#14940) 2025-12-13 10:52:06 -08:00
Yuhao Yang b62fe8504c fix nightly vlm ci : restore original eval for requests without regex (#14875) 2025-12-10 23:13:25 -08:00
Yuhao Yang c1bd5ee8c5 Revert transformers to 4.57.1 (#14801) 2025-12-10 11:04:36 -08:00
Yuhao Yang 02f1e81e2d Revert "fix: checking if tokenizer is in cache before downloading from HF" (#14808) 2025-12-10 01:14:35 -08:00
Yuhao Yang 793c98afaf handling incomplete rope_scaling config ci after transformers upgrade (#14784) 2025-12-09 22:56:16 -08:00
Yuhao Yang 15bc8cbd74 fix rope parameter initialization error caused by transformers v5.0 update (#14745) 2025-12-09 10:51:26 -08:00