18694 Commits
Author SHA1 Message Date
jacky.chengandBingxu Chen 180bb2624f [AMD] Fix CI RuntimeError: opentelemetry package is not installed (#23940)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-04-29 18:02:44 +08:00
hhwxw d9270b8c6a fix(moe): relocate orphan tuned configs after #23019 (#24004) 2026-04-29 02:00:13 -07:00
Brian 6c7b242181 mimo v2.5 pro sglang-jax cookbook (#23936) 2026-04-29 16:58:41 +08:00
Sam (Kesen Li) 73e93bebd6 [1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954) 2026-04-29 01:45:48 -07:00
danielafrimi 8327270c72 [Kernel] Support FlashInfer TRTLLM-Gen fused MoE for non-gated FP4 & FP8 (Nemotron) (#21321) 2026-04-29 01:28:21 -07:00
Liangsheng Yin bd448e51bd [Spec] Split accept_length into num_accepted_drafts and num_accepted_tokens (#23962) 2026-04-29 00:02:22 -07:00
Xinyuan Tong 832b4f59ed [Bench] fix MMMU answer-extraction regex dropping multi-line responses (#23864) 2026-04-29 14:48:49 +08:00
shuwenn 2c41ef4c93 [HiCache] feat: add draft KV cache backing for L2/L3 (#21125) 2026-04-28 23:47:31 -07:00
Mick 71d2227a78 [diffusion] CI: update ground truth with official output (#23714) 2026-04-29 13:51:31 +08:00
JieTang66andJieTang 5a94088178 Fix typo: page_first_kv_spilt -> page_first_kv_split (#23983)
Co-authored-by: JieTang <tangjie66@huawei.com>
2026-04-29 12:08:05 +08:00
ZeyuanChen2000 23cfa9e4c8 [NPU] fix rope_theta get error for baichuan2-13b-chat model (#21543) 2026-04-29 12:01:03 +08:00
Yuhao Yang b437f6be48 model: Nemotron-omni-v3-alias (#23857) 2026-04-29 11:08:23 +08:00
Qiaolin Yu f57ec8d6ef [spec decoding] add extra attribute 'spec_hidden_size' (#23890) 2026-04-28 19:54:50 -07:00
Lianmin Zheng 2a771a40ac Add engine_type label to tokenizer manager metrics (#23978) 2026-04-28 19:52:58 -07:00
Lianmin Zheng d66eb3a91b docs: update contribution guide with coding style guidelines (#23977) 2026-04-28 19:51:58 -07:00
Rahul Vijayaraghavan 1aee04e7df [XPU] Support apply_router_weight_on_input for Llama4 for fused_experts (#22654)
merge this one as it is xpu only change.
2026-04-29 10:44:48 +08:00
ccullen-cert 1e14bd6f36 Fix for CVE-2026-5760 (#23660) 2026-04-28 19:39:44 -07:00
Baizhou Zhang 4e885baa9b docs(cookbook): add H200 (FP4) deployment option for DeepSeek-V4 (#23980) 2026-04-28 19:38:53 -07:00
Kangyan-ZhouandClaude Opus 4.7 feec1ac7f9 ci: clean up stale-CUDA mooncake variant in install_extra_deps (#23960)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 19:32:38 -07:00
Hubert Lu e5da200d0a [AMD] Fix Aiter RMSNorm layout handling (#23974) 2026-04-28 19:28:46 -07:00
Chandrakant KhandelwalandMa Mingfei 0ac23cffac Add intel_xpu as backend for GptOssForCausalLM, enabled for bf16 models with torch native backend (#12771)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-29 10:21:52 +08:00
iridiumineandiridiumine 08699bb1b2 [NPU] Fix DeepEP LL dispatch BF16 flag and skip triton kernel on NPU for Qwen3.5 (#23815)
Co-authored-by: iridiumine <iridiumine@users.noreply.github.com>
2026-04-29 10:19:24 +08:00
MingxuZhandMa Mingfei 2d27c38f13 Update XPU Docker runtime stack & hf_home config (#23820)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-29 10:03:11 +08:00
Xiaoyu Zhang 13afe8acdf [codex] Enable Qwen3-Next MoE all-reduce fusion (#23619) 2026-04-29 09:11:35 +08:00
Lianmin Zheng 14b4e6fa69 Support --model as alias for --model-path in CLI (#23894) 2026-04-28 17:23:27 -07:00
shuwenn 233048212a [HiCache][SPEC] fix: normalize storage prefetch key (#23631) 2026-04-28 15:53:20 -07:00
zijiexia 387c932dfc [Docs] update Docker image for Nemotron 3 Nano Omni (#23968) 2026-04-28 15:08:34 -07:00
Khoa Pham ddcacaf1bd Fix failing test_nvidia_nemotron_3_nano by fixing test_grouped_topk (#23874) 2026-04-28 15:03:58 -07:00
Alex NailsandClaude Opus 4.7 345fecc547 fix(bench): wire request_func in bench_long_context ContextWorkloadGenerator (#23898)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 14:45:51 -07:00
Liangsheng Yin cf0061da43 [Spec] Fix spec_accept_rate and unify accept/draft naming (#23530) 2026-04-28 14:40:04 -07:00
shuwennandClaude Opus 4.6 3e1c5e1b74 [HiCache][SPEC] fix: empty key after page alignment in match_prefix (#23387)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-28 14:06:55 -07:00
Lewisand百麒 9814cc89ce [Fix] NVFP4 qwen3.5 quant error fix by add packed_modules_mapping (#23471)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
2026-04-28 13:36:09 -07:00
914ef7c7f3 Fix multimodal /v1/embeddings Jinja chat template handling (#20835)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-04-28 13:05:45 -07:00
Qingfu Wen dc1eac4903 [MUSA][Diffusion] Fix fa3 API on MT MUSA (#23646) 2026-04-28 13:01:35 -07:00
Khoa PhamandClaude Opus 4.7 826f2d0620 chore(codeowners): add @kpham-sgl as owner for gemma4 files (#23916)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 11:43:32 -07:00
AlonKejzman 66ea0aee7f tokenizer: Add fastokens support (#23753) 2026-04-28 11:43:10 -07:00
zijiexia ad785a2299 [Docs] add Nemotron 3 Nano Omni cookbook (#23907) 2026-04-28 10:24:40 -07:00
Xinyuan Tong e458a9248f docs: enable MiMo V2.5 MTP cookbook path (#23945) 2026-04-28 10:22:19 -07:00
Xinyuan Tong 3fce8f2009 [Docs] add cookbook for Ling-2.6 family (#23947) 2026-04-29 00:42:04 +08:00
jacky.cheng d95715ec65 [AMD] Fix CI test_diffusion_generation[flux_2_image_t2i_2_gpus] (#23944) 2026-04-28 23:06:29 +08:00
Yuhao Yang 4e1ef6b3cf [Docs] Add single-node H200 DeepSeek-V4-Pro low-latency recipe (#23943) 2026-04-28 23:03:26 +08:00
Mick 144038fbae [diffusion] chore: change default seed to 42 (#23836) 2026-04-28 20:39:23 +08:00
Muqi Li 69a71219cb feat: tiny improve fp8_gemm tune usage (#23912) 2026-04-28 07:47:46 -04:00
Xiaoyu Zhang 7824903417 [SKILL] Sync SGLang skill docs (#23921) 2026-04-28 17:05:36 +08:00
Yinzuo Jiang 71160e4ddb feat(observability): add OpenTelemetry tracing for pipeline parallelism (#23169)
Signed-off-by: Yinzuo Jiang <jiangyinzuo@foxmail.com>
2026-04-28 17:05:23 +08:00
Xun Sun 9a53ab3d6d [6/N] (Elastic EP) Recover failed ranks (#15771) 2026-04-28 00:44:26 -07:00
b8a2dcd300 fix: resolve tensor file overwrite between target and draft models (#21694)
Co-authored-by: jiangguangya <jiangguangya@baidu.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-04-27 23:40:49 -07:00
zijiexia 7ce2352323 [docs] fix sglang serve --model-path in cookbooks (#23905) 2026-04-27 21:54:21 -07:00
Xiaoyu Zhang 6fbad22feb Remove smoke wording from tests and comments (#23355) 2026-04-28 12:05:27 +08:00
JoyFuture 1a55646dcd [Feature] Xiaomi MiMo-V2.5-Pro day0 support (#23808) 2026-04-28 11:43:29 +08:00