Commit Graph
12113 Commits
Author SHA1 Message Date
Qiaolin Yu 79dbfe4505 Use spec v2 by default (#21062) 2026-04-29 13:40:42 -07:00
Alex NailsandClaude Opus 4.7 c3ab5bec7d ci: consolidate rust + protoc install across workflows (#23700)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 13:39:02 -07:00
Baizhou Zhang b3ead32d3c [minor] Remove incorrect note after supporting w4a16 moe for DeepSeek V4 (#24035) 2026-04-29 13:26:13 -07:00
Zhongdongming Dai 7389743d85 feat: Support modelexpress p2p RDMA transfer (#23105) 2026-04-29 12:57:40 -07:00
jsheng_Linkedin db84a8ebbb [Model] Qwen3ForPooledOutput: forward get_input_embeddings to inner model (#23434) 2026-04-29 12:25:06 -07:00
3272af2f00 [Apple Silicon] [MLX] MLX decode partial overlap scheduling for generation (async eval) (#22416)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-04-29 12:21:14 -07:00
Kangyan-ZhouandClaude Opus 4.7 d4040e7010 [CI] Broaden stage-b-test-1-gpu-large runner pool to H100 + H200 (#24080)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 12:18:10 -07:00
Yihao Wang 903e46d848 [Bench] fix bench_hf.py KeyError + reduce print spam + add --limit (#24079) 2026-04-29 11:43:37 -07:00
Xinyuan Tong 1376761841 fix(moe): repair dead import in fused_moe_native after MoE refactor (#24069) 2026-04-29 11:14:52 -07:00
Jeongho Shin 530b497a48 [BugFix] correct host leaf status check from evicted to backuped (#23537) 2026-04-29 10:26:00 -07:00
shuwennandClaude Opus 4.7 03147f66b8 ci: add /rerun-group to rerun all registered tests in a group (#24023)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 10:24:16 -07:00
Xinyuan Tong 4cf109bbd1 debug followup (#24058) 2026-04-29 23:03:27 +08:00
Qi Yuhangandgemini-code-assist[bot] 3f7c95d6cc [JIT Kernel][1/2]Migrate MXFP8 Group GEMM & Quant into JIT (#23833)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-29 22:50:09 +08:00
Mick eeb7b5c433 [diffusion] fix: align encoder of flux klein with official (#24008) 2026-04-29 22:42:16 +08:00
YC Yen-Ching Tsengandbingxche 155e333039 [AMD] Update AMD CI workflow concurrency group (#24065)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-29 22:39:18 +08:00
Mick e5c2a9b6cf [diffusion] fix: improve LTX2.3 reference accuracy controls (#24022) 2026-04-29 21:39:27 +08:00
Xinyuan Tong 1279ae0787 Bugfix (#24027) 2026-04-29 21:13:51 +08:00
AndyLi429 4c1eefca4f [NPU] ascend backend support qwen3 moe attention cp (#21685) 2026-04-29 19:25:17 +08:00
Zheng Wengang ae0c036c24 [BugFix][EPD] fix embedding req_id transfer error (#23481) 2026-04-29 18:56:03 +08:00
jacky.chengandBingxu Chen 180bb2624f [AMD] Fix CI RuntimeError: opentelemetry package is not installed (#23940)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-04-29 18:02:44 +08:00
hhwxw d9270b8c6a fix(moe): relocate orphan tuned configs after #23019 (#24004) 2026-04-29 02:00:13 -07:00
Brian 6c7b242181 mimo v2.5 pro sglang-jax cookbook (#23936) 2026-04-29 16:58:41 +08:00
Sam (Kesen Li) 73e93bebd6 [1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954) 2026-04-29 01:45:48 -07:00
danielafrimi 8327270c72 [Kernel] Support FlashInfer TRTLLM-Gen fused MoE for non-gated FP4 & FP8 (Nemotron) (#21321) 2026-04-29 01:28:21 -07:00
Liangsheng Yin bd448e51bd [Spec] Split accept_length into num_accepted_drafts and num_accepted_tokens (#23962) 2026-04-29 00:02:22 -07:00
Xinyuan Tong 832b4f59ed [Bench] fix MMMU answer-extraction regex dropping multi-line responses (#23864) 2026-04-29 14:48:49 +08:00
shuwenn 2c41ef4c93 [HiCache] feat: add draft KV cache backing for L2/L3 (#21125) 2026-04-28 23:47:31 -07:00
Mick 71d2227a78 [diffusion] CI: update ground truth with official output (#23714) 2026-04-29 13:51:31 +08:00
JieTang66andJieTang 5a94088178 Fix typo: page_first_kv_spilt -> page_first_kv_split (#23983)
Co-authored-by: JieTang <tangjie66@huawei.com>
2026-04-29 12:08:05 +08:00
ZeyuanChen2000 23cfa9e4c8 [NPU] fix rope_theta get error for baichuan2-13b-chat model (#21543) 2026-04-29 12:01:03 +08:00
Yuhao Yang b437f6be48 model: Nemotron-omni-v3-alias (#23857) 2026-04-29 11:08:23 +08:00
Qiaolin Yu f57ec8d6ef [spec decoding] add extra attribute 'spec_hidden_size' (#23890) 2026-04-28 19:54:50 -07:00
Lianmin Zheng 2a771a40ac Add engine_type label to tokenizer manager metrics (#23978) 2026-04-28 19:52:58 -07:00
Lianmin Zheng d66eb3a91b docs: update contribution guide with coding style guidelines (#23977) 2026-04-28 19:51:58 -07:00
Rahul Vijayaraghavan 1aee04e7df [XPU] Support apply_router_weight_on_input for Llama4 for fused_experts (#22654)
merge this one as it is xpu only change.
2026-04-29 10:44:48 +08:00
ccullen-cert 1e14bd6f36 Fix for CVE-2026-5760 (#23660) 2026-04-28 19:39:44 -07:00
Baizhou Zhang 4e885baa9b docs(cookbook): add H200 (FP4) deployment option for DeepSeek-V4 (#23980) 2026-04-28 19:38:53 -07:00
Kangyan-ZhouandClaude Opus 4.7 feec1ac7f9 ci: clean up stale-CUDA mooncake variant in install_extra_deps (#23960)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 19:32:38 -07:00
Hubert Lu e5da200d0a [AMD] Fix Aiter RMSNorm layout handling (#23974) 2026-04-28 19:28:46 -07:00
Chandrakant KhandelwalandMa Mingfei 0ac23cffac Add intel_xpu as backend for GptOssForCausalLM, enabled for bf16 models with torch native backend (#12771)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-29 10:21:52 +08:00
iridiumineandiridiumine 08699bb1b2 [NPU] Fix DeepEP LL dispatch BF16 flag and skip triton kernel on NPU for Qwen3.5 (#23815)
Co-authored-by: iridiumine <iridiumine@users.noreply.github.com>
2026-04-29 10:19:24 +08:00
MingxuZhandMa Mingfei 2d27c38f13 Update XPU Docker runtime stack & hf_home config (#23820)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-29 10:03:11 +08:00
Xiaoyu Zhang 13afe8acdf [codex] Enable Qwen3-Next MoE all-reduce fusion (#23619) 2026-04-29 09:11:35 +08:00
Lianmin Zheng 14b4e6fa69 Support --model as alias for --model-path in CLI (#23894) 2026-04-28 17:23:27 -07:00
shuwenn 233048212a [HiCache][SPEC] fix: normalize storage prefetch key (#23631) 2026-04-28 15:53:20 -07:00
zijiexia 387c932dfc [Docs] update Docker image for Nemotron 3 Nano Omni (#23968) 2026-04-28 15:08:34 -07:00
Khoa Pham ddcacaf1bd Fix failing test_nvidia_nemotron_3_nano by fixing test_grouped_topk (#23874) 2026-04-28 15:03:58 -07:00
Alex NailsandClaude Opus 4.7 345fecc547 fix(bench): wire request_func in bench_long_context ContextWorkloadGenerator (#23898)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 14:45:51 -07:00
Liangsheng Yin cf0061da43 [Spec] Fix spec_accept_rate and unify accept/draft naming (#23530) 2026-04-28 14:40:04 -07:00
shuwennandClaude Opus 4.6 3e1c5e1b74 [HiCache][SPEC] fix: empty key after page alignment in match_prefix (#23387)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-28 14:06:55 -07:00