Liangsheng Yin
|
a8759dd9af
|
Fix killall.py crash when sglang is not yet installed (#21797)
|
2026-03-31 17:40:58 -07:00 |
|
Qiaolin Yu
|
5f6250769a
|
Reduce redundant speculative decoding CI tests (#21779)
|
2026-03-31 17:40:20 -07:00 |
|
Liangsheng Yin
|
b6fe0cca99
|
Switch MooncakeSpec to EAGLE3 + Llama-3.1 (#21794)
|
2026-03-31 17:12:20 -07:00 |
|
Liangsheng Yin
|
d047d41bad
|
Increase hicache eval to 200 examples (#21791)
|
2026-03-31 16:58:44 -07:00 |
|
Liangsheng Yin
|
7932e4c3e6
|
Remove redundant test_moe_eval_accuracy_large (#21787)
|
2026-03-31 16:45:03 -07:00 |
|
Liangsheng Yin
|
7581d814ae
|
Add CompletionSampler for non-chat eval in run_eval (#21785)
|
2026-03-31 16:33:07 -07:00 |
|
Yilong Zhao
|
1f7cee81da
|
[moe] add customized option to moe-a2a-backend (#21786)
|
2026-03-31 16:32:47 -07:00 |
|
Mohammad Miadh Angkad
|
883ba640b2
|
[CI] Remove more redundant PCG tests (#21554)
|
2026-03-31 16:25:30 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
8c66f4a90f
|
Add Trivy vulnerability scanning to nightly dev Docker builds (#21772)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-31 16:09:11 -07:00 |
|
Liangsheng Yin
|
53d8aa23ae
|
Cache nvidia wheels locally to skip repeated 830 MB downloads in CI (#21778)
|
2026-03-31 16:06:09 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
f60f2ccc10
|
[Fix] Fall back to triton MOE for GPT-OSS on Blackwell with driver >= 595 (#21780)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-31 15:52:10 -07:00 |
|
weireweire
|
9191b02eda
|
Fix cuda graph max bs capture upper bound (#21005)
|
2026-03-31 15:20:56 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
3c91ebdf55
|
[2/n] lora - Shared outer experts and support qwen3_30b_a3b_instruct (#21466)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 14:06:23 -07:00 |
|
Liangsheng Yin
|
f4505e2ee3
|
Fix ineffective is_base_mistral CI patch for HF API rate limiting (#21729)
|
2026-03-31 12:54:34 -07:00 |
|
Trevor Morris
|
b91f78d255
|
[bugfix] Fix rope theta config for MiniMax after transformers v5 update (#21241)
|
2026-03-31 11:37:03 -07:00 |
|
Michael
|
8d919bbd44
|
[AMD] Fix Handle missing rope_theta in get_rope_config for Grok-1 (#21518)
|
2026-03-31 10:58:12 -07:00 |
|
Zhangheng
|
91048b2a8e
|
[HiMambaTree]: Optimize mamba host lock mechanism (#21750)
|
2026-03-31 21:52:24 +08:00 |
|
 R0CKSTARandMick
|
e67dbf257a
|
[diffusion] fix: fix Wan2.2-I2V-A14B video max size issue(#21390)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-31 21:49:53 +08:00 |
|
 MickandClaude Opus 4.6
|
7790645b82
|
[diffusion] UX: replace deprecated ORJSONResponse with orjson_response (#21755)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-31 21:41:33 +08:00 |
|
JD
|
20d07c4384
|
Fix remote weight info nnode>1 and dp>1 (#17389)
|
2026-03-31 21:17:18 +08:00 |
|
Shangming Cai
|
ca2b2130ba
|
[PD] Tiny cleanup after KVReceiver refactor (#21760)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-03-31 21:07:57 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
c7adca9992
|
Fix kimi-linear launch server error (#21752)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-03-31 21:07:08 +08:00 |
|
Ke Bao
|
dbc97456ad
|
Enable evict swa with piecewise cuda graph (#21754)
|
2026-03-31 20:07:16 +08:00 |
|
Ke Bao
|
91ab75b505
|
[CI] Fix ring test timeout (#21751)
|
2026-03-31 18:26:08 +08:00 |
|
 weireweireandWeiliangl User
|
4455d17619
|
[PD] Refactor Disagg Conn and Fix Hang with total_request/total_tokens Balancing (#21299)
Co-authored-by: Weiliangl User <weiliangl@login-node.hosted.internal>
|
2026-03-31 18:01:50 +08:00 |
|
Ke Bao
|
acd37d8701
|
[CI] Fix rerun-test suite detection to skip commented registrations (#21753)
|
2026-03-31 18:00:53 +08:00 |
|
R0CKSTAR
|
6c03ae6fe2
|
[diffusion] fix: fix typo (#21746)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-03-31 17:51:46 +08:00 |
|
 xiaoqiandxiaoqi.31
|
a6a8b9b376
|
bugfix(model):fix deepstack index out of range error (#21727)
Co-authored-by: xiaoqi.31 <xiaoqi.31@jd.com>
|
2026-03-31 02:41:47 -07:00 |
|
Ke Bao
|
2456889f98
|
Rename rerun-ut to rerun-test (#21747)
|
2026-03-31 17:31:55 +08:00 |
|
Ke Bao
|
08d3c1a134
|
Fix disaggregation hybrid attention ci (#21745)
|
2026-03-31 16:22:54 +08:00 |
|
Baizhou Zhang
|
d52757fe97
|
[CI]Remove msgm-en and mmlu tests which cause timeout (#21733)
|
2026-03-31 01:10:05 -07:00 |
|
Thomas Wang
|
5628e908ae
|
[AMD] Use tgemm.mm for MoEGate router gemm in deepseek_v2.py (#21657)
|
2026-03-31 00:55:40 -07:00 |
|
xiazhahe
|
b4cb31f698
|
[NPU] fix conflict between empty_cache and use_mem_pool (#21507)
|
2026-03-31 15:37:33 +08:00 |
|
Mohammad Miadh Angkad
|
dd9c9c1b8e
|
Add explicit disable flag for FlashInfer allreduce fusion (#21446)
|
2026-03-31 00:15:44 -07:00 |
|
Yuhao Yang
|
68a4573627
|
[diffusion] fix: fix Flux.2 with tp(#21664)
|
2026-03-31 14:14:59 +08:00 |
|
jacky.cheng
|
8ba992411d
|
[AMD] Fix CI multimodal-gen-test-1-gpu-amd for gen model (#21621)
|
2026-03-30 23:02:20 -07:00 |
|
Jincong Chen
|
03e4f2858d
|
[Perf]Remove H2D for Qwen3.5 SpecV2 (#20864)
|
2026-03-31 11:54:58 +08:00 |
|
 Lewisand百麒
|
33e725b052
|
[Fix] Update supported custom_mem_pool types for mooncake (#21728)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
|
2026-03-31 11:18:30 +08:00 |
|
Xiaoyu Zhang
|
505eb312ec
|
Revert "DeepSeek-R1-0528-w4a8: DeepEP Low Latency Dispatch Adopts FP8 Communication" (#21719)
|
2026-03-31 10:22:01 +08:00 |
|
 Alison ShaoandAlison Shao
|
9b6bee2d40
|
Fix human-eval CI install on 5090 runners (#21714)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-03-30 18:53:09 -07:00 |
|
DarkSharpness
|
4e480982fa
|
[misc] multiprocess compilation to speed up test (#21483)
|
2026-03-31 08:56:37 +08:00 |
|
 Alison ShaoandAlison Shao
|
3650bfb199
|
Remove flashinfer wheel cache cleanup that deletes other versions (#21711)
Co-authored-by: Alison Shao <alison.shao@MacBook-Pro-D2W773R9CD.local>
|
2026-03-30 16:47:04 -07:00 |
|
kk
|
67c295b5f5
|
[AMD] fix performance regression issue when run gpt-oss with "--context-length 13824" (#21691)
|
2026-03-30 16:30:16 -07:00 |
|
jhchouuu
|
4b8456e266
|
[AMD][MoRI] bump MoRI to v0.1.0 (#21673)
|
2026-03-30 14:44:11 -07:00 |
|
 Zhai FeiyueandHaiShaw
|
daf697afda
|
[AMD] Add SGLANG_DISAGGREGATION_NUM_PRE_ALLOCATE_REQS env var for configurable KV transfer overlap (#20410)
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-03-30 14:37:16 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
d6029de6ad
|
[Bugfix][NPU] Skip FRACTAL_NZ format for MoE weights with unaligned dimensions (#21209)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-03-30 23:22:17 +03:00 |
|
Vedant V Jhaveri
|
4a9ffc3ab6
|
fix nemotron capture for non attention layers (#21436)
|
2026-03-30 12:50:49 -07:00 |
|
Yuxuan Zhang
|
ad064c2f4e
|
[GLM-V and GLM-4.7] Cast to FP32 before gate projection for GLM model. (#21660)
|
2026-03-30 12:25:27 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) yuefeng Wuandgemini-code-assist[bot]
|
a20d12ae96
|
[diffusion][doc]: add ring sp performance benchmark page (#20998)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-30 20:26:05 +03:00 |
|
Makcum888e
|
f4b0e9c64a
|
[diffusion] [NPU] support ring attention on NPU with FA (#21383)
|
2026-03-30 20:10:55 +03:00 |
|