18694 Commits
Author SHA1 Message Date
Thomas Wang 729b74d8dd [AMD] Fix GLM-5 fp8 KV quant path dispatch on MI300 (#22314) 2026-04-07 21:16:02 -07:00
Alison ShaoandAlison Shao 36f05810c9 [CI] Move manual-only nightly tests out of test/registered/ (#22298)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
2026-04-07 21:03:52 -07:00
yuefeng Wu 4e4b4ac153 [NPU] enable index Cache for npu (#21502) 2026-04-08 11:45:17 +08:00
Alex NailsandClaude Opus 4.6 493ec91cbe [CI] Fix stage-b-test-1-gpu-large (0) timeout by reordering LoRA tests and using tokenizer from cache (#22292)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-07 20:00:44 -07:00
Alison ShaoandAlison Shao 86e4542f35 Use dedicated runner label for deepep 8-GPU tests (#22309)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
2026-04-07 19:58:54 -07:00
1c5c6dad5e [tiny] Fix TOCTOU race in pause-aware weight update locking (#22304)
Co-authored-by: maocheng23 <maocheng@berkeley.edu>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-07 18:54:28 -07:00
Mick eca62ab8f4 UX: clean loggings (#22174) 2026-04-08 09:46:38 +08:00
Qiaolin YuandLiangsheng Yin 117508dcd7 Switch eagle_infer_beta to EAGLE3 (#22303)
Co-authored-by: Liangsheng Yin <hnyls2002@users.noreply.github.com>
2026-04-07 18:43:48 -07:00
maocheng23andClaude Opus 4.6 6c2a759a04 [fix] Fix writer lock deadlock in update_weights_from_ipc during pause_generation (#22290)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-07 18:32:56 -07:00
Liangsheng Yin 8c3d80eabe Only upload CUDA coredumps on test failure (#22301) 2026-04-07 18:07:28 -07:00
Kangyan-Zhou dd73e9a62e Revert "[CI] Update nightly test models for H200/B200 (#22288)" (#22297) 2026-04-07 17:04:06 -07:00
f6fc39569a [CI] Migrate mgsm_en eval to gsm8k to remove openaipublic dependency (#21931)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-04-07 16:29:20 -07:00
Trevor Morris 7546d04c81 [NVIDIA] Enable FP4 flashinfer trtllm routed moe (#21240) 2026-04-07 16:16:29 -07:00
Liangsheng Yin 0e2a0260a1 Add fast-fail to multimodal-gen CI (#22284) 2026-04-07 15:56:12 -07:00
Kangyan-ZhouandClaude Opus 4.6 e6652309c4 [CI] Update nightly test models for H200/B200 (#22288)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-07 15:44:52 -07:00
Thomas Wang 671fe73961 Reduce unnecessary kernels and copies in the NSA indexer (#22232) 2026-04-07 15:37:08 -07:00
f08726fd56 [Feature] Add DFLASH speculative decoding support (#22077)
Co-authored-by: Jian Chen <141193260+jianc99@users.noreply.github.com>
Co-authored-by: Zhijian Liu <5782437+zhijian-liu@users.noreply.github.com>
Co-authored-by: Richard Gong <8001209+gongy@users.noreply.github.com>
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com>
Co-authored-by: xm:D <38322020+xiaomin-d@users.noreply.github.com>
2026-04-07 14:48:51 -07:00
YC Yen-Ching Tseng e14876742a [AMD] Fix test_kimi_k25_mxfp4.py : stage-c-test-large-8-gpu-amd-mi35x (linux-mi35x-gpu-8, 1) (#22188) 2026-04-07 13:48:37 -07:00
Liangsheng Yin cc35714b03 [tiny] migrate /get_server_info; print accept length in accuracy tests (#22282) 2026-04-07 13:08:35 -07:00
Rain Jiang 1a8eb890f6 Kernels community fa3 (#20796) 2026-04-07 12:48:44 -07:00
0c204fbd57 [HiSparse] Optimize the scheduling of decode backup. (#21932)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-04-07 10:34:58 -07:00
khalilzhk 6131fb5882 [NPU] enable mla prepare fused kernel only when being mla attn (#22024) 2026-04-08 00:49:16 +08:00
Ke Bao be42fbbbd7 Support HTTP2 server (#21700) 2026-04-08 00:42:52 +08:00
shuwenn ec5742f4ab fix: Auto-correct page_size for Mamba no_buffer radix cache mode (#20538) 2026-04-08 00:19:31 +08:00
Henson-Zh-Aliandhzh0425 727a182067 [Mamba] eliminate D2H if tracking mamba states (#20522)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-08 00:17:26 +08:00
YAMY 5ae00ecd48 [Disagg][NIXL] Support Mamba state slice transfer for heterogeneous TP (Step 2/2 for Qwen3.5) (#22240) 2026-04-07 23:47:31 +08:00
Mick e7bc23cdab [diffusion] CI: fix consistency check (#22251) 2026-04-07 23:43:18 +08:00
Ke Bao fae90abf6e Move ring test to nightly (#22267) 2026-04-07 21:56:39 +08:00
Yujun Dong 233f3e31bf fix(pcg,mm): fix zeroing of input_embeds when replay PCG (#22229) 2026-04-07 20:33:17 +08:00
Xingyu Liu 98f38b14df Add registration API for external linear attention backend (#21983)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
2026-04-07 02:47:40 -07:00
Nicolas Castet 490fa9fa44 [Perf] Restore torch.compile fusion for topk postprocessing (#21771) 2026-04-07 01:38:38 -07:00
Zhangheng 3d3a32c0b9 [HiSparse]: Add readme docs for HiSparse Feature (#22238) 2026-04-07 00:39:24 -07:00
YAMY 3148742ddb [Disagg][NIXL] Fix heterogeneous TP KV transfer for non-MLA models (same logic with mooncake, Step 1/2 for Qwen3.5 support) (#22145) 2026-04-07 14:52:02 +08:00
Michael ba78f6e0ef [AMD] Add Qwen3.5-397B FP8 nightly perf benchmarks for MI30x and MI35x (#21669) 2026-04-06 23:46:00 -07:00
amote-i 3f7dfba419 fix qwen2_5_math_rm_72b (#21295) 2026-04-07 14:36:57 +08:00
Aditya SharmaandXinyuan Tong f6e85676b5 model: support qwen3-asr (#22073)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-04-07 13:27:05 +08:00
Chang Min BarkandR0CKSTAR a757c1e3fb [Apple Silicon] [MLX] Add mlx and mlx-lm dependencies (#22162)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-04-07 11:36:43 +08:00
2813cb6d9a [New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
2026-04-06 20:24:44 -07:00
Liangsheng Yin be0277f9a0 [Spec][Ngram] Add output-as-corpus accept length benchmark for external SAM (#22199) 2026-04-06 19:09:52 -07:00
jianzhao-xuandJianzhao Xu 73fc87a74f fix(grok): adapt huihui-ai/grok-2 (#21522)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
2026-04-07 10:04:41 +08:00
Lianmin ZhengandClaude Opus 4.6 494bb86169 Cache sub-objects in __getitem__ to ensure identity stability (#22184)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 18:53:38 -07:00
Prozac614anddaiweitao ef2d4013d7 [diffusion] CI: add consistency test (#15236)
Co-authored-by: daiweitao <dwti614707404@163.com>
2026-04-07 09:50:23 +08:00
Liangsheng Yin e4b1366a46 [Spec][Ngram] Support multiple SAMs with dynamic HTTP API (#22203) 2026-04-06 18:49:22 -07:00
Liangsheng Yin 49cb7d546e Move hash utils out of hicache_storage to break CUDA import chain (#22214) 2026-04-06 18:16:40 -07:00
AichenF 5e2b0f860c [diffusion] perf: replace Conv3d with reshape + F.linear in PatchEmbed (#21014) 2026-04-07 09:12:59 +08:00
shadowxz109 ae38b24cc3 [NPU] Support dp-attention for MiniMax2.5 (#20919) 2026-04-07 08:55:37 +08:00
Trevor Morris 5cc246e095 Fix extra calls to get_numa_node_if_available to clean up logs (#21781) 2026-04-06 16:18:40 -07:00
Trevor Morris 56266de624 [CI] Add basic unit test for Minimax-M2.5 (#21792) 2026-04-06 15:48:33 -07:00
Alison ShaoandAlison Shao 6f1412f4f5 [CI] Relax transformers MMLU threshold from 0.65 to 0.64 (#22210)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
2026-04-06 15:32:09 -07:00
Lianmin Zheng a80961333b Clean up req_time_stats: reduce overhead and simplify (#22186) 2026-04-06 14:20:51 -07:00