Mick
|
4ac6fa0d87
|
[diffusion] fix: fix loading multiple ckpts with different precision for a same module (#22360)
|
2026-04-09 02:44:19 +08:00 |
|
Yihao Wang
|
a5ed507a16
|
[refactor] [asr] add transcription adapter for extensible ASR models support (#22181)
|
2026-04-09 01:19:37 +08:00 |
|
Yihao Wang
|
ae8da14ea3
|
[fix] [whisper] ensure inputs are moved to the correct device before processing. (#22293)
|
2026-04-08 23:45:42 +08:00 |
|
 Xiaoyu ZhangandMick
|
b5b2dbe05f
|
[Diffusion] Add diffusion NVFP4 scaled-mm correctness test (#22127)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-08 22:07:24 +08:00 |
|
Xiaoyu Zhang
|
ea119adc90
|
Refactor auto benchmark unit tests and fix CI bug (#22270)
|
2026-04-08 21:54:41 +08:00 |
|
zhaozx-cn
|
33c9cc8994
|
[NPU] fix qwen3.5 video processor (#22266)
|
2026-04-08 21:13:29 +08:00 |
|
 Alex NailsandClaude Opus 4.6
|
931dbceadc
|
[CI] Set RUNAI_STREAMER_MEMORY_LIMIT=0 for stage-b-test-1-gpu-small (#22346)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-08 02:23:35 -07:00 |
|
 Alison ShaoandAlison Shao
|
2ad5e6df12
|
[CI] Relax gpt-oss 4GPU accuracy threshold from 0.60 to 0.58 (#22237)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-08 02:20:23 -07:00 |
|
Fergus
|
413913763f
|
fix: wrap _import_static_state in inference_mode to fix resume on Blackwell (#21035)
|
2026-04-08 02:03:39 -07:00 |
|
 Vladislav Nosivskoyandhzh0425
|
79c82c5c42
|
[HiCache] Fix write_backup return type when parent not backed up (#22185)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-04-08 16:42:57 +08:00 |
|
Sundara Raman Ramachandran
|
712c8c5051
|
[Score API] Add SequenceClassification Model support (#22118)
|
2026-04-08 01:30:58 -07:00 |
|
Baizhou Zhang
|
213af1d4f7
|
Add CI tests for GLM-5 (#22285)
|
2026-04-08 01:05:36 -07:00 |
|
HuangJi
|
c3c13dd5e3
|
[diffusion] fix: make warmup image initialization rank-safe (#21817)
|
2026-04-08 15:51:09 +08:00 |
|
 Bingxu Chenandbingxche
|
de0cfed159
|
[AMD] Fix DLPack Error in Aiter flydsl GEMM by Detaching MoE Gate Weight (#22262)
Co-authored-by: bingxche <binxche@amd.com>
|
2026-04-07 23:42:10 -07:00 |
|
Артем Савкин
|
cd373667cd
|
[Bugfix] [NPU] Qwen3.5 with quantization fix (#21692)
|
2026-04-08 09:15:48 +03:00 |
|
Michael
|
db60a620db
|
[AMD] Add GLM-5-FP8 nightly performance benchmarks for MI30x and MI35x (#21710)
|
2026-04-07 22:43:14 -07:00 |
|
Thomas Wang
|
729b74d8dd
|
[AMD] Fix GLM-5 fp8 KV quant path dispatch on MI300 (#22314)
|
2026-04-07 21:16:02 -07:00 |
|
 Alison ShaoandAlison Shao
|
36f05810c9
|
[CI] Move manual-only nightly tests out of test/registered/ (#22298)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-07 21:03:52 -07:00 |
|
yuefeng Wu
|
4e4b4ac153
|
[NPU] enable index Cache for npu (#21502)
|
2026-04-08 11:45:17 +08:00 |
|
 Alex NailsandClaude Opus 4.6
|
493ec91cbe
|
[CI] Fix stage-b-test-1-gpu-large (0) timeout by reordering LoRA tests and using tokenizer from cache (#22292)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 20:00:44 -07:00 |
|
 Alison ShaoandAlison Shao
|
86e4542f35
|
Use dedicated runner label for deepep 8-GPU tests (#22309)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-07 19:58:54 -07:00 |
|
 
|
1c5c6dad5e
|
[tiny] Fix TOCTOU race in pause-aware weight update locking (#22304)
Co-authored-by: maocheng23 <maocheng@berkeley.edu>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 18:54:28 -07:00 |
|
Mick
|
eca62ab8f4
|
UX: clean loggings (#22174)
|
2026-04-08 09:46:38 +08:00 |
|
 Qiaolin YuandLiangsheng Yin
|
117508dcd7
|
Switch eagle_infer_beta to EAGLE3 (#22303)
Co-authored-by: Liangsheng Yin <hnyls2002@users.noreply.github.com>
|
2026-04-07 18:43:48 -07:00 |
|
 maocheng23andClaude Opus 4.6
|
6c2a759a04
|
[fix] Fix writer lock deadlock in update_weights_from_ipc during pause_generation (#22290)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 18:32:56 -07:00 |
|
Liangsheng Yin
|
8c3d80eabe
|
Only upload CUDA coredumps on test failure (#22301)
|
2026-04-07 18:07:28 -07:00 |
|
Kangyan-Zhou
|
dd73e9a62e
|
Revert "[CI] Update nightly test models for H200/B200 (#22288)" (#22297)
|
2026-04-07 17:04:06 -07:00 |
|
 
|
f6fc39569a
|
[CI] Migrate mgsm_en eval to gsm8k to remove openaipublic dependency (#21931)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-07 16:29:20 -07:00 |
|
Trevor Morris
|
7546d04c81
|
[NVIDIA] Enable FP4 flashinfer trtllm routed moe (#21240)
|
2026-04-07 16:16:29 -07:00 |
|
Liangsheng Yin
|
0e2a0260a1
|
Add fast-fail to multimodal-gen CI (#22284)
|
2026-04-07 15:56:12 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
e6652309c4
|
[CI] Update nightly test models for H200/B200 (#22288)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 15:44:52 -07:00 |
|
Thomas Wang
|
671fe73961
|
Reduce unnecessary kernels and copies in the NSA indexer (#22232)
|
2026-04-07 15:37:08 -07:00 |
|
    
|
f08726fd56
|
[Feature] Add DFLASH speculative decoding support (#22077)
Co-authored-by: Jian Chen <141193260+jianc99@users.noreply.github.com>
Co-authored-by: Zhijian Liu <5782437+zhijian-liu@users.noreply.github.com>
Co-authored-by: Richard Gong <8001209+gongy@users.noreply.github.com>
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com>
Co-authored-by: xm:D <38322020+xiaomin-d@users.noreply.github.com>
|
2026-04-07 14:48:51 -07:00 |
|
YC Yen-Ching Tseng
|
e14876742a
|
[AMD] Fix test_kimi_k25_mxfp4.py : stage-c-test-large-8-gpu-amd-mi35x (linux-mi35x-gpu-8, 1) (#22188)
|
2026-04-07 13:48:37 -07:00 |
|
Liangsheng Yin
|
cc35714b03
|
[tiny] migrate /get_server_info; print accept length in accuracy tests (#22282)
|
2026-04-07 13:08:35 -07:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
 
|
0c204fbd57
|
[HiSparse] Optimize the scheduling of decode backup. (#21932)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-04-07 10:34:58 -07:00 |
|
khalilzhk
|
6131fb5882
|
[NPU] enable mla prepare fused kernel only when being mla attn (#22024)
|
2026-04-08 00:49:16 +08:00 |
|
Ke Bao
|
be42fbbbd7
|
Support HTTP2 server (#21700)
|
2026-04-08 00:42:52 +08:00 |
|
shuwenn
|
ec5742f4ab
|
fix: Auto-correct page_size for Mamba no_buffer radix cache mode (#20538)
|
2026-04-08 00:19:31 +08:00 |
|
 Henson-Zh-Aliandhzh0425
|
727a182067
|
[Mamba] eliminate D2H if tracking mamba states (#20522)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-04-08 00:17:26 +08:00 |
|
YAMY
|
5ae00ecd48
|
[Disagg][NIXL] Support Mamba state slice transfer for heterogeneous TP (Step 2/2 for Qwen3.5) (#22240)
|
2026-04-07 23:47:31 +08:00 |
|
Mick
|
e7bc23cdab
|
[diffusion] CI: fix consistency check (#22251)
|
2026-04-07 23:43:18 +08:00 |
|
Ke Bao
|
fae90abf6e
|
Move ring test to nightly (#22267)
|
2026-04-07 21:56:39 +08:00 |
|
Yujun Dong
|
233f3e31bf
|
fix(pcg,mm): fix zeroing of input_embeds when replay PCG (#22229)
|
2026-04-07 20:33:17 +08:00 |
|
Xingyu Liu
|
98f38b14df
|
Add registration API for external linear attention backend (#21983)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
|
2026-04-07 02:47:40 -07:00 |
|
Nicolas Castet
|
490fa9fa44
|
[Perf] Restore torch.compile fusion for topk postprocessing (#21771)
|
2026-04-07 01:38:38 -07:00 |
|
Zhangheng
|
3d3a32c0b9
|
[HiSparse]: Add readme docs for HiSparse Feature (#22238)
|
2026-04-07 00:39:24 -07:00 |
|
YAMY
|
3148742ddb
|
[Disagg][NIXL] Fix heterogeneous TP KV transfer for non-MLA models (same logic with mooncake, Step 1/2 for Qwen3.5 support) (#22145)
|
2026-04-07 14:52:02 +08:00 |
|
Michael
|
ba78f6e0ef
|
[AMD] Add Qwen3.5-397B FP8 nightly perf benchmarks for MI30x and MI35x (#21669)
|
2026-04-06 23:46:00 -07:00 |
|