Thomas Wang
|
d15a2dc72c
|
[AMD] dpsk-v4 swa loc cache support (#26931)
|
2026-06-01 22:37:07 -07:00 |
|
 
|
4226a6f13a
|
[AMD] Fix GPT-OSS MXFP4 accuracy on ROCm AITER path (#26884)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-06-01 22:30:43 -07:00 |
|
 Khoa PhamandCursor
|
08526c7fca
|
[Spec] FrozenKVMTP fold assistant seed into captured draft graph (#25539)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-01 22:27:25 -07:00 |
|
Kurkur
|
f27fa0da93
|
[NPU][Docs] Kimi-K2.5 best practice (#26774)
|
2026-06-02 13:14:14 +08:00 |
|
shuwenn
|
f40b2ca9d3
|
chore(CODEOWNERS): add allocator/ owners and @alphabetc1 to mem_cache (#27003)
|
2026-06-01 22:07:15 -07:00 |
|
 Ethan ZHUandZhangheng
|
594ec6335d
|
[Bug Fix][HiCache] Drop @lru_cache on UnifiedTreeNode.get_prefix_hash_values (#26939)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-06-02 12:38:28 +08:00 |
|
  
|
1c0019da75
|
[Docs] GLM-4.7 cookbook: add NVIDIA Blackwell (B200, GB200) + NVFP4 sections (#26384)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-01 20:47:51 -07:00 |
|
 popsiclexuandpopsiclexu
|
951fa05a09
|
[MoE] Support BF16 standard A2A with DeepGEMM runner (#26473)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
|
2026-06-01 20:40:38 -07:00 |
|
 Teng MaandZijie Xia
|
b562da0d9f
|
[PD] docs: clarify disaggregation IB device formats (#25521)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-06-02 11:38:33 +08:00 |
|
ybyang
|
9fe8b72912
|
Speed up DeepGEMM JIT warmup with per-PP-rank parallel compile (#26567)
|
2026-06-01 19:51:27 -07:00 |
|
 
|
0574d2b8a5
|
[NVIDIA] [GDN] Enable FlashInfer MTP verify on SM100+ (Blackwell) (#23273)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-01 18:56:42 -07:00 |
|
Liangsheng Yin
|
54143264bf
|
ci: disable cross-job fast-fail for run_all_tests dispatch (#26990)
|
2026-06-01 18:32:41 -07:00 |
|
  
|
98a1b58c47
|
docs(cookbook): port popular model usage guides into cookbook pages (#25813)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-06-01 17:41:49 -07:00 |
|
Liangsheng Yin
|
5e63200064
|
ci: full parallelism for run_all_tests dispatch (#26986)
|
2026-06-01 17:40:43 -07:00 |
|
Liangsheng Yin
|
f6d0beaca8
|
Revert "Support spec v2 tree drafting (eagle topk>1) with page_size==1" (#26981)
|
2026-06-01 17:16:44 -07:00 |
|
Glen Liu
|
167272e785
|
[LoRA] add lora chunked req test and fix (#23179)
|
2026-06-01 16:25:27 -07:00 |
|
 chenkaiyueandZhiqiang Xie
|
dff45411da
|
[HiCache] Prevent KV cache data loss when radix tree node is split b… (#16946)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-06-01 15:58:06 -07:00 |
|
Qiaolin Yu
|
4151a04d1a
|
[Perf][Spec Decoding] Skip cat/topk/sort/gather in draft_forward for topk=1 (#26424)
|
2026-06-01 15:37:47 -07:00 |
|
Liangsheng Yin
|
1d4ee060c2
|
Support spec v2 tree drafting (eagle topk>1) with page_size==1 (#26866)
|
2026-06-01 15:37:20 -07:00 |
|
Baizhou Zhang
|
2fdae94e46
|
docs: update RTX PRO 6000 deployment snippet (#26968)
|
2026-06-01 14:34:27 -07:00 |
|
Yongfei Xu
|
5700790c05
|
DeepSeek V4: Support context parallelism with fused MoE (non-DeepEP) (#24947)
|
2026-06-01 14:25:43 -07:00 |
|
Qiaolin Yu
|
3bce192bd2
|
[misc] update adaptive spec decoding code owners (#26965)
|
2026-06-01 14:08:09 -07:00 |
|
 eeechoandClaude Opus 4.6
|
524ba10eda
|
feat: SM120 (Blackwell Desktop) support for DeepSeek-V4 inference (#24692)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-06-01 14:05:20 -07:00 |
|
Liangsheng Yin
|
dfa1af99f5
|
Fix kill_process_tree reap wait crashing on pidfd EINVAL (#26964)
|
2026-06-01 13:56:59 -07:00 |
|
 
|
a0670b5ba3
|
[SPEC] feat: add adaptive speculative decoding metrics (#25940)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com>
|
2026-06-01 13:53:30 -07:00 |
|
 Shu Wangandzijiexia
|
106092123f
|
Update Qwen3-Coder docs_new NVIDIA guidance (#24435)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-06-01 13:38:34 -07:00 |
|
Khoa Pham
|
da01f2974e
|
[Log] include max_token_num and hidden_dim in FlashInfer workspace init log (#26605)
|
2026-06-01 13:26:52 -07:00 |
|
Mick
|
f2beb7bc76
|
[diffusion] improve: avoid cosmos3 cpu float video postprocess (#26956)
|
2026-06-02 04:12:01 +08:00 |
|
 Khoa PhamandClaude Opus 4.8
|
cb8a103b81
|
chore: add @pyc96 as codeowner for FrozenKVMTP module (#26953)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-01 13:06:58 -07:00 |
|
Mick
|
9a8ab2d22b
|
[diffusion] fix: align cosmos3 text packing with official pipeline (#26950)
|
2026-06-02 02:07:17 +08:00 |
|
 
|
86afa21ca7
|
feat: optional caller-supplied mm_hashes on GenerateReqInput (#25300)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
|
2026-06-01 20:04:37 +02:00 |
|
   
|
f6a5a1b59c
|
[RL+VLM] Avoid retokenization drift for pre-tokenized (token-id) VLM requests (#26555)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: root <root@slurm-h200-209-231.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-01 09:58:14 -07:00 |
|
Mick
|
1988a2c9ea
|
[diffusion] feat: improve cosmos3 serve API support (#26926)
|
2026-06-02 00:53:39 +08:00 |
|
Mick
|
ed24e3aae8
|
[diffusion] feat: speed up png image output saving (#26947)
|
2026-06-02 00:43:02 +08:00 |
|
Ke Bao
|
f59bbef841
|
Split SWA leaf to one window on insert (#26919)
|
2026-06-01 23:46:55 +08:00 |
|
 Lukas HumbelandClaude Opus 4.7
|
d8a5a25c36
|
Refactor NIXL hicache. Add O_DIRECT support (#25173)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-01 17:28:53 +02:00 |
|
 
|
89feb18eb9
|
[diffusion] feat: allow --dit-cpu-offload with --dit-layerwise-offload (#26925)
Co-authored-by: Yiqi Yang <yiqi.yang@kiwiar.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-06-01 23:17:53 +08:00 |
|
Ke Bao
|
693adabff7
|
Fix Mamba2Metadata dropping has_mamba_track_mask (#26877)
|
2026-06-01 22:03:59 +08:00 |
|
zhaozx-cn
|
fd16e05252
|
[NPU] fix npu profiler (#24835)
|
2026-06-01 21:44:24 +08:00 |
|
Bi Xue
|
6965fe0eec
|
[sgl] Window-aware LRU refresh for SWA prefix cache in unified cache (#26615)
|
2026-06-01 19:35:18 +08:00 |
|
Chi McIsaac
|
931765e23e
|
Do not cap DeepSeek V4 PD prefill by SWA pool size (#26607)
|
2026-06-01 19:29:45 +08:00 |
|
yiheng
|
1f8d3c7a42
|
[Speculative] [NPU] Adaptive-SD NPU support (#25644)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
|
2026-06-01 19:19:58 +08:00 |
|
Liangsheng Yin
|
1bff7a290f
|
Refactor EAGLE infer tests: shared fixture + kits + overlap matrix (#26871)
|
2026-06-01 03:55:01 -07:00 |
|
 Bingxu ChenandClaude Opus 4.8
|
89410b380b
|
[AMD] Pin compressed-tensors==0.15.0 to fix ROCm nightly build (#26879)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-01 17:33:00 +08:00 |
|
giang_ng_tr
|
2394dede0e
|
[EPD][Perf] Async image preprocessing and cross-request ViT batching for encode_server (#25669)
|
2026-06-01 16:52:12 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
bc36231d65
|
[KDA] Support KDA packed decode (#26586)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-06-01 16:52:01 +08:00 |
|
amote-i
|
d078cb72bd
|
[NPU] [DOC] clarify Ascend NPU exclusive supported values for speculative args (#26903)
|
2026-06-01 16:40:11 +08:00 |
|
  
|
b14fba17d9
|
[AMD] make bypass-fastfail label also disable within-suite fast-fail (#26909)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Bingxu Chen <Bingxu.Chen@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-06-01 16:17:59 +08:00 |
|
Ke Bao
|
cdd06011a1
|
Make unified tree SWA hicache tests faithful to write-through backup (#26870)
|
2026-06-01 16:08:57 +08:00 |
|
Mengxuan Xiong
|
ff642ed936
|
[MoE] Extend kimi_k2_moe_fused_gate to support 256 experts (MiMo V2 Flash) (#26303)
|
2026-06-01 16:03:58 +08:00 |
|