shuwenn
|
108a183f6b
|
[mem_cache][6/N] refactor: move MHA host-pool into pool_host/mha.py (#30249)
|
2026-07-08 20:12:52 +08:00 |
|
 
|
fefc1743a9
|
Cute-DSL FP8 MQA logits (#25220)
Co-authored-by: Mindy Li <11663212+limin2021@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-06 23:34:07 -07:00 |
|
 Chetan Kumar VermaandMa Mingfei
|
b3ab56545b
|
Add Accuracy Benchmark for OCR models (#25364)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-06 16:18:43 +08:00 |
|
Xiaoyu Zhang
|
b276a9acee
|
chore: cleanup garbage code (#29770)
|
2026-07-02 16:14:01 +08:00 |
|
Liangsheng Yin
|
909123ddb8
|
[misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591)
|
2026-06-28 18:38:28 -07:00 |
|
 Aditya KamatandMick
|
1589603114
|
model: support baidu unlimited-ocr (#29186)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-27 23:36:19 +08:00 |
|
Xinyuan Tong
|
7c23d2255a
|
[minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712)
|
2026-06-22 13:10:43 -07:00 |
|
giang_ng_tr
|
b36360dc5b
|
[AMD][Perf] Split-KV flash-decode attention for EAGLE target-verify (Triton backend) (#27382)
|
2026-06-18 19:11:10 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
3340f4e3da
|
[GDN][KDA][mem_cache] int8 checkpoint pool for the linear-attn prefix cache (#28185)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-06-17 20:41:46 -07:00 |
|
huangtingwei
|
b5bcd76a41
|
[HiCache & JIT Kernel] Refactoring HiCache Write-Back Kernel (#21631)
|
2026-06-15 19:44:27 -07:00 |
|
Mohammad Miadh Angkad
|
91c63aeb4d
|
Fix stale CUDA graph benchmark and docs refs (#28041)
|
2026-06-13 21:51:42 -07:00 |
|
Yuhao Yang
|
aea0e30853
|
[4/N] Qwen3.5Opt: Overlap mamba verify update with draft extend (#26924)
|
2026-06-13 20:29:20 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
518e35fae7
|
[KDA] Add CuteDSL Prefill Kernel on SM100 (#27488)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-06-10 21:25:19 +08:00 |
|
 
|
c2eae96c56
|
MSCCL++ Integration (#22734)
Co-authored-by: Caio Rocha <caiorocha@microsof.com>
Co-authored-by: empyreus <rjsouza1995@gmail.com>
|
2026-06-08 21:13:13 -07:00 |
|
Hubert Lu
|
72929c7000
|
[AMD] Enable AITER custom all-gather on ROCm (#25093)
|
2026-06-02 15:57:37 -07:00 |
|
 Xiaoyu ZhangandBBuf
|
3ea1ba5b15
|
[GDN] Optimize prefill QKV split dispatch (#26206)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
|
2026-06-02 16:48:31 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
bc36231d65
|
[KDA] Support KDA packed decode (#26586)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-06-01 16:52:01 +08:00 |
|
Jialin Ouyang
|
98bc6f3c22
|
API Perf: Replace pydantic per-element validation with C loop validation (#26355)
|
2026-05-27 02:04:07 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
d34d4d9f5f
|
[GDN] Support SM100 CuTeDSL GDN Prefill Kernel (#26200)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-26 15:38:29 +08:00 |
|
 Jialin Ouyangandjialino
|
06c23d55b5
|
perf: migrate Req token-id storage to array.array('q') in Scheduler (#25098)
Co-authored-by: jialino <jialino@fb.com>
|
2026-05-22 10:51:07 -07:00 |
|
Pai Liu
|
cf1fd26d16
|
benchmark/lora: make number of LoRA adapters configurable (#25363)
|
2026-05-20 18:23:13 -07:00 |
|
  
|
ba2ffcf156
|
Add DeepSeekV4 fused MoE Triton autotune support (#25569)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
|
2026-05-18 21:35:32 +08:00 |
|
 sglang-botandClaude Opus 4.7
|
0a2615df24
|
chore: add vLLM SPDX copyright headers to ported files (#25182)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-13 15:17:30 -07:00 |
|
 Brayden Zhongandb8zhong
|
1d80a1a9fe
|
Use Cute-DSL NVFP4 quantization kernels (#23745)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-11 00:40:02 -07:00 |
|
RunningLeon
|
335dbd60b4
|
Support Intern-S2-Preview (#24875)
|
2026-05-10 22:17:30 +08:00 |
|
sky
|
c8bc23522f
|
Refactor: decouple segment tracking from comm registration (#21392)
Signed-off-by: wangfakang <fakangwang@gmail.com>
|
2026-05-06 17:07:58 +08:00 |
|
 egvenediktovandronnie_zheng
|
83bf5d6869
|
[NPU]TP Communications compression For Qwen3 models for NPU (#20520)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-02 14:29:11 +03:00 |
|
Clint
|
87dad74b33
|
Add benchmark/hicache/bench_warm_cache.py for exact warm-cache shared-prefix benchmarking (#24083)
|
2026-04-30 23:48:32 -07:00 |
|
Yihao Wang
|
0acc569edd
|
[Bench] extend MMMU answer extractor with explicit-commit patterns (#24084)
|
2026-04-30 19:08:08 -07:00 |
|
Yihao Wang
|
903e46d848
|
[Bench] fix bench_hf.py KeyError + reduce print spam + add --limit (#24079)
|
2026-04-29 11:43:37 -07:00 |
|
hhwxw
|
d9270b8c6a
|
fix(moe): relocate orphan tuned configs after #23019 (#24004)
|
2026-04-29 02:00:13 -07:00 |
|
Xinyuan Tong
|
832b4f59ed
|
[Bench] fix MMMU answer-extraction regex dropping multi-line responses (#23864)
|
2026-04-29 14:48:49 +08:00 |
|
 Alex NailsandClaude Opus 4.7
|
345fecc547
|
fix(bench): wire request_func in bench_long_context ContextWorkloadGenerator (#23898)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-28 14:45:51 -07:00 |
|
Muqi Li
|
69a71219cb
|
feat: tiny improve fp8_gemm tune usage (#23912)
|
2026-04-28 07:47:46 -04:00 |
|
Xiaoyu Zhang
|
6fbad22feb
|
Remove smoke wording from tests and comments (#23355)
|
2026-04-28 12:05:27 +08:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
6cf0b004ca
|
[MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-21 18:32:05 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
48daa831ea
|
[KDA] Fuse gate+cumsum and reuse chunk index for KDA (#23038)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-04-21 17:54:20 +08:00 |
|
 shuwennandQiaolin-Yu
|
b65799cf83
|
[SPEC][1/N] feat: add adaptive speculative_num_steps for EAGLE topk=1 (#21599)
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
|
2026-04-20 14:25:04 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
5f7aee726a
|
refactor(moe): de-duplicate triton MoE runner path into shared helpers (#23019)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-17 17:05:13 -07:00 |
|
 Hubert LuandHAI
|
edaa5973d4
|
[AMD][No-Merge] Simplify fused allreduce + RMSNorm and remove hidden_dim allowlist (#21986)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-04-11 23:47:08 -07:00 |
|
 satyamk7054andSatyam Kumar
|
059b287e25
|
Add offline auto-tuning for LoRA CSGMV kernel (#20391)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2026-04-10 13:10:43 -07:00 |
|
 Aditya SharmaandXinyuan Tong
|
f6e85676b5
|
model: support qwen3-asr (#22073)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-07 13:27:05 +08:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)    
|
2813cb6d9a
|
[New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
|
2026-04-06 20:24:44 -07:00 |
|
Xiaoyu Zhang
|
f3f7711dac
|
Fix Python 3.11 f-string lint error in deepgemm Blackwell benchmark (#22108)
|
2026-04-04 21:15:22 +08:00 |
|
harrisonlimh
|
9fa12d605a
|
Add dsv3 router gemm benchmark on blackwell (#17707)
|
2026-04-04 01:18:01 -07:00 |
|
Xiaoyu Zhang
|
ee9d922f5a
|
Revert "[Kernel] Fuse temperature + softmax in sampling for decode speedup" (#22046)
|
2026-04-03 21:32:08 +08:00 |
|
Mook
|
7a59e05dd1
|
[Kernel] Fuse temperature + softmax in sampling for decode speedup (#20501)
|
2026-04-02 12:46:36 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
03a87068ea
|
[KDA] Fuse scaled_dot_kkt + solve_tril + recompute_w_u for KDA (#21604)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-03-31 20:57:27 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
f0303fd07e
|
[Intel GPU] Enable DeepSeek R1 inference on XPU (#18461)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
|
2026-03-29 22:35:59 -07:00 |
|