YAMY
|
5f12839591
|
[Fix] Support Kimi-K3 ModelOpt mixed NVFP4/FP8 checkpoint (#35077)
|
2026-08-19 08:13:45 -07:00 |
|
 Shuwen Wangandhzh0425
|
41c018a9ec
|
[UnifiedTree] feat: support runtime attach/detach (#35269)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-08-19 22:49:08 +08:00 |
|
 Rohit Kumar SinghandSingh
|
3e5ce26c2d
|
fix: fix transcription & audio-understanding for ASR/audio/speech models (#32611)
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
|
2026-08-19 18:44:23 +08:00 |
|
 
|
f446e853e7
|
[AMD] DeepSeek-V4: route decode wo_a bf16 batched matmul to aiter batched_gemm_bf16 (#33313)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-08-19 03:04:48 -07:00 |
|
karverma-amd
|
ce1830c59b
|
[AMD] DeepSeek-V4 MI355X: eliminate bpreshuffle fp8-scale relayout copy in dense w8a8 linear (#33165)
|
2026-08-19 03:02:40 -07:00 |
|
 HeYaoandHeYao
|
f22442d3a4
|
Add three new test cases (#35502)
Co-authored-by: HeYao <heyao@example.com>
|
2026-08-19 17:59:45 +08:00 |
|
Shangming Cai
|
adca19c497
|
[PD] Deferred decode-side KV release for the NIXL backend (#35360)
|
2026-08-19 17:29:45 +08:00 |
|
YAMY
|
aa215e5523
|
[PD] Overlap prefill DP-rank bootstrap queries (#35071)
|
2026-08-19 17:25:52 +08:00 |
|
 Khoa PhamandCursor
|
0e4a09480c
|
[HiCache] Support DCP with DSpark (#35221)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-19 01:42:19 -07:00 |
|
     
|
8a1e6e4e46
|
Qwen3.8-27B Model Support (#34859)
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
|
2026-08-19 16:31:43 +08:00 |
|
Liangsheng Yin
|
ccbe380028
|
[CI] Trim the base-c 4-gpu-h100 stage from 5 shards to 4 (#35407)
|
2026-08-19 00:48:07 -07:00 |
|
 MickandClaude Opus 5
|
c0c87e0547
|
VLM: feed the packed qkv projection output to vision backends uncopied (#35336)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-19 14:39:52 +08:00 |
|
 Alison ShaoandXinyuan Tong
|
3391ab3712
|
[Constrained] Support MistralCommon tokenizers in the XGrammar backend (#35215)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-08-19 12:34:48 +08:00 |
|
Shuwen Wang
|
a72d8d29d3
|
Fix HiCache PP sync test fixture (#35446)
|
2026-08-19 12:25:17 +08:00 |
|
 Jimmy ShongandClaude Fable 5
|
c863760ae1
|
[Fix] DCP: advertise the logical KV-event block size (#35298)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-18 20:55:11 -07:00 |
|
 Muqi LiandCodex
|
88f6074392
|
feat(api): add sglext_spec (#33518)
Signed-off-by: Muqi Li <muqi1029@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
|
2026-08-19 11:54:18 +08:00 |
|
luoroger37
|
3b065a56b0
|
[HiCache] Batch PP write and load completion sync (#33473)
|
2026-08-19 11:27:27 +08:00 |
|
HuangJi
|
ee1f2e8dfd
|
[diffusion][Minimax H3]support subblock sparse attention on SM90 (#34680)
|
2026-08-19 10:31:45 +08:00 |
|
Zhiqiang Xie
|
977412ae61
|
[HiCache] Buffer-only mode for HiCache host memory layer (#34798)
|
2026-08-18 19:21:24 -07:00 |
|
Mick
|
4cef72faee
|
[diffusion] refactor: reuse srt qwen vision and text modules (#35006)
|
2026-08-19 10:12:44 +08:00 |
|
Chunyuan WU
|
58c5bee3ac
|
Fix DP attention on CPU (#12961)
|
2026-08-19 09:56:21 +08:00 |
|
Mick
|
ef490853bb
|
quant: extract shared checkpoint quant metadata resolver (#35172)
|
2026-08-19 08:26:41 +08:00 |
|
 MickandClaude Opus 5
|
77fc5c128e
|
[perf] overlap page preprocessing, pack the vit, enable prefill CUDA graph for paddle-ocr (#35318)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-19 08:20:55 +08:00 |
|
  
|
c14312a664
|
[Spec] DFlash2: local convolution + candidate selector (#35371)
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-08-18 17:07:28 -07:00 |
|
Baizhou Zhang
|
cfc6dfb364
|
Apply latest DeepEP branch (#34923)
|
2026-08-18 16:24:50 -07:00 |
|
 cctryandcctry
|
37c09ff3d8
|
[Memory] Borrow CUDA graph pool storage for EAGLE sampling (#35375)
Co-authored-by: cctry <cctry@fb.com>
|
2026-08-18 16:19:41 -07:00 |
|
Liangsheng Yin
|
7ebaa98f81
|
[Fix] Assert the page-aligned SWA evict floor on both PD decode prealloc paths (#35396)
|
2026-08-18 15:48:55 -07:00 |
|
Liangsheng Yin
|
79dfef390b
|
[Spec] Page-align the DFLASH decode KV reservation (#35265)
|
2026-08-18 13:36:04 -07:00 |
|
 Yuwei AnandClaude Opus 5
|
955704544c
|
[Fix] Skip padded state slots in the chunked GDN kernel (#33431)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-18 13:07:19 -07:00 |
|
  
|
307a90f6d3
|
Stop losing Kimi-K3 tool calls to reasoning, constraint conflicts, and truncation (#34881)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-08-18 12:42:56 -07:00 |
|
 Wei YangandWei Yang
|
8bb106cee9
|
Fix NIXL cleaner grouping for hybrid cache keys (#35130)
Co-authored-by: Wei Yang <yawei@microsoft.com>
|
2026-08-18 13:55:51 -05:00 |
|
Hanming Lu
|
526af15845
|
[Metrics] Discount queued prefill load by recent cache hits when waiting-queue matching is off (#35248)
|
2026-08-18 11:27:21 -07:00 |
|
 Xingyu Liuandxingyuliu
|
7dcaf11987
|
[Fix] Select custom all-reduce v2 by topology capability (#35061)
Co-authored-by: xingyuliu <xingyuliu@fb.com>
|
2026-08-18 10:30:10 -07:00 |
|
Ke Bao
|
480033def0
|
Refactor kv cache event mixin into a recorder (#35164)
|
2026-08-19 01:20:17 +08:00 |
|
jain-ria
|
83d7d45330
|
fix: preserve output logprobs without input logprobs (#34627)
Signed-off-by: jain-ria <riajain@NVIDIA.com>
|
2026-08-18 12:07:13 -05:00 |
|
 Shangming CaiandClaude Opus 4.8
|
97dedd1ce9
|
[PD] Deferred decode-side KV release for aborts mid-transfer (#35049)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-08-18 23:56:33 +08:00 |
|
pllimax
|
499e90a125
|
[NPU CI] Reorganize test output/log directory structure with workflow context (#33685)
|
2026-08-18 23:46:00 +08:00 |
|
paulzhang-tm
|
0065fbfae1
|
[Scheduler] Cap prefill-delayer queue target by admission capacity (#35191)
|
2026-08-18 21:40:13 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 5
|
ae6945e112
|
[kernels] Reorganize ops/diffusion by operator domain behind a lazy facade (#35114)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-18 20:37:43 +08:00 |
|
 HeYaoandHeYao
|
7605529bdf
|
Add deepseek_v4_flash_w8a8_8p_in32k_out1k_50ms (#35162)
Co-authored-by: HeYao <heyao@example.com>
|
2026-08-18 19:41:25 +08:00 |
|
karverma-amd
|
24d625698d
|
[AMD] feat(moe): fold padded-topk_ids fill into fused shared-experts append+remap (#31370)
|
2026-08-18 02:32:57 -07:00 |
|
vijay-kodamalla
|
880ab72f34
|
test: extend NVFP4 Marlin tests to SM120 (#34327)
|
2026-08-18 16:24:00 +08:00 |
|
mohbasit
|
fc0b95e7ba
|
Profiling Enhancements [2/3]: detailed execution step annotations (#24911)
|
2026-08-18 01:09:07 -07:00 |
|
Lianmin Zheng
|
a779a2a2a5
|
[Chore] Move version tag helper to release scripts (#35196)
|
2026-08-18 00:11:12 -07:00 |
|
Jiajun Li
|
e6df23f3c2
|
refactor: rename chat response token IDs (#35225)
|
2026-08-17 23:07:43 -07:00 |
|
Shuwen Wang
|
0077f84d37
|
[mem_cache][8/N] refactor: move MambaPoolHost to pool_host.mamba (#31180)
|
2026-08-18 05:22:45 +00:00 |
|
Colin Z
|
ea27e3ddab
|
[AMD] Fix Quark Shared Experts Fusion Gate after load-time-override Removal (#35200)
|
2026-08-17 22:10:01 -07:00 |
|
 
|
8ea5229d42
|
[AMD] Add the Kimi-K3 MI35x perf benchmarks in nightly (#34985)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
|
2026-08-17 21:23:51 -07:00 |
|
Ke Bao
|
d528192bf9
|
Skip inkling sheared bias under batch invariance (#35161)
|
2026-08-18 12:22:58 +08:00 |
|
 amd-danli103andThomas Wang
|
d01812d89e
|
[AMD] Optimize KIMI-K3 with Triton MLA decode kernel by tuning the stage-1 geometry for gfx950 (#34580)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-08-17 21:16:14 -07:00 |
|