Mick
|
5bb981dce1
|
[diffusion] chore: read the cgroup this process is actually in (#35707)
|
2026-08-21 08:48:50 +08:00 |
|
Nan Jiang
|
f825d72936
|
[Sampling] Restore finite top-k requirement for sampling masks (#35205)
|
2026-08-20 16:55:25 -07:00 |
|
Jimmy Shong
|
14795dcb1a
|
[docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev image tag (#35767)
|
2026-08-20 23:43:55 +00:00 |
|
 ethcheandEthan Che
|
67f6ad61d9
|
fix(kernel) Fix Helion small-token prefill bug (#35197)
Co-authored-by: Ethan Che <eche@meta.com>
|
2026-08-20 16:41:59 -07:00 |
|
Yihao Wang
|
ba8e601358
|
[docs] fix note formatting in sglang-d documentation (#35761)
|
2026-08-20 16:39:39 -07:00 |
|
Jimmy Shong
|
1a138e13b9
|
[docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753)
|
2026-08-20 23:34:44 +00:00 |
|
 Shenxiu LiuandQiaolin Yu
|
779e593bd1
|
Fix _GenerationStreamAccumulator logprob_end off-by-one under retract (#26510)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
|
2026-08-20 16:03:13 -07:00 |
|
 Wenkai DuandHubert Lu
|
5a7b26c636
|
[AMD] [sgl-kernel] Bypass caches for peer traffic in ROCm custom all-reduce (#32832)
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
|
2026-08-20 15:24:18 -07:00 |
|
Jiajun Li
|
a4ef828207
|
fix(openai): avoid duplicate routed expert in response when return_meta_info = True (#35323)
|
2026-08-20 15:21:58 -07:00 |
|
Richard Gong
|
94907f05c4
|
Add CI permissions for four contributors (#35600)
|
2026-08-20 15:18:18 -07:00 |
|
Liangsheng Yin
|
0149f56e84
|
[CI] Gate /rerun-test on commenter trust and remove /rerun-stage (#35750)
|
2026-08-20 15:05:52 -07:00 |
|
Baizhou Zhang
|
92eeed41d7
|
[Docker] Defer CUDA 13 NCCL override until after dependency resolution (#35756)
|
2026-08-20 14:48:07 -07:00 |
|
 ishandhananiandAlex Nails
|
0f744b6848
|
feat: make mm_inputs msgpack-native (#29656)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-20 14:30:07 -07:00 |
|
Liangsheng Yin
|
5a100d9086
|
[misc] Trim restating comments and docstrings in srt/managers (#35622)
|
2026-08-20 14:18:40 -07:00 |
|
Lee Nau
|
ad367d72b0
|
[Kimi K3] Select FlashInfer MXFP4 for SM107 auto MoE (#35554)
|
2026-08-20 14:09:31 -07:00 |
|
Jimmy Shong
|
d9f6861359
|
[docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663)
|
2026-08-20 13:26:55 -07:00 |
|
 
|
eac91ac362
|
[Fix] Land the decode mamba checkpoint depth on the tree page under DCP (#35412)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-08-20 12:15:42 -07:00 |
|
 Junlin Wuandronnie_zheng
|
308bc1228b
|
📝 [NPU] Clean up quantization comments (#34829)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-08-20 22:03:58 +03:00 |
|
Connor Carpenter
|
61fa64ae7e
|
feat(grpc): expose KV event discovery metadata (#35714)
Signed-off-by: Connor Carpenter <connorc@nvidia.com>
|
2026-08-20 13:39:24 -05:00 |
|
Chao Shi
|
2ef0fe4669
|
TP/PP Consensus checker (#34406)
|
2026-08-21 01:36:03 +08:00 |
|
Shangming Cai
|
23cb04093c
|
fix(multimodal): keep LLaVA image fetch off the CPU-preprocess timeout budget (flaky test_mixed_batch) (#35700)
|
2026-08-21 01:26:30 +08:00 |
|
Ke Bao
|
ba97cc6397
|
Skip empty linear-attention state buffers in PD transfer (#35689)
|
2026-08-21 01:00:50 +08:00 |
|
R0CKSTAR
|
81df6f2c57
|
[MUSA] Harden CI dependencies and diffusion warmup (#35610)
|
2026-08-20 09:40:24 -07:00 |
|
Mick
|
be373395b4
|
[diffusion] feat: support out-of-tree models and pipelines (#35713)
|
2026-08-21 00:33:34 +08:00 |
|
Mick
|
7f8f030000
|
[diffusion] feat: let every layerwise component be configurable (#35688)
|
2026-08-20 22:38:05 +08:00 |
|
Xiaoyu Zhang
|
04444ee352
|
[diffusion] Refresh eager optimization skills and benchmark safeguards (#35679)
|
2026-08-20 22:03:57 +08:00 |
|
Shuwen Wang
|
9b249a25a1
|
test: switch the Inkling-Small NVFP4 deterministic suite to DSPARK (#35293)
|
2026-08-20 21:54:52 +08:00 |
|
silencejade
|
b03ac355e7
|
[NPU] [FIX] Fix non-contiguous parameter issue in FIA operator (#34936)
|
2026-08-20 20:43:40 +08:00 |
|
Estrella-xx
|
c98f1ccedb
|
[NPU]Ensure tensors allocated by empty_like are contiguous (#34935)
|
2026-08-20 20:40:57 +08:00 |
|
Mohammad Miadh Angkad
|
a4ffb996db
|
[Fix] Keep deterministic GDN prefill on Triton (#35632)
|
2026-08-20 19:38:53 +08:00 |
|
Mick
|
82c6fc2db9
|
[diffusion] quant: support pruned safetensors checkpoints for minimax-h3 (#35418)
|
2026-08-20 19:34:14 +08:00 |
|
Mick
|
97efc0507c
|
[diffusion] feat: plan pinned host memory against the cgroup cap not the machine (#35641)
|
2026-08-20 19:32:29 +08:00 |
|
Jimmy Shong
|
710267dc4c
|
[Quant] Load compressed-tensors kv_cache_scheme scales (#35455)
|
2026-08-20 19:17:59 +08:00 |
|
Mick
|
cf3813f4ce
|
[diffusion] feat: add weight source reader (#35668)
|
2026-08-20 18:37:29 +08:00 |
|
Mick
|
17313cf4b2
|
[diffusion] CI: add minimax-h3 ref2va audio consistency coverage and guard peak vram (#35511)
|
2026-08-20 17:22:05 +08:00 |
|
Mick
|
f1b9a1f42a
|
[diffusion] feat: support unverified short edge instead of rejecting it for minimax-h3 (#35664)
|
2026-08-20 16:52:44 +08:00 |
|
 Bingxu ChenandCursor
|
06ad7b2b0d
|
[AMD][CI] Run Both ROCm 7.2.4 and ROCm 7.2.0 Images on Nightly Test AMD (#35603)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-20 01:37:23 -07:00 |
|
Bingxu Chen
|
f386e2a471
|
[AMD][CI] Default the ROCm 7.2 PR gate to ROCm 7.2.4 Image (#35602)
|
2026-08-20 01:36:26 -07:00 |
|
 
|
21c88f8625
|
[diffusion] quant: support gguf (#35370)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-20 15:46:34 +08:00 |
|
YAMY
|
ae23423b46
|
Split TRTLLM MHA decode batches by KV sequence length (#34888)
|
2026-08-20 00:44:26 -07:00 |
|
Mick
|
b8996a5ab2
|
[diffusion] fix: keep large vocab tables in host memory under layerwise offload (#35626)
|
2026-08-20 15:20:06 +08:00 |
|
YC Yen-Ching Tseng
|
7ba3430365
|
Fix Grok-2 nightly: derive image-understanding capability from is_multimodal (#33730)
|
2026-08-19 23:56:51 -07:00 |
|
Baizhou Zhang
|
d287880a7a
|
Update deepep for SBO feature (#35450)
|
2026-08-19 23:37:13 -07:00 |
|
   
|
0bda0b168a
|
[Fix]: exclude SM120 from attn-res TMA dispatch (#35361)
Co-authored-by: 1BIN4 <1741738350@qq.com>
Co-authored-by: L-Ark <fliangae@connect.ust.hk>
Co-authored-by: Chikati <jxudn@connect.ust.hk>
Co-authored-by: mengzili <zilim@ust.hk>
|
2026-08-20 14:32:31 +08:00 |
|
Shangming Cai
|
628674a5c7
|
Remove unused MOONCAKE_COMPILE_ARG argument from Dockerfile (#35649)
|
2026-08-20 14:19:52 +08:00 |
|
 cctryandcctry
|
32d98aad13
|
[HiCache] Allow a retraction host pool smaller than the device pool (#35543)
Co-authored-by: cctry <cctry@fb.com>
|
2026-08-19 22:59:57 -07:00 |
|
  
|
50dae2d99d
|
Amd/dsv4 shared experts fusion top6 (#32340)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-08-19 22:57:58 -07:00 |
|
Michael
|
02b93e7e01
|
[AMD] Add GLM-5.2 MI35x nightly accuracy and perf benchmark (#32570)
|
2026-08-19 22:49:43 -07:00 |
|
 siyuandliusy58
|
58e327480a
|
update codeowner (#34802)
Co-authored-by: liusy58 <liusy58@smail.nju.edu.cn>
|
2026-08-19 22:43:15 -07:00 |
|
Mohammad Miadh Angkad
|
ba433bb462
|
[Docs] Update contribution guide (#35419)
|
2026-08-19 22:31:24 -07:00 |
|