Zhiqiang Xie
3dc91366ac
[HiCache] write_back: reclaim duplicated host copy first under host pressure ( #33777 )
2026-08-07 15:59:59 -07:00
Baizhou Zhang
eb3cc879e0
Install DeepEP from release wheels ( #33932 )
2026-08-07 15:38:44 -07:00
b2f9603f93
[bugfix] Stop/EOS inside a spec accept run beats the max_new_tokens finish ( #33758 )
...
Signed-off-by: Shiyan Deng <dsy842974287@meta.com >
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com >
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com >
Co-authored-by: Hanming Lu <hanminglu@meta.com >
2026-08-07 15:18:14 -07:00
Khoa Pham and Claude Opus 5
07297049e9
config: route DCP topology reads through get_parallel() ( #33925 )
...
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
2026-08-07 14:53:54 -07:00
Brayden Zhong and Brayden Zhong
b3ee679467
[MXFP8] Use FlashInfer CUTLASS for dense GEMM on SM120, delete Triton path ( #33208 )
...
Co-authored-by: Brayden Zhong <brayden@radixark.ai >
2026-08-07 14:30:43 -07:00
Dmitrii Sergeev
699fcdc936
Fix _pa_swa_prefill_lens off-by-one in FlashAttentionBackend ( #33379 )
2026-08-07 14:06:48 -07:00
Sam Shleifer and Claude Fable 5
62a28197c0
Autotune flashinfer extend buckets at warmup ( #32556 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-07 14:06:05 -07:00
Zhangheng
1480687cff
[CP]: Support CP V2 Strategy for dsv4 ( #33532 )
2026-08-07 14:03:27 -07:00
Sam Shleifer
8e7d361def
[perf] Compute input logprobs without materializing the full-vocab log-softmax ( #31958 )
2026-08-07 14:01:08 -07:00
Oguz Ulgen and Yinghai Lu
7f6b4cb94b
Add CUDA VMM multimodal feature transport ( #33899 )
...
Co-authored-by: Yinghai Lu <yinghai@meta.com >
2026-08-07 13:39:54 -07:00
3c51e29deb
Responses support ( #32689 )
...
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Harmya Bhatt <harmyacs@gmail.com >
Co-authored-by: harmya <harmya@modal.com >
Co-authored-by: Xinyuan <xinyuan@radixark.com >
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
2026-08-07 13:21:46 -07:00
Zhangheng
12de7fb1f6
Remove the HiMambaRadixTree that is no longer in use ( #33468 )
2026-08-07 23:01:40 +08:00
Ke Bao
bc8c037041
Fix vae fast path test after the gate refactor ( #33983 )
2026-08-07 20:45:19 +08:00
danielafrimi
4020bc95a7
Fix Nemotron W4A16 NVFP4 MoE backend ( #33543 )
...
Signed-off-by: dafrimi <dafrimi@nvidia.com >
2026-08-07 12:00:28 +00:00
Ke Bao
0756a1d2b0
Move SWA chunk-cap hatch tests into the registered suite ( #33975 )
2026-08-07 17:49:07 +08:00
Baizhou Zhang
5e60363960
Fix prefill CP graph overflow with larger bucket search ( #33906 )
2026-08-07 01:33:14 -07:00
Liangsheng Yin
7395ee833e
[CI] Share VLM engines and prune launch matrices on the per-commit H100/H200 suites ( #33944 )
2026-08-07 01:22:22 -07:00
Leon Gao
4d4f8023c4
[srt] Batch scheduler cache frees ( #33475 )
2026-08-06 21:49:59 -07:00
Liangsheng Yin
afa79330b8
[misc] Remove break-graph debug log; reclaim pid-less /dev/shm leaks in CI ( #33929 )
2026-08-06 21:18:01 -07:00
f9e6888b5a
[Distributed] Propagate semantic group names to PyTorch process groups ( #32900 )
...
Co-authored-by: jipengtian <jipengtian@xiaohongshu.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-08-06 20:38:12 -07:00
Nikhil Kulkarni and Claude Opus 5
fe6a05a8e8
fix: preserve priority for batched embedding requests ( #32977 )
...
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
2026-08-06 19:26:18 -07:00
Yuwei An
db8f3cdd11
fix(gdn): skip the -1 padding sentinel in the chunked extend kernel ( #33810 )
2026-08-06 18:48:47 -07:00
Hao Zhang and zhisbug
af7c62e337
Fix paged SWA retraction resume accounting ( #33794 )
...
Co-authored-by: zhisbug <1654062+zhisbug@users.noreply.github.com >
2026-08-06 16:32:54 -07:00
valechen
18e6c61c21
[AMD] perf: compact Triton extend-attention for ragged prefill (AMD/HIP-only) ( #29677 )
2026-08-06 14:46:10 -07:00
Brayden Zhong
dd7e4c91e2
Fix Mistral-Large-3 EAGLE draft skipping DeepseekV2Model.__init__ ( #33785 )
2026-08-06 13:59:28 -07:00
YAMY
2fc557254b
fix(PP): size the mamba pool per pipeline stage, not per whole model ( #33666 )
2026-08-06 13:10:43 -07:00
Mohammad Miadh Angkad and Brayden Zhong
434e646282
[Deps] Upgrade CUDA PyTorch stack to 2.13 ( #28836 )
...
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca >
2026-08-06 12:08:44 -07:00
Ziang Li and Brayden Zhong
4ad990ba7d
[ModelOpt FP4] Support online MoE weight quantization ( #33115 )
...
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca >
2026-08-06 11:01:55 -07:00
YAMY and Shangming Cai
05c7ebf64c
[Disagg][StagingBuffer][2/2] Support radix cache ( #30545 )
...
Co-authored-by: Shangming Cai <csmthu@gmail.com >
2026-08-06 23:59:35 +08:00
Mohammad Miadh Angkad
8a1637a479
Fix serving benchmark post-warmup cache flush race ( #33663 )
2026-08-06 15:27:10 +00:00
Xiaoyu Zhang and Claude Fable 5
591cfb0881
[diffusion] FLUX.2 bit-exact residual-gate fast path (H200 klein-4B 50-step denoise -1.2%) ( #33823 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-06 22:54:40 +08:00
Xiaoyu Zhang and Claude Fable 5
dd98c9572a
[diffusion] Generalize the FLUX.2 VAE decoder fast path to AutoencoderKL (Z-Image / FLUX.1) behind quality=high ( #33818 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-06 22:53:15 +08:00
Peng Wu
21225aba3d
[Scheduler] Fix to restrict the SWA chunk-cap escape hatch to true head-of-line livelock ( #32700 )
2026-08-06 22:09:04 +08:00
Mohammad Miadh Angkad
1a15cf1536
Gate multimodal feature transport by model capability ( #33653 )
2026-08-06 06:48:32 -07:00
Xiaoyu Zhang and Claude Fable 5
3654740347
[diffusion] ERNIE-Image bit-exact fused RMSNorm+scale/shift (H200 1024^2 e2e 15.63 -> 15.00 s, denoise -3.3%) ( #33854 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-06 19:58:44 +08:00
Xiaoyu Zhang and Claude Fable 5
295784723a
[diffusion] Ideogram 4: fuse RMSNorm modulate/gate chains via the Z-Image Triton suite behind quality=high (H200 e2e -2.9%/-3.4%) ( #33822 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-06 19:57:34 +08:00
Xiaoyu Zhang and Claude Fable 5
eff6a11350
[diffusion] FLUX.1 bit-exact residual-gate fast path + tanh-GELU epilogue behind quality=high (H200 e2e -1.1% lossless / -4.3% high) ( #33819 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-06 19:56:38 +08:00
f8f2870a84
Profiling Enhancements [1/3]: cuda graph profile traces ( #24370 )
...
Co-authored-by: Basit <mohbasit@ctr2-alola-ctrl-01.amd.com >
Co-authored-by: HAI <hixiao@gmail.com >
2026-08-06 03:19:35 -07:00
Yoray Zack
48f1b14fc7
test: fix NIXL EP Mooncake FT test ( #32638 )
2026-08-06 18:00:18 +08:00
YAMY and Lee Nau
8b29c90218
[NVIDIA] Enable CuTe DSL BF16 GEMM on SM107 ( #33617 )
...
Co-authored-by: Lee Nau <lnau@nvidia.com >
2026-08-06 02:06:08 -07:00
Bingxu Chen
dfe53232d7
[AMD] Move test_load_weights_from_remote_instance.py to extra CI ( #33809 )
2026-08-06 01:57:31 -07:00
Baizhou Zhang
0e584529f5
[CI] Remove profiling from nightly tests ( #33832 )
2026-08-06 01:16:08 -07:00
Xinyuan Tong
31c1e5943f
Facade DSA index-cache: MTP topk-reuse state + index-K storage ( #28609 )
2026-08-06 00:34:31 -07:00
Baizhou Zhang
5d1a0c7129
Revert "Warn on risky serving-time Triton work" ( #33826 )
2026-08-05 23:36:10 -07:00
32e5d788bd
[mm] rust-server: native multimodal processing for Qwen VL (integrate sglang-mm, e2e) ( #32365 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-08-05 23:35:50 -07:00
8e11feb68e
[HiCache] Support packed and sidecar draft caches for MTP/EAGLE/DSpark ( #30393 )
...
Co-authored-by: hjzhang <hjzhang89.gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org >
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com >
2026-08-06 14:31:11 +08:00
Xiaoyu Zhang and Claude Fable 5
b6876fc652
[diffusion] ERNIE-Image bit-exact residual-gate fast path (H200 1024^2 e2e 16.17 -> 15.75 s) ( #33734 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-08-06 13:43:09 +08:00
Liangsheng Yin
9bd1461757
[CI] Bound the CUDA graph capture range in test launches and lift the spec fixture's admission cap ( #33776 )
2026-08-05 20:00:30 -07:00
Mohammad Miadh Angkad
f33f6a522f
Relax GDN ReplaySSM fold test for Triton 3.7 ( #33780 )
2026-08-05 19:44:14 -07:00
Cheng Wan
4ea227fa91
config: the draft runner carries its own attention backend
...
`build_draft_tp_worker` built a `ServerArgs` variant whose only job was to make
four config reads answer with the draft's backend instead of the target's, and
published it for the duration of the build so the bags agreed. The backend is a
per-runner fact — target and draft coexist in one process — so it moves onto the
runner, and the variant and the construction-time publish both go away.
`ModelRunner` takes `draft_attention_backend` and resolves the runner's effective
value once (`resolve_draft_attention_backend`: the algorithm's resolved backend,
else `--speculative-draft-attention-backend`, else None for a target runner);
`TpModelWorker` threads it to both runner constructions.
`resolve_attention_backend_strs` reads it off the runner, and `ModelRunner`
stamps the resolved pair *before* building backends so a backend can read it
while it constructs — which is what the FlashInfer KV-access check needs now that
it no longer asks the config. `configure_kv_cache_dtype` and the draft backend
factory read the runner too.
One latent bug falls out: the non-hybrid branch of the backend build ignored the
resolved pair and re-read `server_args.attention_backend`, which is why the
variant had to set that field as well as the split pair. It now uses the value
that was resolved for the runner.
`draft_server_args_overrides` and the `preserve_config()` publish switch are
deleted; with them goes the last production `ServerArgs.derive` outside
pre-publish config building, and the last construction-time publish. The
chunked-prefix gate the target resolved simply stays in the bags, since nothing
re-projects them.
2026-08-05 19:32:24 -07:00