Commit Graph
13401 Commits
Author SHA1 Message Date
a0670b5ba3 [SPEC] feat: add adaptive speculative decoding metrics (#25940)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com>
2026-06-01 13:53:30 -07:00
Shu Wangandzijiexia 106092123f Update Qwen3-Coder docs_new NVIDIA guidance (#24435)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-01 13:38:34 -07:00
Khoa Pham da01f2974e [Log] include max_token_num and hidden_dim in FlashInfer workspace init log (#26605) 2026-06-01 13:26:52 -07:00
Mick f2beb7bc76 [diffusion] improve: avoid cosmos3 cpu float video postprocess (#26956) 2026-06-02 04:12:01 +08:00
Khoa PhamandClaude Opus 4.8 cb8a103b81 chore: add @pyc96 as codeowner for FrozenKVMTP module (#26953)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 13:06:58 -07:00
Mick 9a8ab2d22b [diffusion] fix: align cosmos3 text packing with official pipeline (#26950) 2026-06-02 02:07:17 +08:00
86afa21ca7 feat: optional caller-supplied mm_hashes on GenerateReqInput (#25300)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-06-01 20:04:37 +02:00
f6a5a1b59c [RL+VLM] Avoid retokenization drift for pre-tokenized (token-id) VLM requests (#26555)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: root <root@slurm-h200-209-231.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-01 09:58:14 -07:00
Mick 1988a2c9ea [diffusion] feat: improve cosmos3 serve API support (#26926) 2026-06-02 00:53:39 +08:00
Mick ed24e3aae8 [diffusion] feat: speed up png image output saving (#26947) 2026-06-02 00:43:02 +08:00
Ke Bao f59bbef841 Split SWA leaf to one window on insert (#26919) 2026-06-01 23:46:55 +08:00
Lukas HumbelandClaude Opus 4.7 d8a5a25c36 Refactor NIXL hicache. Add O_DIRECT support (#25173)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 17:28:53 +02:00
89feb18eb9 [diffusion] feat: allow --dit-cpu-offload with --dit-layerwise-offload (#26925)
Co-authored-by: Yiqi Yang <yiqi.yang@kiwiar.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-01 23:17:53 +08:00
Ke Bao 693adabff7 Fix Mamba2Metadata dropping has_mamba_track_mask (#26877) 2026-06-01 22:03:59 +08:00
zhaozx-cn fd16e05252 [NPU] fix npu profiler (#24835) 2026-06-01 21:44:24 +08:00
Bi Xue 6965fe0eec [sgl] Window-aware LRU refresh for SWA prefix cache in unified cache (#26615) 2026-06-01 19:35:18 +08:00
Chi McIsaac 931765e23e Do not cap DeepSeek V4 PD prefill by SWA pool size (#26607) 2026-06-01 19:29:45 +08:00
yiheng 1f8d3c7a42 [Speculative] [NPU] Adaptive-SD NPU support (#25644)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
2026-06-01 19:19:58 +08:00
Liangsheng Yin 1bff7a290f Refactor EAGLE infer tests: shared fixture + kits + overlap matrix (#26871) 2026-06-01 03:55:01 -07:00
Bingxu ChenandClaude Opus 4.8 89410b380b [AMD] Pin compressed-tensors==0.15.0 to fix ROCm nightly build (#26879)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 17:33:00 +08:00
giang_ng_tr 2394dede0e [EPD][Perf] Async image preprocessing and cross-request ViT batching for encode_server (#25669) 2026-06-01 16:52:12 +08:00
Yuan Luoandluoyuan.luo bc36231d65 [KDA] Support KDA packed decode (#26586)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-06-01 16:52:01 +08:00
amote-i d078cb72bd [NPU] [DOC] clarify Ascend NPU exclusive supported values for speculative args (#26903) 2026-06-01 16:40:11 +08:00
b14fba17d9 [AMD] make bypass-fastfail label also disable within-suite fast-fail (#26909)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Bingxu Chen <Bingxu.Chen@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-06-01 16:17:59 +08:00
Ke Bao cdd06011a1 Make unified tree SWA hicache tests faithful to write-through backup (#26870) 2026-06-01 16:08:57 +08:00
Mengxuan Xiong ff642ed936 [MoE] Extend kimi_k2_moe_fused_gate to support 256 experts (MiMo V2 Flash) (#26303) 2026-06-01 16:03:58 +08:00
YC Yen-Ching Tseng f60710a1d7 [AMD] Fix stage-b-test-large-8-gpu-mi35x-disaggregation-amd : switch CACHE_HOST to a fresh path to fix "No space left on device" (#26905) 2026-06-01 16:00:58 +08:00
ZeyuanChen2000 1d7e2f6fb8 [NPU] fix normal DeepEP mode num_tokens_per_rdma_rank error caused by none (#22972) 2026-06-01 15:32:14 +08:00
WingEdge777 20f47cfe8e fix : add sglang script as entry bin for runtime docker image (#26895) 2026-06-01 09:22:37 +02:00
Shangming Cai afd2d0b2f4 [PP][Bugfix] Handle input_ids assignment in prepare_for_extend (#26883) 2026-06-01 14:43:50 +08:00
Lianmin Zheng 53b8378307 Fix weights_checker checksum for 0-dim tensors and multi-GPU (#26863) 2026-05-31 21:27:03 -07:00
Mick 4b0453f814 [diffusion] CI: infer diffusion test sampling params from task type (#26530) 2026-06-01 11:57:41 +08:00
Lianmin ZhengandJaewon a779791b3f Add random-ids dataset, round-robin expert simulation, and kill_process_tree logging (#26862)
Co-authored-by: Jaewon <52840625+jaewonlee-fb@users.noreply.github.com>
2026-05-31 20:50:29 -07:00
Bruce Changlong Xu 11411aa49d [tokenizer] Surface scheduler load info (num_running_reqs / num_waiting_reqs) in meta_info (#24000) 2026-06-01 11:45:35 +08:00
liuxianglong17 3aaf8f115e fix test cases failed in nightly pipeline (#26714) 2026-06-01 11:33:47 +08:00
Qiaolin Yu 118465f5b5 [attn backend] Make spec_v2 seq_lens_cpu optional in trtllm_mla backend (#26824) 2026-05-31 20:29:50 -07:00
Chandrakant Khandelwal 61cc70e8aa Fixed incorrect indexing for slot 0 compatibility (#26481) 2026-06-01 10:46:21 +08:00
shadowxz109 4d20dc44fc 【NPU】add MiniMax2.5 best practice docs (#26725) 2026-06-01 10:09:12 +08:00
jundu 1ee189831f [CI] Bump xeon PR test unit tests timeout to 60 minutes (#26682) 2026-06-01 09:13:09 +08:00
huangtingwei 373cadc92e [bugfix] mooncake store double-tag bug fix (#26569) 2026-05-31 15:15:38 -07:00
shuwenn c06220159b [mem_cache][1/N] refactor: split allocator.py into allocator/ subpackage (#26675) 2026-05-31 20:31:29 +08:00
Ke Bao 972fbf7711 Skip flaky mamba extra_buffer disagg test (#26838) 2026-05-31 15:59:03 +08:00
Liangsheng Yin 585baa97f7 [core] Compute token_type_ids in ForwardBatch.init_new (#26797) 2026-05-31 00:54:49 -07:00
ybyang 1eadb7a173 Fix multi-tokenizer batch request output routing (health stuck at 503) (#26831) 2026-05-31 00:31:23 -07:00
Bruce Changlong Xu 376635c1e3 Fix routed-experts device buffer overflow under DP attention (#26123) 2026-05-30 19:11:06 -07:00
fzyzcjy f220c72929 Add periodic KV-canary stats logging and kernel-run-counter health check (#26821) 2026-05-31 10:00:19 +08:00
fzyzcjy 7dd19ae3d8 Add a sliding-window-attention divergence reporter for the KV-canary (#26820) 2026-05-31 09:59:28 +08:00
fzyzcjy ae9db7ff4b Add the KV-canary perturb modes and PD-disaggregation e2e tests (#26819) 2026-05-31 09:59:09 +08:00
fzyzcjy 6be4b32d8d Add token-id verification to the KV-canary (#26818) 2026-05-31 09:58:51 +08:00
fzyzcjy 0ca610a6df Add real-data KV verification to the KV-canary (#26817) 2026-05-31 09:58:32 +08:00