 
|
13537f8e20
|
Unskip Marlin NVFP4 tests (#27589)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: shaunkotek <shaunkotek@users.noreply.github.com>
|
2026-06-16 11:58:22 -07:00 |
|
Xinyuan Tong
|
33f205d8c5
|
docs(cookbook): fix GLM-5.2 thinking toggle kwarg + document reasoning effort (#28454)
|
2026-06-16 18:17:34 +00:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
799584e173
|
fix: get_processor fails when --tokenizer-path lacks model config.json (#25643)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-16 20:25:05 +03:00 |
|
 feliang-gitandxutizhou
|
92b42c8d8a
|
LPLB: linear-programming load balancer for MoE expert parallelism (#24515)
Co-authored-by: xutizhou <xutingz@nvidia.com>
|
2026-06-16 10:19:42 -07:00 |
|
Xinyuan Tong
|
00081a00d5
|
docs(cookbook): tune GLM-5.2 MTP to 5-1-6 and simplify launch flags (#28448)
|
2026-06-17 01:18:34 +08:00 |
|
Zhangheng
|
78b6a4fabf
|
[UnifiedTree]: Clean up some unused dead code. (#28389)
|
2026-06-16 21:53:35 +08:00 |
|
Xinyuan Tong
|
0cb6183432
|
docs(cookbook): add GLM-5.2 deployment cookbook (#28437)
|
2026-06-16 21:49:25 +08:00 |
|
 
|
265202cda2
|
fix(openai): validate assistant tool call arguments before chat template (#28035)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-06-16 12:21:23 +00:00 |
|
 Yinghai LuandLianmin Zheng
|
fcca4611fa
|
[CAR] Let custom allreduce support VMM based allocation (#27593)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-16 04:49:05 -07:00 |
|
 sglang-botandClaude Opus 4.8
|
12ebb35439
|
docs: refresh README News section and add Modal to adoption list (#28330)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-16 03:45:10 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
25e696aa8d
|
Fix Stage B CUDA CI (#28367)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-16 02:39:17 -07:00 |
|
 Wang, FangYuanandThomas Wang
|
a362ba9da3
|
[AMD] Feat: Add prefill context parallel support for deepseek v4 unified kv attention (#27928)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-06-16 02:00:51 -07:00 |
|
jacky.cheng
|
149fabcca7
|
[AMD] Fuse sigmoid + mul into single Triton kernel for shared expert gating (#27636)
|
2026-06-16 01:16:58 -07:00 |
|
YC Yen-Ching Tseng
|
102392df5b
|
[AMD][Fix] Skip EPLB topk remap when global server args are unset (#28404)
|
2026-06-16 01:13:22 -07:00 |
|
Bingxu Chen
|
0fc2bc4a8b
|
[AMD] Test DeepSeek V4 FlashMLA backend variants nightly (#28290)
|
2026-06-16 01:03:26 -07:00 |
|
Xiaoyu Zhang
|
c5b9106c1a
|
[perf] Use default torch compile mode for Wan2.2 T2V A14B (#28304)
|
2026-06-16 15:45:00 +08:00 |
|
Baizhou Zhang
|
77f327cb6e
|
[2/n] [CP] Add context parallel strategy abstractions (#27313)
|
2026-06-16 00:20:04 -07:00 |
|
Zhangheng
|
6c908b3a3a
|
[UnifiedTree]: Replace anonymous tuples with NamedTuples in UnifiedRadixCache (#28375)
|
2026-06-16 14:54:04 +08:00 |
|
Teng Ma
|
175336ff73
|
[Chore] update codeowner for mooncake store (#28377)
|
2026-06-16 14:51:34 +08:00 |
|
 Junlin Wuandronnie_zheng
|
2a8ea70059
|
✨ [llm][npu][quant] Add W8A8 MXFP8 quantization support for Qwen3 Dense on Ascend NPU (#22352)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-16 09:45:18 +03:00 |
|
Michael
|
72d962be88
|
[AMD] Fix jit-kernel-unit-test-amd: activation.cuh ROCm build + per_token CUDA-only (R165) (#27947)
|
2026-06-15 23:44:26 -07:00 |
|
Liangsheng Yin
|
556cf54d47
|
[Perf] Avoid per-decode-step host sync in min_new_tokens penalty (#28397)
|
2026-06-15 23:32:28 -07:00 |
|
 jy-song-hubandMick
|
637c9f780b
|
[diffusion] fix: add precision consistency layer (#27088)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-16 14:21:01 +08:00 |
|
 Qiaolin Yuandshuwenn
|
e068355831
|
[spec decoding] supports step 0 in adaptive spec decoding (updating draft kv cache without draft decoding) (#23994)
Co-authored-by: shuwenn <2508695655@qq.com>
|
2026-06-15 22:21:26 -07:00 |
|
Thomas Wang
|
800aaefc9e
|
[AMD] Annotate ATOM source for imported v4 unified attention kernels (#28392)
|
2026-06-15 22:12:12 -07:00 |
|
iridiumine
|
486ec150d0
|
[NPU] Add NPU fallback for fused Triton gating kernels (#28293)
|
2026-06-16 11:37:05 +08:00 |
|
 
|
ed9024e09f
|
[AMD] Fix AITER Scout workflow permissions (#28360)
Co-authored-by: bingxche <bingxche@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-16 11:21:39 +08:00 |
|
Shu Wang
|
32685874f3
|
Reenable MNNVL backend for FlashInfer allreduce fusion (#23402)
|
2026-06-15 20:19:15 -07:00 |
|
shuwenn
|
b23477af44
|
bench: infer tokenizer from serving model info (#28195)
|
2026-06-15 20:03:12 -07:00 |
|
zhangxiaolei
|
063ab89ac1
|
DeepSeek-V4 Online Compress support MTP (#26471)
|
2026-06-15 19:56:07 -07:00 |
|
huangtingwei
|
b5bcd76a41
|
[HiCache & JIT Kernel] Refactoring HiCache Write-Back Kernel (#21631)
|
2026-06-15 19:44:27 -07:00 |
|
Mick
|
a4a8a614b1
|
[diffusion] UX: suppress noisy diffusers torchao warning (#28317)
|
2026-06-16 10:32:21 +08:00 |
|
Michael
|
c6d9d73fd6
|
[Spec][test] fix(kv_canary): assert draft-extend-v2 oracle tokens in token_oracle test (#28325)
|
2026-06-15 19:14:09 -07:00 |
|
Cheng Wan
|
448af67a98
|
ci: run AMD and NPU PR tests on PRs not targeting main (#28368)
|
2026-06-15 19:10:51 -07:00 |
|
 sglang-botandsglang-bot
|
407d3a91db
|
docs: sync LMSYS SGLang blog cards (#28364)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-06-15 18:30:13 -07:00 |
|
Mick
|
01e45762ba
|
[diffusion] feat: use srt custom allreduce for tp groups (#28324)
|
2026-06-16 09:22:47 +08:00 |
|
 Zhiyao JiangandXinyu Jiang
|
2dd449ce5e
|
[AMD-miles] add amd-miles daily docker build workflow (#27765)
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
|
2026-06-16 08:22:46 +08:00 |
|
YAMY
|
b3be2e7402
|
[dsv4] Pad MLA decode q-heads to 64 (not full n_heads) for FlashMLA head64 kernel (#27954)
|
2026-06-15 17:18:10 -07:00 |
|
Liangsheng Yin
|
14f6348524
|
[Fix] Demote OpenAIServingResponses init failure log to one-line WARNING (#28349)
|
2026-06-15 16:58:33 -07:00 |
|
Hanming Lu
|
81dcb00673
|
[Spec v2] Use decode kernel for TRT-LLM MHA draft extend (#28241)
|
2026-06-15 16:57:08 -07:00 |
|
 Kangyan-ZhouandClaude Fable 5
|
cad43d3212
|
[CI] Reclaim leaked /dev/shm segments on server startup (#28089)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-06-15 16:14:08 -07:00 |
|
 zijiexiaandClaude Opus 4.8
|
7221be2cec
|
feat(cookbook): MTP --max-running-requests callout + skill sync (#28340)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-15 15:53:01 -07:00 |
|
Baizhou Zhang
|
4c0457f440
|
[misc] Update codeowner (#28342)
|
2026-06-15 14:42:01 -07:00 |
|
Yi Zhong
|
3a0dd69f8e
|
Minor refactorings to the LFM2.5 cookbook for accuracy (#28072)
|
2026-06-15 14:39:10 -07:00 |
|
jinhaosong-source
|
30d8ee87b0
|
Add OrcaRouter usage example (#28004)
|
2026-06-15 21:15:26 +00:00 |
|
Jia Guo
|
4ed698a491
|
fix(fa3): no NaN embeddings with fa_skip_kv_cache under piecewise CUDA graph (#27343)
|
2026-06-15 13:46:22 -07:00 |
|
YAMY
|
f870bf1ed0
|
[dsv4] Prewarm MHC prenorm kernel at startup (#27986)
|
2026-06-15 13:26:26 -07:00 |
|
 Lianmin ZhengandIan O'Connell
|
7e629a2f8c
|
Allow overriding tokenizer path in benchmark harness (#28280)
Co-authored-by: Ian O'Connell <ianoc@meta.com>
|
2026-06-15 13:07:50 -07:00 |
|
 
|
33719cfb31
|
[PD] Optimize SWA allocation (#28085)
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-15 11:01:55 -07:00 |
|
 YC Yen-Ching Tsengandbingxche
|
19e85868f6
|
[AMD] Point AITER scout at amd/aiter-ci (#28313)
Co-authored-by: bingxche <bingxche@amd.com>
|
2026-06-15 22:50:16 +08:00 |
|