 Mohammad Miadh AngkadandMohammad Angkad
|
52c191da52
|
[Deps] Retire the CUDA 12 lane (#38404)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-10 16:58:09 -07:00 |
|
 sglang-botandsglang-bot
|
887c401e15
|
docs: sync LMSYS SGLang blog cards (#36773)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-09-10 18:37:14 +00:00 |
|
 zijiexiaandClaude Opus 5
|
4b7331fb77
|
Make the remaining DeepSeek-V4.1 NVIDIA cells start (#38861)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-10 11:04:53 -07:00 |
|
 Yuhao YangandClaude Code
|
a37ded1693
|
[Cookbook] DeepSeek-V4.1: add the HiCache L2 knob to the Playground (#38844)
Co-authored-by: Claude Code <noreply@anthropic.com>
|
2026-09-10 18:20:34 +08:00 |
|
 MickandMick Qian
|
c9c26d56b2
|
[diffusion] docs: sync snapshot and minimax-h3 subblock features (#38784)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-10 17:23:43 +08:00 |
|
 zijiexiaandClaude Opus 5
|
5caafd2118
|
Fix the DeepSeek-V4.1 reasoning example and mark the B300 cells verified (#38839)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-10 01:49:44 -07:00 |
|
Xinyuan Tong
|
9a2f17f41d
|
Add INT4 and FP4 lanes to the Ling-3.0-flash-VL cookbook (#38527)
|
2026-09-10 15:23:50 +08:00 |
|
   
|
1b77f498a0
|
[NVIDIA] Support flashinfer Mega Moe (#31470)
Co-authored-by: djns99 <40156487+djns99@users.noreply.github.com>
Co-authored-by: 云挚 <ningyunxiao.nyx@antgroup.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
|
2026-09-10 00:22:47 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
c0b790cf7f
|
Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-09-10 15:12:01 +08:00 |
|
 zijiexiaandClaude Opus 5
|
69777c4d36
|
Add DeepSeek-V4.1 Flash cookbook (#38802)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-09 23:15:30 -07:00 |
|
Thomas Wang
|
c415f977b8
|
[AMD] Update v4 args for agentic workload (#38677)
|
2026-09-09 21:41:05 -07:00 |
|
 MickandMick Qian
|
ce555ed82a
|
[diffusion] refactor: refactor utility ownership and document helper placement (#38699)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-10 09:11:01 +08:00 |
|
Alison Shao
|
a7e00b7576
|
[CI] Answer unrecognized slash commands instead of skipping silently (#38736)
|
2026-09-09 18:01:18 -07:00 |
|
William Hu
|
0084030179
|
Add Opt-In for GLM-5.3 Flash breakable prefill CUDA graphs (#38522)
|
2026-09-09 17:02:08 -07:00 |
|
Alison Shao
|
2948a62a6f
|
[CI] Add /run-full-ci and /run-extra-ci slash commands (#38734)
|
2026-09-09 13:56:38 -07:00 |
|
Even Zhou
|
dba34cc964
|
[NPU] Bump memfabric and sgl-kernel-npu versions in docs and pyproject_npu.toml (#38437)
|
2026-09-09 19:53:49 +08:00 |
|
  
|
daf66f6670
|
[Diffusion][SenseNova] support SenseNova-U1.5-8B-MoT (#36606)
Co-authored-by: wuyuefeng <wuyuefeng@noreply.gitcode.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-09-09 10:12:38 +03:00 |
|
Xiaoyu Zhang
|
a27be5ff62
|
[Diffusion] Enable lossless BCG for FLUX.1-dev (#38591)
|
2026-09-09 14:09:21 +08:00 |
|
 MickandMick Qian
|
00a9028e87
|
[diffusion] feat: add explicit snapshot-offload component residency (#38535)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-09 13:44:54 +08:00 |
|
Xiaoyu Zhang
|
03d06a764e
|
[Docs] Add measured JoyEcho H200 residency and BCG recipe (#38534)
|
2026-09-09 11:19:27 +08:00 |
|
Baizhou Zhang
|
e54ff1efb9
|
[CP V1 Deprecation 5/5] Update prefill CP documentation (#36230)
|
2026-09-08 19:27:49 -07:00 |
|
 MickandMick Qian
|
a8e45f16cc
|
[diffusion] feat: support mixed INT8 embeddings and Comfy NVFP4 encoders for minimax-h3 (#38506)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-09 08:36:43 +08:00 |
|
 MickandMick Qian
|
65400bb420
|
[diffusion] model: support MiniMax-H3 singularity hybrid checkpoints (#38455)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-09 08:35:39 +08:00 |
|
Cheng Wan
|
db272201a2
|
[Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement (#38375)
|
2026-09-08 16:42:12 -07:00 |
|
Xinyuan Tong
|
afe90a8bc9
|
Point Ling-3.0-flash-VL cookbook install section at the model image (#38539)
|
2026-09-08 10:51:33 -07:00 |
|
Xinyuan Tong
|
482e9f257b
|
Add Ling-3.0-flash-VL cookbook (#38434)
|
2026-09-08 23:00:24 +08:00 |
|
HZY
|
a6b542813f
|
fix(glm-5.2-nvfp4): bound Mooncake synchronous transfer batches (#32758)
|
2026-09-08 22:14:33 +08:00 |
|
Wuhen Duan
|
dfd9b5c2a4
|
[NPU] Enable non-greedy MTP sampling (#32495)
|
2026-09-08 11:18:01 +08:00 |
|
 
|
fba967ed9c
|
[diffusion] Support diffusion decoder parallel tiling for LTX-2.5 (#36026)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-08 10:01:06 +08:00 |
|
Siju Samuel
|
2358916d5a
|
[Feature][Intel XPU] Add memory saver support for Intel XPU via upstream torch_memory_saver (#29935)
|
2026-09-08 09:38:51 +08:00 |
|
 Kedar PotdarandPo-Han Huang
|
f4bbf12423
|
docs(cookbook): Qwen3.5 FP8 on B200/B300 — trtllm-gen MoE + symm mem (#38374)
Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
|
2026-09-08 09:12:49 +08:00 |
|
 
|
f4b75b5c36
|
docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-07 15:33:03 -07:00 |
|
 zijiexiaandClaude Opus 5
|
e4008de757
|
Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-07 21:30:17 +08:00 |
|
Mick
|
ba6d3df69a
|
[diffusion] doc: document verified GB300 and derived GB200 H3 recipes (#38296)
|
2026-09-07 16:32:29 +08:00 |
|
Mick
|
15d2cbcc90
|
[diffusion] CI: validate every repeated server request (#38185)
|
2026-09-07 14:36:50 +08:00 |
|
 
|
39a80354aa
|
[MUSA] Add installation guide and Dockerfile (#36709)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-09-06 20:13:53 -05:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) faceless voidandgithub-actions[bot]
|
30d0eb2ca9
|
[NPU] Adapt DFlash2 speculative decoding to Ascend NPUs (#35629)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-09-07 09:08:39 +08:00 |
|
Mick
|
f3d05644db
|
[diffusion] docs+skill: document which components to stream under layerwise offload (#35674)
|
2026-09-06 23:12:36 +08:00 |
|
 MickandClaude Fable 5
|
ade1da017f
|
[diffusion] docs: verify the DGX Spark H3 recipe (#37456)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-06 23:10:27 +08:00 |
|
Mick
|
938dc5621d
|
[diffusion] refactor: reuse plain state-dict loading without per-model classes (#38127)
|
2026-09-06 18:39:21 +08:00 |
|
Xiaoyu Zhang
|
d61378af77
|
docs(diffusion): add per-model tuning decision table to performance guide (#38148)
|
2026-09-06 15:26:40 +08:00 |
|
 
|
bd16c22a04
|
[diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-05 18:12:30 +08:00 |
|
Mick
|
0ea8378085
|
[diffusion] feat: support request-scoped skip-softmax attention (#37959)
|
2026-09-05 13:50:17 +08:00 |
|
 Shuwen WangandClaude Opus 5
|
4b44a1cde2
|
[Refactor] Let eviction policies take construction parameters (#37795)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-05 13:48:30 +08:00 |
|
 
|
09f542b23a
|
[CI] Add /rerun-test --changed to rerun every test file a PR modifies (#37618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-04 22:38:11 -07:00 |
|
 
|
92a4d8b5ee
|
Clean logging under --weight-loader-prefetch-checkpoints (#33930)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-09-04 20:05:53 -07:00 |
|
 zijiexiaandClaude Opus 5
|
3b64169f9d
|
[Cookbook] Kimi-K3: add measured B300 1x8 Unified 8k/1k speed numbers (#37878)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-04 17:44:00 -07:00 |
|
  
|
320bdd1ee2
|
[Docs] Document --retraction-policy, --return-hidden-states-mode, --language-model-only (#37989)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: mottopanikeiku <fcetin@hawk.iit.edu>
Co-authored-by: alp <falpercetin@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-05 05:26:28 +08:00 |
|
 alpandXinyuan Tong
|
3e873c2110
|
[Docs] Clarify OpenAI chat template defaults (#32172)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-05 05:22:05 +08:00 |
|
Faradawn Yang
|
c8ba8996c4
|
Update DeepSeek-V4 Pro for B200 FP4 agentic HiCache DSpark (#38026)
|
2026-09-04 11:09:24 -07:00 |
|