Mick
|
4beb157e87
|
[diffusion] doc: define native diffusion model integration contract (#34952)
|
2026-08-15 23:37:34 +08:00 |
|
 
|
5c0ace30c0
|
[diffusion] model: support ltx-2.5 (#34471)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-15 23:36:02 +08:00 |
|
Mick
|
35cefd1c51
|
feat: add safeguards for remote media URLs (#34892)
|
2026-08-15 18:12:15 +08:00 |
|
Colin Z
|
bc7e3ba66c
|
[AMD][Quantization] Online MXFP4 quantization 4/N - NVFP4 to MXFP4 Online Requantization on AMD GPUs (#29328)
|
2026-08-14 21:59:39 -07:00 |
|
  
|
bfb224ff01
|
Add Reasoning-Aware Compression (RAC) pruning recipe for reasoning models (#32414)
Co-authored-by: Ryan Lucas <ryanluc@mit.edu>
Co-authored-by: Kayhan Behdin <kbehdin@linkedin.com>
Co-authored-by: Zhipeng Wang <zwanga@wustl.edu>
|
2026-08-14 15:13:45 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
5e65dd01a7
|
Remove the torchao integration (--torchao-config) (#34304)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-08-14 21:49:11 +08:00 |
|
amote-i
|
fe0c18effd
|
[NPU] [DOC] Add Qwen3.8-Max deployment tutorial on Ascend NPUs (#34836)
|
2026-08-14 20:02:04 +08:00 |
|
 triple-muandMick
|
a86edcdc0a
|
[diffusion] feat: rebuild minimax-h3 adaln outputs on demand (#34650)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-14 15:33:51 +08:00 |
|
Mick
|
46d84f4b48
|
feat(cli): add extensible serve backend plugins (#34753)
|
2026-08-14 13:57:59 +08:00 |
|
Ziang Li
|
9d34c2809f
|
[FlashInfer v0.6.16] Support FlashInfer CuTe DSL NVFP4 MoE quantization (#28354)
|
2026-08-13 17:33:46 -07:00 |
|
 Khoa PhamandCursor
|
652a2709d1
|
[Docs] Add decode context parallelism to advanced features (#34654)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-13 13:51:57 -07:00 |
|
  
|
fad376d3ee
|
[CPU][QUANT] add amx cpu support for auto-round (#29593)
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Signed-off-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: Weiwei Zhang <WeiweiZhang1@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-08-13 15:50:59 +08:00 |
|
 Carrie ChenandBrayden Zhong
|
6a5a9eccaa
|
add flashinfer cute-dsl backend for mxfp8 gemm (#34042)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-08-13 08:50:01 +08:00 |
|
Lukas Humbel
|
40eaf34428
|
fix: make automatic NUMA binding configurable (#30394)
|
2026-08-12 17:40:28 -07:00 |
|
Mick
|
644d55ebfa
|
[diffusion] feat: support native and peft minimax h3 loras (#34359)
|
2026-08-12 17:52:40 +08:00 |
|
Yuzhen Zhou
|
2d76d537e5
|
feat: support deterministic FA4 for GLM-4.7-Flash (#33945)
|
2026-08-12 16:57:34 +08:00 |
|
Mick
|
a9a355774a
|
[diffusion] feat: support dynamically cpu offload components (#34391)
|
2026-08-12 11:38:27 +08:00 |
|
Xiaoyu Zhang
|
a53d3636ce
|
[diffusion][model] Add native SANA-Video T2V support (#32921)
|
2026-08-12 10:07:24 +08:00 |
|
gongwei1027
|
2c07ca5e8d
|
[Fix] Allow flashinfer_sparse_mla DSA backend for HiSparse on SM120 FP8 KV (#33075)
|
2026-08-11 15:05:04 -07:00 |
|
Mick
|
8267d76c2c
|
[VLM] replace deprecated image processor use_fast (#34175)
|
2026-08-12 00:14:07 +08:00 |
|
Jinyan Yi
|
f148eb6e6e
|
Add Hunyuan3 On Ascend Doc (#30223)
|
2026-08-11 21:20:30 +08:00 |
|
 Zhiqiang XieandTingwei Huang
|
5469faec45
|
HiSparse: shared-index (IndexShare) plan-then-IO swap-in prefetch (#34329)
Co-authored-by: Tingwei Huang <huangtingwei9988@gmail.com>
|
2026-08-11 01:58:28 -07:00 |
|
 
|
d07ac32d05
|
[diffusion] feat: support --served-model-name in sglang serve (#34228)
Co-authored-by: TobyMint <tobymint@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-10 22:27:09 +08:00 |
|
Mick
|
8ba9385097
|
[diffusion] chore: optimize model weight loading (#34064)
|
2026-08-10 20:07:48 +08:00 |
|
  
|
a6c34df044
|
Muse Glimmer Cookbook (#34271)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: Jimmy Shong <jimmysh341@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-08-10 10:21:13 +00:00 |
|
Feng Su
|
fb3d1419fd
|
[tracing] sglang tracing v2: support exporting tracing data asynchronously (#30023)
|
2026-08-10 15:23:05 +08:00 |
|
    
|
2969ab3d41
|
[MLX] Window-bounded SWA KV storage and in-graph sampling (#34166)
Co-authored-by: Siming Deng <siming_deng_stat@163.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
Co-authored-by: Jiminator <Jiminator@users.noreply.github.com>
Co-authored-by: damahua <damahua@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-09 21:28:51 -07:00 |
|
Mick
|
169783d42f
|
[diffusion] chore: make torch.compile opt-in for speed mode (#34173)
|
2026-08-10 10:22:16 +08:00 |
|
Liangsheng Yin
|
7c90840bad
|
[CI] Key scheduled CUDA suites by runner_config instead of hand-written jobs (#34186)
|
2026-08-09 16:44:53 -07:00 |
|
WenhaoZhang
|
51470b376f
|
[diffusion] feat: support sol-attn sparse attention backend for h3 (#33702)
|
2026-08-09 16:26:06 +08:00 |
|
 
|
548ff545c5
|
[diffusion] fix: guard sage attention sm90 bindings (#34107)
Co-authored-by: RunFMe <RunFMe@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-08 21:32:40 +08:00 |
|
 DarkSharpnessandClaude Fable 5
|
4ad5bb5d9a
|
[jit_kernel] Move JIT kernels into namespace sglang (#33400)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 16:10:15 +08:00 |
|
 
|
f64328c7f6
|
[diffusion] feat: support quant-videogen prq kv-cache quantization (memory-saving) for causal-dit (#32581)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-08 12:56:54 +08:00 |
|
 MickandClaude Fable 5
|
24c84dfa68
|
[diffusion] doc: add parallelism overview (#33704)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 11:36:06 +08:00 |
|
 MickandClaude Fable 5
|
a25c330eb1
|
[diffusion] feat: cross-node sequence parallelism (Ulysses x Ring) (#33327)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 10:56:28 +08:00 |
|
 
|
bc148dfdc8
|
[diffusion] feat: make scheduler rpc deadlines explicit (#33965)
Co-authored-by: suoyf <suoyf@nscc-tj.cn>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-07 21:33:23 +08:00 |
|
 Lennox FuandMick
|
7af3d000f2
|
[diffusion] feat: gate /health and /health_generate on warmup completion and add liveness endpoint (#33787)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-07 17:59:23 +08:00 |
|
amote-i
|
470807ef74
|
[NPU] [DOC] Upgrade recommendeded sglang version on Ascend NPU (#33976)
|
2026-08-07 17:45:00 +08:00 |
|
Mick
|
c2657cc4bf
|
[diffusion] refactor: gate fast vae paths by quality (#33849)
|
2026-08-07 12:39:55 +08:00 |
|
WenhaoZhang
|
914644e81c
|
[diffusion] fix: fix 4/8-step distilled minimax-h3 turbo lora merge (#33875)
|
2026-08-07 12:38:37 +08:00 |
|
 MickandClaude Fable 5
|
1e08b865f9
|
[diffusion] feat: support K/V-gather style sequence parallel (CP-like) attention (#32667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 09:39:28 +08:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
434e646282
|
[Deps] Upgrade CUDA PyTorch stack to 2.13 (#28836)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-08-06 12:08:44 -07:00 |
|
 Ziang LiandBrayden Zhong
|
4ad990ba7d
|
[ModelOpt FP4] Support online MoE weight quantization (#33115)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-08-06 11:01:55 -07:00 |
|
Mick
|
7195b8e4c7
|
[diffusion] refactor: validate and document spectrum controls (#33851)
|
2026-08-06 23:23:11 +08:00 |
|
Mick
|
44bde3911a
|
[diffusion] fix: resolve IPC A2A peers from process groups (#33848)
|
2026-08-06 23:22:41 +08:00 |
|
Mick
|
c212a6938c
|
[diffusion] chore: retire released warmup and decoder flags (#33850)
|
2026-08-06 23:02:32 +08:00 |
|
 
|
f8f2870a84
|
Profiling Enhancements [1/3]: cuda graph profile traces (#24370)
Co-authored-by: Basit <mohbasit@ctr2-alola-ctrl-01.amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-08-06 03:19:35 -07:00 |
|
 MickandClaude Fable 5
|
604d3561b0
|
[diffusion] feat: data-parallel serving (--dp-size) (#33725)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 11:50:13 +08:00 |
|
 Khoa PhamandClaude Opus 5
|
beabc5949b
|
Enable MoE deferred finalize by default and drop its expert_weights dtype workaround (#33618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-05 17:56:47 -07:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4c0a8940fa
|
[Kernel] Unify BaseFusedOp and MultiPlatformOp dispatch (#33205)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 08:52:09 +08:00 |
|