Mick
|
dd8c5849af
|
[diffusion] refactor: move dit execution capabilities to runtime models (#34249)
|
2026-08-11 18:20:43 +08:00 |
|
Mick
|
aeab1de1de
|
[diffusion] optimization: support cuda graph for Pi-0.5 prefix encoding (#34256)
|
2026-08-11 09:14:53 +08:00 |
|
Mick
|
ba5183fe10
|
[diffusion] UX: fix CI server warmup progress logging (#34301)
|
2026-08-11 09:09:32 +08:00 |
|
Mick
|
418975ba64
|
[EPD] feat: pipeline owner-only multimodal preprocessing (#34206)
|
2026-08-11 09:08:05 +08:00 |
|
Mick
|
ec9babe36c
|
[diffusion] fix: fix h3 rank-local fsdp qkv loading (#34294)
|
2026-08-10 23:59:00 +08:00 |
|
Mick
|
3e2a26708b
|
[diffusion] chore: expose architecture config at the dit runtime boundary (#34248)
|
2026-08-10 20:08:45 +08:00 |
|
Mick
|
8ba9385097
|
[diffusion] chore: optimize model weight loading (#34064)
|
2026-08-10 20:07:48 +08:00 |
|
Mick
|
443b62db57
|
fix(vlm): stream-order cuda-ipc feature pool lifecycle and streamline multimodal transport module (#33949)
|
2026-08-10 18:47:19 +08:00 |
|
Mick
|
955569a2dc
|
[diffusion] feat: expose cosmos3 policies through the Action API (#34243)
|
2026-08-10 18:16:20 +08:00 |
|
Mick
|
169783d42f
|
[diffusion] chore: make torch.compile opt-in for speed mode (#34173)
|
2026-08-10 10:22:16 +08:00 |
|
Mick
|
c20e99bd22
|
fix(vlm): preserve Kimi-K3 GPU JPEG accuracy (#34163)
|
2026-08-10 09:42:52 +08:00 |
|
Mick
|
22e003580b
|
[Kimi K3] optimize: preprocess cpu-transport images on the vision owner (#33921)
|
2026-08-09 16:13:20 +08:00 |
|
Mick
|
db75dfe10f
|
fix: always capture default prefill CUDA graph (#33352)
|
2026-08-08 19:24:49 +08:00 |
|
Mick
|
cf2d4fd679
|
docs: clarify K3 VLM feature transport (#34099)
|
2026-08-08 19:23:00 +08:00 |
|
Mick
|
d747bd052e
|
feat(vlm): auto-select cuda vmm on multi-node mnnvl (#33936)
|
2026-08-08 16:00:58 +08:00 |
|
Mick
|
db3898fec1
|
fix: avoid piecewise prefill graph for trtllm_mla (#32785)
|
2026-08-08 16:00:10 +08:00 |
|
 MickandClaude Fable 5
|
24c84dfa68
|
[diffusion] doc: add parallelism overview (#33704)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 11:36:06 +08:00 |
|
 MickandClaude Fable 5
|
a25c330eb1
|
[diffusion] feat: cross-node sequence parallelism (Ulysses x Ring) (#33327)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 10:56:28 +08:00 |
|
 MickandClaude Fable 5
|
52afe87a08
|
[diffusion] fix: stop runai-model-streamer's rank-discovery collective from firing on independent per-rank loads (#33969)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 08:54:08 +08:00 |
|
Mick
|
5ca734fc3d
|
[diffusion] UX: speed up tp and fsdp checkpoint loading (#33960)
|
2026-08-07 19:23:30 +08:00 |
|
 MickandClaude Fable 5
|
13938fed3f
|
[diffusion] feat: make ring admission a backend capability (#33928)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 17:55:44 +08:00 |
|
 MickandClaude Fable 5
|
28b43bf693
|
[diffusion] perf: build qwen's masked varlen metadata host-side (#33954)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 17:55:15 +08:00 |
|
 MickandClaude Fable 5
|
85d611a055
|
[diffusion] fix: scope the masked-path replicated guard to sp runs (#33953)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 15:45:50 +08:00 |
|
Mick
|
9aadacfc53
|
[diffusion] CI: exercise the default sp selection in CI (#33931)
|
2026-08-07 14:20:19 +08:00 |
|
Mick
|
c2657cc4bf
|
[diffusion] refactor: gate fast vae paths by quality (#33849)
|
2026-08-07 12:39:55 +08:00 |
|
 MickandClaude Fable 5
|
6dc77e490d
|
[diffusion] chore: route zimage and hunyuanvideo attention through USPAttention (#33923)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 12:36:23 +08:00 |
|
 MickandClaude Fable 5
|
698f019a5a
|
[diffusion] chore: derive h3 attention admission from backend capabilities (#33707)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 12:35:15 +08:00 |
|
 MickandClaude Fable 5
|
1e08b865f9
|
[diffusion] feat: support K/V-gather style sequence parallel (CP-like) attention (#32667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 09:39:28 +08:00 |
|
Mick
|
9ee658d4f6
|
[diffusion] CI: fix output-rank test fixture (#33878)
|
2026-08-07 09:33:50 +08:00 |
|
Mick
|
7195b8e4c7
|
[diffusion] refactor: validate and document spectrum controls (#33851)
|
2026-08-06 23:23:11 +08:00 |
|
Mick
|
44bde3911a
|
[diffusion] fix: resolve IPC A2A peers from process groups (#33848)
|
2026-08-06 23:22:41 +08:00 |
|
Mick
|
c212a6938c
|
[diffusion] chore: retire released warmup and decoder flags (#33850)
|
2026-08-06 23:02:32 +08:00 |
|
Mick
|
183bd80add
|
[diffusion] chore: centralize entrypoint API hygiene (#33845)
|
2026-08-06 21:53:40 +08:00 |
|
Mick
|
2132cdef16
|
[diffusion] chore: consolidate pipeline core hygiene (#33843)
|
2026-08-06 21:51:34 +08:00 |
|
Mick
|
45dfd80674
|
[diffusion] refactor: simplify disaggregation transport hygiene (#33844)
|
2026-08-06 21:50:32 +08:00 |
|
 MickandClaude Fable 5
|
bfce378e5f
|
[diffusion] feat: capture-safe pynccl all-to-all (#33775)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 15:41:05 +08:00 |
|
 MickandClaude Fable 5
|
604d3561b0
|
[diffusion] feat: data-parallel serving (--dp-size) (#33725)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 11:50:13 +08:00 |
|
Mick
|
99709f734d
|
[VLM] split multimodal scheduling from mm_utils (#32415)
|
2026-08-05 20:24:12 +08:00 |
|
Mick
|
b058dc9106
|
[diffusion] fix: reject ring parallelism where it would silently miscompute (#33353)
|
2026-08-04 11:15:23 +08:00 |
|
Mick
|
0ba46c88e5
|
[diffusion] CI: add minimax-h3 2-gpu consistency coverage (#33281)
|
2026-08-03 22:40:08 +08:00 |
|
      
|
70fe2e0dd5
|
[diffusion] model: support minimax-h3 (#33275)
Co-authored-by: zhenaozhenfu <zhenaozhenfu@minimaxi.com>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Zijie Xia <zijie_xia@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: chao-xue <877184285@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-02 22:32:37 +08:00 |
|
 MickandClaude Fable 5
|
754b692afc
|
[diffusion] optimization: support cuda-ipc zero-staging all-to-all for 2-rank Ulysses (#31854)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-31 19:35:48 +08:00 |
|
Mick
|
a149717308
|
feat: log multimodal encoder DP tradeoffs (#30903)
|
2026-07-31 08:50:20 +08:00 |
|
Mick
|
b129e8a299
|
[diffusion] docs: surface diffusion AR and PE guides (#32932)
|
2026-07-30 21:05:54 +08:00 |
|
Mick
|
db3da62333
|
[diffusion] feat: unify encoder folding and batch data-parallel encoding (#30211)
|
2026-07-30 20:15:22 +08:00 |
|
Mick
|
22faf9fef8
|
embedding: centralize capabilities and complete OpenAI compatibility (#32481)
|
2026-07-30 10:28:52 +08:00 |
|
Mick
|
2aa86e9130
|
[diffusion] docs: add diffusion cookbook model tags (#32836)
|
2026-07-30 10:03:55 +08:00 |
|
Mick
|
22151edca1
|
[diffusion] optimization: accelerate CUDA video output finalization (#32784)
|
2026-07-29 22:04:20 +08:00 |
|
 MickandClaude Sonnet 5
|
67c2258906
|
[diffusion] fix: fix dual-DiT models crash with (1,)-placeholder weights after compile-time offload (#32743)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
2026-07-29 20:44:13 +08:00 |
|
Mick
|
da5528db30
|
fix(vlm): materialize Qwen3-VL features on the vision device (#31596)
|
2026-07-29 15:25:04 +08:00 |
|
Mick
|
9a03bebf13
|
docs(kimi-k3): clarify VLM compatibility (#32661)
|
2026-07-29 11:53:04 +08:00 |
|
Mick
|
70ea37e7e0
|
vlm: reject moss vision metadata mismatches (#31957)
|
2026-07-29 07:09:54 +08:00 |
|
Mick
|
84cdfde5b2
|
[diffusion] fix: fix diffusion output stability on mps (#30017)
|
2026-07-28 21:55:03 +08:00 |
|
Mick
|
161fffedfc
|
docs: clarify diffusion stage reuse guidance (#32639)
|
2026-07-28 19:52:52 +08:00 |
|
Mick
|
75017c3fa0
|
[diffusion] fix: keep fused qk-norm-rope out of dynamo tracing (#31849)
|
2026-07-28 19:51:00 +08:00 |
|
Mick
|
a24906a091
|
[diffusion] feat: add dynamic cuDNN SDPA attention backend (#30090)
|
2026-07-28 19:50:24 +08:00 |
|
Mick
|
08af5aea57
|
optimize: optimize EmbeddingGemma prefill performance (#32383)
|
2026-07-27 17:34:29 +08:00 |
|
Mick
|
abb8f4b5e3
|
model: support EmbeddingGemma (#32375)
|
2026-07-27 10:40:47 +08:00 |
|
Mick
|
1054060ef1
|
perf: speed up marlin moe with occupancy-aware launch specialization (#31552)
|
2026-07-25 19:38:11 +08:00 |
|
Mick
|
4682ded472
|
vlm: parallelize multimodal preprocessing with customized worker num (#31438)
|
2026-07-21 08:44:58 +08:00 |
|
Mick
|
2eed35d738
|
perf: avoid temporary VLM encoder gather padding (#31301)
|
2026-07-20 12:54:29 +08:00 |
|
Mick
|
6a25dd7b5f
|
fix: warm up Kimi VLM vision encoder at startup (#31298)
|
2026-07-20 08:50:57 +08:00 |
|
Mick
|
d4801be447
|
fix: fix vlm cuda graph shape stability (#30868)
|
2026-07-19 22:35:51 +08:00 |
|
Mick
|
573c075fef
|
CI: synchronize prefill graph test fixtures (#31665)
|
2026-07-18 18:48:05 +08:00 |
|
Mick
|
6c6175fabd
|
perf: avoid excessive prefill CUDA graph padding (#31487)
|
2026-07-18 16:25:30 +08:00 |
|
Mick
|
42a058c760
|
optimize: avoid fla l2-norm recompilation by token count (#31558)
|
2026-07-18 07:58:36 +08:00 |
|
Mick
|
85ac56c823
|
docs: simplify diffusion new model guide (#30109)
|
2026-07-17 19:39:38 +08:00 |
|
Mick
|
681c223570
|
refactor: wrap split backends once on full-attention backends (#31439)
|
2026-07-17 19:15:04 +08:00 |
|
Mick
|
24a8944e15
|
fix: enable Kimi multimodal breakable prefill cuda graph replay (#31391)
|
2026-07-17 19:13:54 +08:00 |
|
Mick
|
7d0fd5101d
|
optimization: shard kimi dp image feature transport and misc optimizations (#31227)
|
2026-07-16 20:31:42 +08:00 |
|
Mick
|
d9003dd452
|
fix: skip unsafe automatic prefill graph capture (#31204)
|
2026-07-16 09:27:38 +08:00 |
|
Mick
|
947a14d617
|
feat: unify multimodal feature transport (#30904)
|
2026-07-15 17:42:38 +08:00 |
|
Mick
|
43124cdd90
|
fix: fix image benchmark backend parity (#30867)
|
2026-07-15 10:11:22 +08:00 |
|
Mick
|
04af94d150
|
fix: avoid tilelang cuda runtime pollution (#30870)
|
2026-07-14 22:30:27 +08:00 |
|
Mick
|
43241b7f3f
|
[diffusion] model: support fal Ideogram V4 Fast and Instant (#31177)
|
2026-07-14 19:51:35 +08:00 |
|
Mick
|
23b2c6f1ce
|
docs: fix diffusion cookbook overview cards (#31101)
|
2026-07-14 10:40:42 +08:00 |
|
Mick
|
33f83011e0
|
fix: fix Kimi-VL encoder parallelism (#30869)
|
2026-07-14 08:44:06 +08:00 |
|
Mick
|
7da30f4e55
|
feat: enable piecewise prefill graph for Kimi K2.5/K2.7 (#30889)
|
2026-07-13 08:37:30 +08:00 |
|
Mick
|
f1c247edf9
|
profile: add vlm prefill profiler ranges (#30871)
|
2026-07-12 14:07:10 +08:00 |
|
Mick
|
bce3fc987d
|
perf: reuse MoonViT FA3 max-seqlen metadata (#30878)
|
2026-07-12 14:05:21 +08:00 |
|
Mick
|
a358abd651
|
chore: update vlm moe config and tune scripts (#30866)
|
2026-07-12 08:35:59 +08:00 |
|
Mick
|
af66370d81
|
bench: support random image resolutions (#30879)
|
2026-07-12 08:28:56 +08:00 |
|
Mick
|
649ce5dd3d
|
model: support Pi0.5 (#30633)
|
2026-07-11 07:50:58 +08:00 |
|
Mick
|
559854fe6a
|
[diffusion] docs: sync cookbook and log hygiene (#30791)
|
2026-07-10 22:55:14 +08:00 |
|
Mick
|
7090a49198
|
update codeowners (#30788)
|
2026-07-10 22:35:34 +08:00 |
|
Mick
|
4a8e1b07a2
|
[diffusion] refactor: reorganize runtime utility and server_args modules (#30447)
|
2026-07-10 15:53:00 +08:00 |
|
Mick
|
5ce5e1ee3e
|
[Diffusion] Revert CPU AMX optimizations (#30716)
|
2026-07-10 09:09:38 +08:00 |
|
Mick
|
6c1fb8a937
|
[diffusion] fix: fix ragged-caption dynamic-batching accuracy bug in ernie-Image (#30241)
|
2026-07-07 08:41:51 +08:00 |
|
Mick
|
5f98f62a8a
|
[diffusion] perf: tp-shard every text/image encoder across the full DiT replica (any parallelism) (#30086)
|
2026-07-06 14:48:07 +08:00 |
|
Mick
|
a37bc2456d
|
[diffusion] refactor: consolidate diffusion weight load planning (#30118)
|
2026-07-05 11:59:50 +08:00 |
|
Mick
|
763c6bf372
|
[diffusion] perf: add unified SP shard helpers and zero-copy tail-pad attention (#30107)
|
2026-07-04 23:57:39 +08:00 |
|
Mick
|
36fc0093d6
|
[diffusion] fix: shut down diffusion workers on serve exit (#30110)
|
2026-07-04 20:42:39 +08:00 |
|
Mick
|
03962d4238
|
[diffusion] feat: enable compile warmup for vae decode (#29306)
|
2026-07-04 15:25:26 +08:00 |
|
Mick
|
5af1f949ca
|
[diffusion] CI: prefer official diffusion consistency GT (#29831)
|
2026-07-04 10:09:33 +08:00 |
|
Mick
|
486bcb48e7
|
[diffusion] CI: fix AMD diffusion CI import (#30039)
|
2026-07-03 22:26:08 +08:00 |
|
Mick
|
42acfd1550
|
[diffusion] feat: performance_mode=speed enables torch.compile by default (#30016)
|
2026-07-03 19:12:17 +08:00 |
|
Mick
|
3c1adddff9
|
[diffusion] refactor: refactor cuda attention backend resolver (#29852)
|
2026-07-02 21:34:54 +08:00 |
|
Mick
|
119b76567d
|
[diffusion] feat: add --offload-during-compile to fit max-autotune on tight-memory GPUs (#29862)
|
2026-07-02 19:39:19 +08:00 |
|
Mick
|
03b9278da0
|
[diffusion] CI: tighten multimodal-gen consistency thresholds (#29824)
|
2026-07-02 01:32:51 +08:00 |
|
Mick
|
79f334b1aa
|
[diffusion] CI: add 5090 job (#29791)
|
2026-07-01 19:09:43 +08:00 |
|