Commit Graph
100 Commits
Author SHA1 Message Date
Mick d97b796c16 [diffusion] chore: reuse SRT CLIP encoder blocks (#35004) 2026-08-17 19:51:51 +08:00
Mick e9ad8102a2 [diffusion] chore: reuse SRT SigLIP in Pi0.5 (#34992) 2026-08-17 19:33:36 +08:00
Mick 0e178c3d22 [diffusion] chore: reuse srt siglip vision model (#34988) 2026-08-17 09:16:17 +08:00
Mick d3589a7251 [diffusion] CI: tighten NVIDIA perf baselines (#35016) 2026-08-16 20:54:50 +08:00
Mick 3d3194f6c3 vlm: cache kimi-k3 per-image processor artifacts (#34404) 2026-08-16 19:51:13 +08:00
Mick 968b355f12 vlm: streamline vision sdpa reshapes (#34991) 2026-08-16 19:05:53 +08:00
Mick 2ee0d38a85 [diffusion] chore: refresh docs, retire stale knobs, and fix nightly attribution (#34663) 2026-08-16 15:41:08 +08:00
Mick a54de989c8 [diffusion] chore: speed up minimax-h3 vae decode on 2×h100 (#34817) 2026-08-16 15:38:21 +08:00
Mick e9fe58139f [diffusion] refactor: unify component residency controls (#34736) 2026-08-16 11:24:48 +08:00
Mick d269a28b47 [diffusion] refactor: route minimax h3 vae attention through native backends (#34949) 2026-08-16 10:07:26 +08:00
Mick 4f9da62547 [diffusion] chore: use native hunyuan3d paint and delight models (#34980) 2026-08-16 10:03:48 +08:00
Mick d106e8b23a [diffusion] chore: use native ernie prompt enhancer (#34951) 2026-08-16 09:59:51 +08:00
Mick 19e3bd6391 [diffusion] chore: use native qwen3-vl vision encoder (#34945) 2026-08-16 09:58:46 +08:00
Mick eb6b773149 [diffusion] chore: use native qwen2.5-vl generation (#34896) 2026-08-16 09:57:39 +08:00
Mick 4beb157e87 [diffusion] doc: define native diffusion model integration contract (#34952) 2026-08-15 23:37:34 +08:00
Mick e331baaaa8 [diffusion] chore: scope attention backend fallback (#34891) 2026-08-15 22:00:46 +08:00
Mick a64791b312 test: restore GLM-4.1V nightly latency threshold (#34811) 2026-08-15 18:13:19 +08:00
Mick 35cefd1c51 feat: add safeguards for remote media URLs (#34892) 2026-08-15 18:12:15 +08:00
Mick 46d84f4b48 feat(cli): add extensible serve backend plugins (#34753) 2026-08-14 13:57:59 +08:00
MickandClaude Opus 5 969921b32d [diffusion] chore: track minimax-h3 in the nightly diffusion benchmark (#34655)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 15:12:26 +08:00
Mick e8c7dddfa0 [VLM] add content-addressed preprocessing cache infrastructure (#34398) 2026-08-13 14:09:07 +08:00
Mick 69bf601e3c fix: restore VLM nightly regression coverage (#34662) 2026-08-12 21:47:16 -07:00
MickandClaude Opus 5 a318e16956 [diffusion] feat: publish an index of nightly comparison runs (#34652)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 11:50:10 +08:00
Mick ad47dde65c [diffusion] optimize: fuse cosmos qk norm, rope, and kv packing (#34275) 2026-08-12 23:18:16 +08:00
Mick dc5f6c4883 [diffusion] optimize: stream and parallelize bit-exact video output saves (#34564) 2026-08-12 20:03:37 +08:00
Mick 9701cc138c [diffusion] optimize: optimize bit-exact h3 reference video ingress (#34563) 2026-08-12 20:02:49 +08:00
Mick b3bffef70a [diffusion] UX: suppress noisy worker startup warnings (#34512) 2026-08-12 17:53:45 +08:00
Mick 644d55ebfa [diffusion] feat: support native and peft minimax h3 loras (#34359) 2026-08-12 17:52:40 +08:00
Mick a9a355774a [diffusion] feat: support dynamically cpu offload components (#34391) 2026-08-12 11:38:27 +08:00
Mick 2be9773a21 [diffusion] doc: update cosmos3 edge and distilled cookbook (#34497) 2026-08-12 10:51:11 +08:00
MickandClaude Opus 5 81c88da1ab [diffusion] fix: nightly diffusion benchmark passes the retired --warmup flag (#34423)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:25:14 +08:00
Mick a2d723820e [diffusion] fix: fix model-driven dit layerwise offload auto policy (#34401) 2026-08-12 09:08:31 +08:00
Mick 8267d76c2c [VLM] replace deprecated image processor use_fast (#34175) 2026-08-12 00:14:07 +08:00
Mick dd8c5849af [diffusion] refactor: move dit execution capabilities to runtime models (#34249) 2026-08-11 18:20:43 +08:00
Mick aeab1de1de [diffusion] optimization: support cuda graph for Pi-0.5 prefix encoding (#34256) 2026-08-11 09:14:53 +08:00
Mick ba5183fe10 [diffusion] UX: fix CI server warmup progress logging (#34301) 2026-08-11 09:09:32 +08:00
Mick 418975ba64 [EPD] feat: pipeline owner-only multimodal preprocessing (#34206) 2026-08-11 09:08:05 +08:00
Mick ec9babe36c [diffusion] fix: fix h3 rank-local fsdp qkv loading (#34294) 2026-08-10 23:59:00 +08:00
Mick 3e2a26708b [diffusion] chore: expose architecture config at the dit runtime boundary (#34248) 2026-08-10 20:08:45 +08:00
Mick 8ba9385097 [diffusion] chore: optimize model weight loading (#34064) 2026-08-10 20:07:48 +08:00
Mick 443b62db57 fix(vlm): stream-order cuda-ipc feature pool lifecycle and streamline multimodal transport module (#33949) 2026-08-10 18:47:19 +08:00
Mick 955569a2dc [diffusion] feat: expose cosmos3 policies through the Action API (#34243) 2026-08-10 18:16:20 +08:00
Mick 169783d42f [diffusion] chore: make torch.compile opt-in for speed mode (#34173) 2026-08-10 10:22:16 +08:00
Mick c20e99bd22 fix(vlm): preserve Kimi-K3 GPU JPEG accuracy (#34163) 2026-08-10 09:42:52 +08:00
Mick 22e003580b [Kimi K3] optimize: preprocess cpu-transport images on the vision owner (#33921) 2026-08-09 16:13:20 +08:00
Mick db75dfe10f fix: always capture default prefill CUDA graph (#33352) 2026-08-08 19:24:49 +08:00
Mick cf2d4fd679 docs: clarify K3 VLM feature transport (#34099) 2026-08-08 19:23:00 +08:00
Mick d747bd052e feat(vlm): auto-select cuda vmm on multi-node mnnvl (#33936) 2026-08-08 16:00:58 +08:00
Mick db3898fec1 fix: avoid piecewise prefill graph for trtllm_mla (#32785) 2026-08-08 16:00:10 +08:00
MickandClaude Fable 5 24c84dfa68 [diffusion] doc: add parallelism overview (#33704)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 11:36:06 +08:00
MickandClaude Fable 5 a25c330eb1 [diffusion] feat: cross-node sequence parallelism (Ulysses x Ring) (#33327)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 10:56:28 +08:00
MickandClaude Fable 5 52afe87a08 [diffusion] fix: stop runai-model-streamer's rank-discovery collective from firing on independent per-rank loads (#33969)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 08:54:08 +08:00
Mick 5ca734fc3d [diffusion] UX: speed up tp and fsdp checkpoint loading (#33960) 2026-08-07 19:23:30 +08:00
MickandClaude Fable 5 13938fed3f [diffusion] feat: make ring admission a backend capability (#33928)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 17:55:44 +08:00
MickandClaude Fable 5 28b43bf693 [diffusion] perf: build qwen's masked varlen metadata host-side (#33954)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 17:55:15 +08:00
MickandClaude Fable 5 85d611a055 [diffusion] fix: scope the masked-path replicated guard to sp runs (#33953)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 15:45:50 +08:00
Mick 9aadacfc53 [diffusion] CI: exercise the default sp selection in CI (#33931) 2026-08-07 14:20:19 +08:00
Mick c2657cc4bf [diffusion] refactor: gate fast vae paths by quality (#33849) 2026-08-07 12:39:55 +08:00
MickandClaude Fable 5 6dc77e490d [diffusion] chore: route zimage and hunyuanvideo attention through USPAttention (#33923)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 12:36:23 +08:00
MickandClaude Fable 5 698f019a5a [diffusion] chore: derive h3 attention admission from backend capabilities (#33707)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 12:35:15 +08:00
MickandClaude Fable 5 1e08b865f9 [diffusion] feat: support K/V-gather style sequence parallel (CP-like) attention (#32667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 09:39:28 +08:00
Mick 9ee658d4f6 [diffusion] CI: fix output-rank test fixture (#33878) 2026-08-07 09:33:50 +08:00
Mick 7195b8e4c7 [diffusion] refactor: validate and document spectrum controls (#33851) 2026-08-06 23:23:11 +08:00
Mick 44bde3911a [diffusion] fix: resolve IPC A2A peers from process groups (#33848) 2026-08-06 23:22:41 +08:00
Mick c212a6938c [diffusion] chore: retire released warmup and decoder flags (#33850) 2026-08-06 23:02:32 +08:00
Mick 183bd80add [diffusion] chore: centralize entrypoint API hygiene (#33845) 2026-08-06 21:53:40 +08:00
Mick 2132cdef16 [diffusion] chore: consolidate pipeline core hygiene (#33843) 2026-08-06 21:51:34 +08:00
Mick 45dfd80674 [diffusion] refactor: simplify disaggregation transport hygiene (#33844) 2026-08-06 21:50:32 +08:00
MickandClaude Fable 5 bfce378e5f [diffusion] feat: capture-safe pynccl all-to-all (#33775)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 15:41:05 +08:00
MickandClaude Fable 5 604d3561b0 [diffusion] feat: data-parallel serving (--dp-size) (#33725)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 11:50:13 +08:00
Mick 99709f734d [VLM] split multimodal scheduling from mm_utils (#32415) 2026-08-05 20:24:12 +08:00
Mick b058dc9106 [diffusion] fix: reject ring parallelism where it would silently miscompute (#33353) 2026-08-04 11:15:23 +08:00
Mick 0ba46c88e5 [diffusion] CI: add minimax-h3 2-gpu consistency coverage (#33281) 2026-08-03 22:40:08 +08:00
70fe2e0dd5 [diffusion] model: support minimax-h3 (#33275)
Co-authored-by: zhenaozhenfu <zhenaozhenfu@minimaxi.com>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Zijie Xia <zijie_xia@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: chao-xue <877184285@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-02 22:32:37 +08:00
MickandClaude Fable 5 754b692afc [diffusion] optimization: support cuda-ipc zero-staging all-to-all for 2-rank Ulysses (#31854)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 19:35:48 +08:00
Mick a149717308 feat: log multimodal encoder DP tradeoffs (#30903) 2026-07-31 08:50:20 +08:00
Mick b129e8a299 [diffusion] docs: surface diffusion AR and PE guides (#32932) 2026-07-30 21:05:54 +08:00
Mick db3da62333 [diffusion] feat: unify encoder folding and batch data-parallel encoding (#30211) 2026-07-30 20:15:22 +08:00
Mick 22faf9fef8 embedding: centralize capabilities and complete OpenAI compatibility (#32481) 2026-07-30 10:28:52 +08:00
Mick 2aa86e9130 [diffusion] docs: add diffusion cookbook model tags (#32836) 2026-07-30 10:03:55 +08:00
Mick 22151edca1 [diffusion] optimization: accelerate CUDA video output finalization (#32784) 2026-07-29 22:04:20 +08:00
MickandClaude Sonnet 5 67c2258906 [diffusion] fix: fix dual-DiT models crash with (1,)-placeholder weights after compile-time offload (#32743)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 20:44:13 +08:00
Mick da5528db30 fix(vlm): materialize Qwen3-VL features on the vision device (#31596) 2026-07-29 15:25:04 +08:00
Mick 9a03bebf13 docs(kimi-k3): clarify VLM compatibility (#32661) 2026-07-29 11:53:04 +08:00
Mick 70ea37e7e0 vlm: reject moss vision metadata mismatches (#31957) 2026-07-29 07:09:54 +08:00
Mick 84cdfde5b2 [diffusion] fix: fix diffusion output stability on mps (#30017) 2026-07-28 21:55:03 +08:00
Mick 161fffedfc docs: clarify diffusion stage reuse guidance (#32639) 2026-07-28 19:52:52 +08:00
Mick 75017c3fa0 [diffusion] fix: keep fused qk-norm-rope out of dynamo tracing (#31849) 2026-07-28 19:51:00 +08:00
Mick a24906a091 [diffusion] feat: add dynamic cuDNN SDPA attention backend (#30090) 2026-07-28 19:50:24 +08:00
Mick 08af5aea57 optimize: optimize EmbeddingGemma prefill performance (#32383) 2026-07-27 17:34:29 +08:00
Mick abb8f4b5e3 model: support EmbeddingGemma (#32375) 2026-07-27 10:40:47 +08:00
Mick 1054060ef1 perf: speed up marlin moe with occupancy-aware launch specialization (#31552) 2026-07-25 19:38:11 +08:00
Mick 4682ded472 vlm: parallelize multimodal preprocessing with customized worker num (#31438) 2026-07-21 08:44:58 +08:00
Mick 2eed35d738 perf: avoid temporary VLM encoder gather padding (#31301) 2026-07-20 12:54:29 +08:00
Mick 6a25dd7b5f fix: warm up Kimi VLM vision encoder at startup (#31298) 2026-07-20 08:50:57 +08:00
Mick d4801be447 fix: fix vlm cuda graph shape stability (#30868) 2026-07-19 22:35:51 +08:00
Mick 573c075fef CI: synchronize prefill graph test fixtures (#31665) 2026-07-18 18:48:05 +08:00
Mick 6c6175fabd perf: avoid excessive prefill CUDA graph padding (#31487) 2026-07-18 16:25:30 +08:00
Mick 42a058c760 optimize: avoid fla l2-norm recompilation by token count (#31558) 2026-07-18 07:58:36 +08:00
Mick 85ac56c823 docs: simplify diffusion new model guide (#30109) 2026-07-17 19:39:38 +08:00