Commit Graph
100 Commits
Author SHA1 Message Date
Mick f7fc2c8592 [diffusion] fix: fix accuracy for some image models (#20679) 2026-03-22 15:11:57 +08:00
Mick 6dfa8a40bc [diffusion] CI: make auxiliary coverage explicit and simplify testcases (#20983) 2026-03-21 20:18:23 +08:00
Mick f15b3338c9 Revert "[Bugfix] Fix GLM-4.6V vision regression in glm4v_moe and glm_ocr" (#20740) 2026-03-18 10:09:50 +08:00
Mickandgemini-code-assist[bot] 5717834f1f [diffusion] refactor: cleanup parallel_state.py (#20760)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-17 21:21:42 +08:00
Mick 5ec49a5309 Revert "[diffusion] CI: use dedicated HF token for accessing restricted models" (#20737) 2026-03-16 20:22:20 -07:00
Mick 474a851ae3 [diffusion] fix: fix sampling params incorrectly override in cli (#20689) 2026-03-17 08:48:10 +08:00
Mick 1eea744855 [diffusion] CI: enable UT (#20690) 2026-03-17 07:44:04 +08:00
Mick 485597e651 [diffusion] fix: fix some sampling args passed via cli are omitted (#20630) 2026-03-16 16:55:30 +08:00
Mick f07529b947 [diffusion] CI: use dedicated HF token for accessing restricted models (#20620) 2026-03-15 19:15:28 +08:00
Mick b638b25b22 [diffusion] UX: suppress excessive logging from httpx and httpcore (#20452) 2026-03-13 14:43:09 +08:00
Mick 154af9e46c update CI_PERMISSIONS.json (#20405) 2026-03-12 19:50:35 -07:00
Mick 8c8a487468 [diffusion] doc: add diffusion-optimal-perf (#20311) 2026-03-11 12:20:09 +08:00
Mick e1f0b3181a [diffusion] fix: adjust convert_hf_to_fp8 to be compatible with more dits (#20281) 2026-03-11 01:21:54 +08:00
Mick 2c183350be [diffusion] fix: fix wrong dit config for qwen-image-edit-plus-2511 (#20123) 2026-03-08 20:08:36 +08:00
Mick d605d811fb update CODEOWNERS (#19969) 2026-03-05 10:19:05 -08:00
Mick 43a9249e7c CI: update CI_PERMISSIONS.json (#19744) 2026-03-02 22:34:41 -08:00
Mick 2e15c015c0 [diffusion] feat: Add --model-id for config resolution; deprecate model_detectors (#19607) 2026-03-02 16:39:53 +08:00
Mick a75840b373 [diffusion] CI: create and refactor UT (#19619) 2026-03-01 19:38:20 +08:00
Mick d098c8dab0 [diffusion] add .claude and update contributing with attitude towards vibe-pr (#19511) 2026-03-01 14:41:55 +08:00
Mick 471acd98b9 [diffusion] logging: improve logging (#19312) 2026-02-25 23:00:35 +08:00
Mick 9840cd3f68 [diffusion] chore: enable sequence shard for wan by default (#19311) 2026-02-25 18:21:44 +08:00
Mick aff2f130ec [diffusion] CI: fix pr-test workflow file to include changes in jit kernels (#19285) 2026-02-25 14:08:55 +08:00
Mick 241ee90164 [diffusion] chore: tiny fix pyproject.toml (#19256) 2026-02-25 11:57:53 +08:00
Mick 0ede5c54a8 [diffusion] logging: improve request and component load logs (#19253) 2026-02-25 09:32:36 +08:00
Mickandzyzshishui 2e053d6eb6 [diffusion] quant: support quant for all dits (#19156)
Co-authored-by: zyzshishui <zyzshishui@gmail.com>
2026-02-24 14:20:54 +08:00
Mick 45095bac70 [diffusion] refactor: rename quantized model path server arg (#19142) 2026-02-22 23:18:35 +08:00
Mick 7d4860bc5e [diffusion] CI: relax perf check threshold (#19154) 2026-02-22 20:02:12 +08:00
Mick 87823722b3 [diffusion] chore: minor cleanups (#19123) 2026-02-22 19:07:25 +08:00
Mick 6503f94211 [diffusion] feat: support passing component path via server args (#19108) 2026-02-21 21:22:47 +08:00
Mick b89ca65789 [diffusion] refactor: reduce redundancy and improve stage api (#19060) 2026-02-21 16:35:47 +08:00
Mick 8d789b5c3d [diffusion] feat: support nunchaku for Z-Image-Turbo and flux.1 (int4) (#18959) 2026-02-20 21:16:08 +08:00
Mick 38a69652e6 [diffusion] logging: log available mem when each stage starts in debug level (#18998) 2026-02-20 19:57:06 +08:00
Mick 3207427d6d [diffusion] CI: enable warmup as default (#19010) 2026-02-19 23:27:23 +08:00
Mick d73f06f091 [diffusion] chore: improve memory usage on consumer-level GPU (#18997) 2026-02-19 21:59:49 +08:00
Mick 420a611275 [diffusion] refactor: unify SamplingParams construction and improve DiffGenerator return types (#18928) 2026-02-18 14:56:58 +08:00
Mick bfe34c90ff Revert "[diffusion] operator: unify rotary embedding impl" (#18929) 2026-02-17 22:56:04 +08:00
Mick de833f9e8e Revert "[diffusion]: Improve layerwise offload buffer reuse and shared-storage handling" (#18866) 2026-02-16 18:00:58 +08:00
Mick d0c94e136a [diffusion] logging: improve peak vram logging (#18865) 2026-02-16 16:44:37 +08:00
Mick 0af9dcc407 [diffusion] refactor: refactor server_args adjust and validate logics (#18863) 2026-02-16 11:49:06 +08:00
Mick 78b4c9e248 [diffusion] fix: avoid saving output for warmup requests (#18867) 2026-02-16 11:48:28 +08:00
3feb48139e [diffusion] quant: add support for svdquant and nunchaku (#18549)
Co-authored-by: AichenF <aichenf@nvidia.com>
Co-authored-by: jianyingzhu <53300651@qq.com>
2026-02-15 20:43:00 +08:00
Mick 37273408eb [diffusion] chore: use batched P2P ops in VAE parallel decoding (#18728) 2026-02-13 22:11:20 +08:00
Mick efdd676d56 [diffusion] refactor: merge redundant default_dtype and param_dtype parameters in FSDP loader (#18789) 2026-02-13 21:18:02 +08:00
Mick efcdda0176 [diffusion] fix: fix fsdp (#18187) 2026-02-10 20:22:20 +08:00
Mick 4f7da5ad0f [diffusion] chore: fix unclean shutdown and resource leaks (#18477) 2026-02-09 22:32:08 +08:00
Mick 6601bc24da [diffusion] chore: revise process title (#18446) 2026-02-09 00:14:06 +08:00
Mick a41aff1243 [diffusion] refactor: group component loaders under the component_loaders/ directory (#18438) 2026-02-08 23:02:27 +08:00
Mick 31d4cd2ffd [diffusion] fix: respect dist_timeout option (#18386) 2026-02-07 20:56:04 +08:00
Mick f218234e4f [diffusion] chore: prohibit Chinese characters usage (#18249) 2026-02-05 09:22:26 +08:00
Mick 36a3e78af9 [diffusion] refactor: move model_stages into stages folder (#18248) 2026-02-05 00:23:31 +08:00
Mick 62004fd2be [diffusion] UX: improve logging (#18122) 2026-02-03 10:35:05 +08:00
Mick c84cd4b5ff [diffusion] fix: fix missing component names for VAELoader (#18069) 2026-02-02 09:48:17 +08:00
Mick 977096ae03 [diffusion] cli: introduce generic attention backend configuration in ServerArgs (#18036) 2026-02-02 09:47:40 +08:00
Mick 1a006c2a0d [diffusion] refactor: split component_loader into component-wise files (#17820) 2026-01-31 20:22:31 +08:00
Mick 2573a262af [diffusion] doc: fix wrong docker run command (#17856) 2026-01-28 14:52:33 +08:00
Mick 88fcd8535f [diffusion] feat: add an arg for controlling the number of prefetched layers in layerwise-offload (#17693) 2026-01-28 09:34:27 +08:00
Mick 1507dc6cdf [diffusion] fix: fix suppressing error log on non-main ranks (#17712) 2026-01-28 09:29:19 +08:00
Mick b105dad5da [diffusion] refactor: remove useless lazy-import cache-dit codes (#17659) 2026-01-25 22:43:22 +08:00
MickandXiaoyu Zhang 51f147ada3 Update CODEOWNERS for multimodal_gen (#17308)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2026-01-19 08:39:20 +08:00
Mick 09491a9bcd cli: support sglang version (#17250) 2026-01-18 13:20:24 +08:00
Mick dc743fe4ba [diffusion] chore: clean srt imports (#17252) 2026-01-17 15:47:49 +08:00
Mick eb768189ee [diffusion] doc: add instruction for adding performance baseline of new model (#17249) 2026-01-17 12:24:59 +08:00
Mick 16831ab6d7 [diffusion] fix: fix using upstream flash_attn on blackwell (#17111) 2026-01-15 22:30:48 +08:00
Mick 68e8d0f68d [diffusion] CI: add testcase for cfg parallel (#17056) 2026-01-15 13:13:31 +08:00
Mick a5348eac4c [diffusion] chore: avoid raising error when output resolution is not optimal (#17030) 2026-01-14 11:36:27 +08:00
Mick 9524040220 [diffusion] chore: refactor warmup logic (#17027) 2026-01-14 11:35:06 +08:00
Mick 7a869045b6 [diffusion] chore: clean excessive document (#16986) 2026-01-13 21:33:50 +08:00
Mick 47d485f35f [diffusion] fix: fix not respecting dit_layerwise_offload server arg (#16252) 2026-01-13 09:29:07 +08:00
Mick 2b42309955 [diffusion] UX: provide solutions for OOM (#16940) 2026-01-13 09:25:27 +08:00
Mick badcd02896 [diffusion] chore: automatically enable dit_layerwise_offload for Wan (#16499) 2026-01-07 10:22:08 +08:00
Mick ca922d4b05 [diffusion] feat: support warmup with resolutions (#16434) 2026-01-06 10:32:18 +08:00
Mick 9a8ba3c189 [diffusion] feat: support warmup with resolutions (#16330) 2026-01-05 10:16:26 +08:00
Mick 2c09de343e [diffusion] improve: skip loading vision module for text encoders (#16304) 2026-01-03 19:30:45 +08:00
Mickandgemini-code-assist[bot] 38d48de93d [diffusion] CI: simplify warmup (#16303)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-03 18:09:38 +08:00
Mickandgemini-code-assist[bot] f26f6c2c99 [diffusion] fix: make lora compatible with layerwise-offload (#16298)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-02 22:33:26 +08:00
Mick 5062537b67 [diffusion] feat: support lightweight e2e warmup for benchmarking (#16213) 2026-01-02 20:10:27 +08:00
Mick 21de3e1406 [diffusion] webui: tiny fix loading output image (#16251) 2026-01-01 14:34:49 +08:00
Mickandgemini-code-assist[bot] 5bf0d862dd [diffusion] CI: fix generate mode and add cli test (#16174)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-31 09:51:48 +08:00
Mick 3449806727 [diffusion] feat: generalize layer-wise-offload to all supported models (#16150) 2025-12-30 22:06:57 +08:00
Mick 1e45320198 [diffusion] improve: tiny improve layerwise offload manager by consolidating weights per layer (#16081) 2025-12-30 11:31:00 +08:00
Mick 26e17f9076 [diffusion] improve: tiny speedup qwen-image-edit-2511 by avoiding unnecessary calculation (#15896) 2025-12-30 10:10:32 +08:00
MickandXiaoyu Zhang 8e08207c18 [diffusion] fix: fix serving with dit-layerwise-offload enabled (#16066)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2025-12-30 00:21:42 +08:00
Mick b5d9fc873b [diffusion] chore: minor refactor by streamlining the VAE class hierarchy (#16069) 2025-12-29 23:37:59 +08:00
Mick d7a3336ebe [diffusion] fix: fix stages not logged when perf_dump_path is provided (#16016) 2025-12-28 23:17:43 +08:00
Mick 9f8e23071a [diffusion] chore: fix default offload setting for image generation model (#15928) 2025-12-28 20:45:33 +08:00
Mick 3881bc8d0b [diffusion] CI: relax threshold by supporting different profiles (#16002) 2025-12-28 20:05:07 +08:00
Mick b4a00ed2d9 [diffusion] chore: clean ComposedPipelineBase (#15937) 2025-12-28 11:43:25 +08:00
Mick 39d56196a0 [diffusion] logging: log available gpu mem while loading and generating (#15936) 2025-12-28 00:34:58 +08:00
Mick aa89c6a7e2 [diffusion] refactor: unify model loading and offloading behavior (#15923) 2025-12-27 16:18:24 +08:00
Mickandgemini-code-assist[bot] b70914969b [diffusion] CI: support returning request id from endpoint (#15844)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-26 23:38:03 +08:00
Mick 8dc6f0fc4d [diffusion] refactor: centralize hardware platform detection and streamline environment variable management (#15842) 2025-12-26 22:16:18 +08:00
Mick a355911909 [diffusion] improve: improve post-processing by moving compute-intensive tasks to GPU (#15822) 2025-12-26 01:29:04 +08:00
Mick 2a8a785634 [diffusion] log: avoid logging in hot path if unnecessary (#15818) 2025-12-25 18:42:52 +08:00
Mick 2c5679f314 [diffusion] refactor: unify the profiling api for all executors (#15718) 2025-12-24 23:26:09 +08:00
Mick dfb5357448 [diffusion] http-server: relax openai image endpoint's strict content_type limit (#15717) 2025-12-24 13:26:13 +08:00
Mick 1ed946680a [diffusion] bench: improve bench_serving by adding more controlling args (#15554) 2025-12-21 13:37:45 +08:00
Mick d7fbe73bf2 [diffusion] chore: minor improvements and typo-fixing (#15556) 2025-12-21 13:37:10 +08:00
Mick bc18cb8631 [diffusion] refactor: change zmq socket type to router for scheduler (#15479) 2025-12-21 00:39:54 +08:00
Mick 41bd76e18b [diffusion] log: fix wrong use of suppress_other_loggers (#15534) 2025-12-20 23:43:55 +08:00
Mick c6ca1b3afc [diffusion] chore: allow all attention backends if not specified (#15530) 2025-12-20 23:14:55 +08:00