--- title: "Performance Optimization" description: "Optimize SGLang diffusion performance with caching, kernels, and profiling." --- This section covers the main performance levers for SGLang Diffusion: attention backends, caching acceleration, and profiling. ## Overview
Optimization Type Description
Cache-DiT Caching Block-level caching with DBCache, TaylorSeer, and SCM
TeaCache Caching Timestep-level caching based on temporal similarity
Attention Backends Kernel Optimized attention implementations (FlashAttention, SageAttention, etc.)
Inference Batching Scheduler Request batching for native diffusion serving
Profiling Diagnostics PyTorch Profiler and Nsight Systems guidance
## Start Here - Use [Attention Backends](./attention_backends) to choose the best backend for your model and hardware. - Use [Inference Batching](./dynamic_batching) to improve throughput for compatible concurrent requests. - Use [Caching Acceleration](./caching-acceleration) to reduce denoising cost with Cache-DiT or TeaCache. - Use [Profiling](./profiling) when you need to diagnose a bottleneck rather than guess. ## Caching at a Glance - [Cache-DiT](./cache_dit) is block-level caching for diffusers pipelines and higher speedup-oriented tuning. - [TeaCache](./teacache) is timestep-level caching built into SGLang model families. ## Current Baseline Snapshot For Ring SP benchmark details, see: - [Ring SP Performance](./ring_sp_performance) ## References - [Cache-DiT Repository](https://github.com/vipshop/cache-dit) - [TeaCache Paper](https://arxiv.org/abs/2411.14324)