Authors
- Victor Shea-Jay Huang Multimedia Laboratory (MMLab), The Chinese University of Hong Kong Shanghai Artificial Intelligence Laboratory
- Le Zhuo Multimedia Laboratory (MMLab), The Chinese University of Hong Kong Shanghai Artificial Intelligence Laboratory
- Yi Xin Shanghai Artificial Intelligence Laboratory
- Zhaokai Wang Shanghai Artificial Intelligence Laboratory
- Fu-Yun Wang Multimedia Laboratory (MMLab), The Chinese University of Hong Kong
- Yuchi Wang Multimedia Laboratory (MMLab), The Chinese University of Hong Kong
- Renrui Zhang Multimedia Laboratory (MMLab), The Chinese University of Hong Kong
- Peng Gao Shanghai Artificial Intelligence Laboratory
- Hongsheng Li Multimedia Laboratory (MMLab), The Chinese University of Hong Kong Shanghai Artificial Intelligence Laboratory CPII under InnoHK
DOI:
https://doi.org/10.1609/aaai.v40i1.37006Abstract
Diffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE—Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs—a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer.How to Cite
Huang, V. S.-J., Zhuo, L., Xin, Y., Wang, Z., Wang, F.-Y., Wang, Y., … Li, H. (2026). TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation. Proceedings of the AAAI Conference on Artificial Intelligence, 40(1), 435–443. https://doi.org/10.1609/aaai.v40i1.37006
Issue
Section
AAAI Technical Track on Application Domains I
















