Photo Aesthetic Assessment via Spatial and Frequency Domain Fusion Using a Vision Transformer Approach

Authors

  • Reza Rachmadan Institut Teknologi Sepuluh Nopember
  • Shintami Chusnul Hidayati Institut Teknologi Sepuluh Nopember

DOI:

https://doi.org/10.59261/jbt.v7i3.737

Keywords:

AADB Datasets, Adaptive Gate Fusion, Concatenation Fusion, Dual-branch, Fusion, Fast Fourier Transform, Image Aesthetic Assessment, Vision Transformer

Abstract

Background: The rapid growth of digital media has increased the demand for automated image aesthetic assessment (IAA) in social media, creative industries, and e-commerce platforms. Although Vision Transformer (ViT)-based models have demonstrated promising performance, most existing approaches rely primarily on spatial representations while overlooking frequency-domain information, which captures complementary characteristics such as sharpness, texture, noise patterns, and bokeh effects.

Objective: This study proposes a Dual-Branch Late Fusion FFT-ViT architecture with concatenation-based fusion as an approach for integrating dual-domain representations to improve photo aesthetic assessment (PAA).

Methods: The proposed architecture consists of two parallel branches. The RGB branch employs a Vision Transformer (ViT-Small) to extract spatial and compositional features, while the FFT branch utilizes ViT-Tiny to capture frequency-domain characteristics associated with texture, sharpness, and image details. The extracted features from both branches are fused using a concatenation strategy before being passed to the regression layer.

Results: The proposed Dual-Branch FFT-ViT with concatenation fusion achieved the best performance, obtaining a PLCC of 0.7336, SRCC of 0.7347, MSE of 0.0193, MAE of 0.1117, and RMSE of 0.1388. Compared with the RGB-only ViT baseline, the proposed model improved the PLCC score by 0.0124, demonstrating the effectiveness of integrating spatial and frequency-domain features for aesthetic score prediction.

Conclusion: This study demonstrates that integrating spatial and frequency-domain representations through a dual-branch Vision Transformer architecture enhances photo aesthetic assessment performance.

References

Anwar, A., Kanwal, S., Tahir, M., Saqib, M., Uzair, M., Rahmani, M. K. I., & Ullah, H. (2021). A Survey on Image Aesthetic Assessment. ArXiv. https://doi.org/10.48550/arXiv.2103.11616

Behrad, F., Tuytelaars, T., & Wagemans, J. (2025). Charm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic Assessment. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7815–7824. https://doi.org/10.1109/CVPR52734.2025.00732

Gao, M., Song, C., Zhang, Q., Zhang, X., Li, Y., & Yuan, F. (2025). Research Progress on Color Image Quality Assessment. Journal of Imaging, 11(9), 307. https://doi.org/10.3390/jimaging11090307

Gonzalez-Naharro, L., Flores, M. J., Martínez-Gómez, J., & Puerta, J. M. (2025). Analysis of image aesthetics assessment as a positive-unlabeled problem. Signal Processing: Image Communication, 117441.

He, S., Ming, A., Zheng, S., Zhong, H., & Ma, H. (2023). EAT: An Enhancer for Aesthetics-Oriented Transformers. Proceedings of the 31st ACM International Conference on Multimedia, 1023–1032. https://doi.org/10.1145/3581783.3611881

He, S., Zhang, Y., Xie, R., Jiang, D., & Ming, A. (2022). Rethinking Image Aesthetics Assessment: Models, Datasets and Benchmarks. Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI-22), 942–948. https://doi.org/10.24963/ijcai.2022/132

Jia, M., Xing, X., He, J., & Xu, X. (2025). Self-Supervised Image Aesthetic Assessment Based on Transformers. International Journal of Semantic Computing, 18(4), 519–536. https://doi.org/10.1142/S1469026824500299

Jiang, X., Zhang, X., Gao, N., & Deng, Y. (2025). When Fast Fourier Transform Meets Transformer for Image Restoration. Computer Vision – ECCV 2024, 15103, 381–402. https://doi.org/10.1007/978-3-031-72995-9_22

Ke, J., Wang, Q., Wang, Y., Milanfar, P., & Yang, F. (2021). MUSIQ: Multi-Scale Image Quality Transformer. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 5128–5137. https://doi.org/10.1109/ICCV48922.2021.00510

Kong, S., Shen, X., Lin, Z., Mech, R., & Fowlkes, C. (2016). Photo Aesthetics Ranking Network with Attributes and Content Adaptation. Computer Vision – ECCV 2016, 9905, 662–679. https://doi.org/10.1007/978-3-319-46448-0_40

Li, L., Huang, Y., Wu, J., Yang, Y., Li, Y., Guo, Y., & Shi, G. (2023). Theme-Aware Visual Attribute Reasoning for Image Aesthetics Assessment. IEEE Transactions on Circuits and Systems for Video Technology, 33(9), 4798–4811. https://doi.org/10.1109/TCSVT.2023.3249185

Li, Z., Yan, X., Wei, X., & Shao, F. (2025). IAACLIP: Image Aesthetics Assessment via CLIP. Electronics, 14(7), 1425. https://doi.org/10.3390/electronics14071425

Sheng, K., Dong, W., Ma, C., Mei, X., Huang, F., & Hu, B.-G. (2018). Attention-Based Multi-Patch Aggregation for Image Aesthetic Assessment. Proceedings of the 26th ACM International Conference on Multimedia, 879–886. https://doi.org/10.1145/3240508.3240554

Shi, J., Gao, P., & Smolic, A. (2024). Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality Token. IEEE Transactions on Multimedia, 26, 4641–4651. https://doi.org/10.1109/TMM.2023.3325719

Shi, T., Chen, C., Wu, Z., Hao, A., & Fang, Y. (2024). Improving Image Aesthetic Assessment via Multiple Image Joint Learning. ACM Transactions on Multimedia Computing, Communications, and Applications, 20(11). https://doi.org/10.1145/3687128

Sun, W., Zhang, W., Cao, Y., Cao, L., Jia, J., Chen, Z., Zhang, Z., Min, X., & Zhai, G. (2024). Assessing UHD Image Quality from Aesthetics, Distortions, and Saliency. ArXiv. https://doi.org/10.48550/arXiv.2409.00749

Talebi, H., & Milanfar, P. (2018). NIMA: Neural Image Assessment. IEEE Transactions on Image Processing, 27(8), 3998–4011. https://doi.org/10.1109/TIP.2018.2831899

Talebi, H., & Milanfar, P. (2021). Learning to Resize Images for Computer Vision Tasks. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 487–496. https://doi.org/10.1109/ICCV48922.2021.00055

Tanawongsuwan, R., & Phongsuphap, S. (2026). Flip-Robust Neural Image Assessment (FR-NIMA) for Spatially Consistent IQA. ECTI Transactions on Computer and Information Technology (ECTI-CIT), 20(2), 303–318.

Xu, L., Xu, J., Yang, Y., Wang, X., Huang, Y.-J., & Li, Y. (2025). CLIP Brings Better Features to Visual Aesthetics Learners. 2025 IEEE International Conference on Multimedia and Expo (ICME), 1–6. https://doi.org/10.1109/ICME59968.2025.11210032

Yang, H., Li, Y., Jin, X., Zhou, X., Shi, P., & Liu, Y. (2025). Aesthetic multi-attributes network for image captioning. Computers and Electrical Engineering, 123, 110103.

Zeng, H., Cao, Z., Zhang, L., & Bovik, A. C. (2019). A Unified Probabilistic Formulation of Image Aesthetic Assessment. IEEE Transactions on Image Processing, 29, 1548–1561. https://doi.org/10.1109/TIP.2019.2941778

Zhang, R., Dong, W., Lu, L., Zhou, Z., Zhang, T., & Niu, W. (2026). Multi-scale wavelet vision transformer with HDR-aware attention for high dynamic range image quality assessment. The Journal of Supercomputing, 82(9), 509.

Zheng, Q., Shi, Y., & Wu, Y. (2025). Image Aesthetic Assessment Based on Multi-level Hierarchical Adaptive Fusion. Chinese Conference on Pattern Recognition and Computer Vision (PRCV), 354–368.

Zhou, X., Jin, X., Lv, J., Huang, H., Mao, M., & Cui, S. (2022). Aesthetic Attributes Assessment of Images with AMANv2 and DPC-CaptionsV2. ArXiv Preprint ArXiv:2208.04522.

Downloads

Published

2026-08-20