Intelligent visual perception beyond classification: a survey of deep learning and foundation models
Intelligent visual perception has advanced from traditional image classification to thorough comprehension of objects, scenes, relations, spatial context, and semantic information. This paper surveys the development of intelligent visual perception, particularly focusing on DL and foundation models. This section provides a comprehensive overview of visual recognition, detection, segmentation, and scene understanding, including key methods such as convolutional neural networks, recurrent neural networks, attention mechanisms, encoder-decoder architectures, generative models, and Vision Transformers. The study also explores feature attribution, visual attention, visual reasoning, and multimodal vision-language models, which combine visual and text information for tasks such as visual question answering, captioning, and visual generation. This study examines recent foundation models incorporating vision-language architectures for their transferability, contextual understanding, and generalized perception capabilities. The study also discusses challenges related to computational efficiency, environmental variability, data requirements, robustness, explainability, and real-time deployment. Finally, it contrasts current developments to identify new approaches to scalable, interpretable, multimodal, and human-centric visual intelligence systems.
- Y. Hu, H. Chi, and H. Duan, “Metaoptics merging computational optics and optical computing toward intelligent visual perception,” Sci. Adv., vol. 12, no. 1, p. eaea8941, 2026, doi: 10.1126/sciadv.aea8941.
- M. N. Paralath, “Towards Interpretable and Attentive Deep Visual Perception Through Iterative Refinement in Convolutional Neural Networks,” in 2025 International Conference on Advances in Next-Gen Computer Science (ICANCS), IEEE, Nov. 2025, pp. 1–6. doi: 10.1109/ICANCS65819.2025.11376668.
- Y. Li, X. Guo, H. Zhang, S. Li, and X. Dai, “Active Visual Perception: Opportunities and Challenges,” pp. 1–8, 2025, doi: 10.48550/arXiv.2512.03687.
- A. Banan, A. Nasiri, and A. Taheri-Garavand, “Deep learning-based appearance features extraction for automated carp species identification,” Aquac. Eng., vol. 89, p. 102053, May 2020, doi: 10.1016/j.aquaeng.2020.102053.
- P. Wang, E. Fan, and P. Wang, “Comparative analysis of image classification algorithms based on traditional machine learning and deep learning,” Pattern Recognit. Lett., vol. 141, pp. 61–67, Jan. 2021, doi: 10.1016/j.patrec.2020.07.042.
- V. Monga, Y. Li, and Y. C. Eldar, “Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing,” IEEE Signal Process. Mag., vol. 38, no. 2, pp. 18–44, 2021, doi: 10.1109/MSP.2020.3016905.
- X. Li et al., “SARPointNet: An Automated Feature Learning Framework for Spaceborne SAR Image Registration,” IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 15, pp. 6371–6381, 2022, doi: 10.1109/JSTARS.2022.3196383.
- M. Alshayeji, J. Al-Buloushi, A. Ashkanani, and S. Abed, “Enhanced brain tumor classification using an optimized multi-layered convolutional neural network architecture,” Multimed. Tools Appl., vol. 80, no. 19, pp. 28897–28917, 2021, doi: 10.1007/s11042-021-10927-8.
- M. Trigka and E. Dritsas, “A Comprehensive Survey of Deep Learning Approaches in Image Processing,” Sensors, vol. 25, no. 2, p. 531, Jan. 2025, doi: 10.3390/s25020531.
- M. R. C. Mukkolakkal, “1.5 Million Messages Per Second on 3 Machines: Benchmarking and Latency Optimization of Apache Pulsar at Enterprise Scale,” in 22nd International Conference on Network and Service Management, 2026, pp. 1–4, Mar. doi: 10.48550/arXiv.2603.29113.
- M. Awais et al., “Foundation Models Defining a New Era in Vision: A Survey and Outlook,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 4, pp. 2245–2264, 2025, doi: 10.1109/TPAMI.2024.3506283.
- S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in Vision: A Survey,” ACM Comput. Surv., vol. 54, no. 10s, pp. 1–41, Jan. 2022, doi: 10.1145/3505244.
- L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, “VisualBERT: A Simple and Performant Baseline for Vision and Language,” no. 2, pp. 1–14, 2019, doi: 10.48550/arXiv.1908.03557.
- A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “MDETR - Modulated Detection for End-to-End Multi-Modal Understanding,” Proc. IEEE Int. Conf. Comput. Vis., pp. 1760–1770, 2021, doi: 10.1109/ICCV48922.2021.00180.
- W. El Amrani, Attention and Beyond: Explainability Techniques for Vision Transformers. 2025.Conference: Knowledge Extraction and ManagementAt: Strasbourg, FranceVolume: Explain'AI
- Z. Liu and B. Jiang, “Revisiting the Ordering of Channel and Spatial Attention : A Comprehensive Study on Sequential and Parallel Designs,” vol. 20, pp. 1–26.
- S. Uppal et al., “Multimodal research in vision and language: A review of current and emerging trends,” Inf. Fusion, vol. 77, pp. 149–171, Jan. 2022, doi: 10.1016/j.inffus.2021.07.009.
- A. Bose, J. Bhumireddy, and N. Naveen, “A Study on Real-time Object Detection using Deep Learning,” pp. 1–34, 2026, doi: 10.48550/arXiv.2602.15926.
- S. Minaee, Y. Y. Boykov, F. Porikli, A. J. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image Segmentation Using Deep Learning: A Survey,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–1, 2021, doi: 10.1109/TPAMI.2021.3059968.
- T. M. Le, Deep Neural Networks for Visual Reasoning. 2022. doi: 10.48550/arXiv.2209.11990.
- S. Li, X. Chen, Y. Liu, D. Dai, C. Stachniss, and J. Gall, “Multi-Scale Interaction for Real-Time LiDAR Data Segmentation on an Embedded Platform,” IEEE Robot. Autom. Lett., vol. 7, no. 2, pp. 738–745, Apr. 2022, doi: 10.1109/LRA.2021.3132059.
- A. Maham and D. E. N. Tashfa, “Deep Learning Perspective of Scene Understanding in Autonomous Robots,” 2025, doi: 10.48550/arXiv.2512.14020.
- T. M. Strat, R. Chellappa, and V. M. Patel, “Vision and Robotics,” AI Mag., vol. 41, no. 2, pp. 49–65, Jun. 2020, doi: 10.1609/aimag.v41i2.5299.
- B. Wei, Y. Zhao, K. Hao, and L. Gao, “Visual Sensation and Perception Computational Models for Deep Learning : State of the art , Challenges and Prospects,” 2021, doi: 10.48550/arXiv.2109.03391.
- X. Zhu, K. Hao, R. Xie, and B. Huang, “Soft sensor based on eXtreme gradient boosting and bidirectional converted gates long short-term memory self-attention network,” Neurocomputing, vol. 434, pp. 126–136, Apr. 2021, doi: 10.1016/j.neucom.2020.12.028.
- Y. Liu, Z. Zhao, C. Wang, and N. Du, “A Study on Image Understanding and Visual-Semantic Alignment Modeling Methods for Intelligent Perception,” in 2026 International Conference on Intelligent Perception and Autonomous Control (IPAC), 2026, pp. 380–383. doi: 10.1109/IPAC69021.2026.11542612.
- H. Zhao, X. Ji, X. Li, Y. Zhang, and L.-Y. Hao, “MMAN: Multi-scale Mixed Attention Network for Underwater Visual Perception Enhancement,” in 2026 41st Youth Academic Annual Conference of Chinese Association of Automation (YAC), IEEE, May 2026, pp. 169–174. doi: 10.1109/YAC71005.2026.11616203.
- G. Yang, Y. Wang, B. Qiu, J.-P. Wu, S. Zheng, and J.-W. Liu, “A Review of Visual Perception and Generation Technologies in Intelligent Transportation Scenarios,” in 2025 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI), 2025, pp. 1–7. doi: 10.1109/ACAI68217.2025.11406561.
- C. Zhang et al., “Global-Mapping-Consistency-Constrained Visual-Semantic Embedding for Interpreting Autonomous Perception Models,” IEEE Open J. Intell. Transp. Syst., vol. 5, pp. 393–408, 2024, doi: 10.1109/OJITS.2024.3418552.
- C. Du, K. Fu, J. Li, and H. He, “Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 9, pp. 10760–10777, Sep. 2023, doi: 10.1109/TPAMI.2023.3263181.
- B.-H. Kwon, B.-H. Lee, J.-H. Cho, and J.-H. Jeong, “Decoding Visual Imagery from EEG Signals using Visual Perception Guided Network Training Method,” in 2022 10th International Winter Conference on Brain-Computer Interface (BCI), IEEE, Feb. 2022, pp. 1–5. doi: 10.1109/BCI53720.2022.9735014.


