Portrait
Pooyan Rahmanzadehgervi
Computer Science Ph.D. Fellow
Auburn University
About Me

I am a Ph.D. candidate in Computer Science at Auburn University, advised by Anh Nguyen, and a Student Researcher at Google DeepMind. My research focuses on deep learning and explainable AI, with an emphasis on understanding and improving vision-language models (VLMs). I am a recipient of the Charles E. Gavin Doctoral Research Fellowship.

My work has shown that today's VLMs still fail at basic visual perception tasks that are trivial for humans, and has been covered by outlets including TechCrunch and ars Technica. Before Auburn, I completed my M.S. in Artificial Intelligence and Robotics at Ferdowsi University of Mashhad, advised by Ahad Harati and S.K. Ghiasi-Shirazi.

Education
  • Auburn University
    Auburn University
    Department of Computer Science and Software Engineering
    Ph.D. Candidate
    Aug. 2022 - present
  • Ferdowsi University of Mashhad
    Ferdowsi University of Mashhad
    M.S. in Artificial Intelligence and Robotics
    Sep. 2017 - Sep. 2021
  • Ferdowsi University of Mashhad
    Ferdowsi University of Mashhad
    B.S. in Computer Engineering - Software
    Sep. 2012 - Sep. 2017
Experience
  • Google DeepMind
    Google DeepMind
    Student Researcher (Ph.D.)
    Aug. 2026 - present
Honors & Awards
  • Charles E. Gavin Research Fellowship
    2022
  • National Organization for Exceptional Talents (NODET)
    National Organization for Exceptional Talents (NODET)
    2008-2012
  • National Organization for Exceptional Talents (NODET)
    National Organization for Exceptional Talents (NODET)
    2005-2008
News
2024
TechTalks: Why vision-language models fail on simple visual tests. Read more
Aug 01
News Bytes: Whatever be the claims, AI models can NOT actually see. Read more
Jul 20
TechSPOT: Study shows the best visual learning models fail at very basic visual identification tests. Read more
Jul 17
ars Technica: Can you do better than top-level AI models on these basic vision tests? Read more
Jul 15
Tech Xplore: Visual abilities of language models found to be lacking depth. Read more
Jul 12
TechCrunch: Are 'visual' AI models actually blind? Read more
Jul 11
Selected Publications
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen

IEEE/CVF International Conference on Computer Vision (ICCV) 2025

We propose a 1-head Transformer Attention Bottleneck (TAB) layer inserted after standard self-attention to constrain how much visual information flows through a VLM, enabling interpretable, user-adjustable intervention and debugging, demonstrated on image-difference captioning across three datasets.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen

IEEE/CVF International Conference on Computer Vision (ICCV) 2025

We propose a 1-head Transformer Attention Bottleneck (TAB) layer inserted after standard self-attention to constrain how much visual information flows through a VLM, enabling interpretable, user-adjustable intervention and debugging, demonstrated on image-difference captioning across three datasets.

Vision language models are blind: Failing to translate detailed visual features into words

Pooyan Rahmanzadehgervi*, Logan Bolton*, Mohammad Reza Taesiri, Anh Totti Nguyen (* equal contribution)

Asian Conference on Computer Vision (ACCV) 2024 Oral

While VLMs like GPT-4o and Gemini 1.5 Pro score highly on many vision benchmarks, on our BlindTest suite of 7 simple low-level vision tasks (e.g., detecting overlapping circles or counting line intersections), four state-of-the-art VLMs average only 58.07% accuracy, far below human-level performance.

Vision language models are blind: Failing to translate detailed visual features into words

Pooyan Rahmanzadehgervi*, Logan Bolton*, Mohammad Reza Taesiri, Anh Totti Nguyen (* equal contribution)

Asian Conference on Computer Vision (ACCV) 2024 Oral

While VLMs like GPT-4o and Gemini 1.5 Pro score highly on many vision benchmarks, on our BlindTest suite of 7 simple low-level vision tasks (e.g., detecting overlapping circles or counting line intersections), four state-of-the-art VLMs average only 58.07% accuracy, far below human-level performance.

All publications