I am a Ph.D. candidate in Computer Science at Auburn University, advised by Anh Nguyen, and a Student Researcher at Google DeepMind. My research focuses on deep learning and explainable AI, with an emphasis on understanding and improving vision-language models (VLMs). I am a recipient of the Charles E. Gavin Doctoral Research Fellowship.
My work has shown that today's VLMs still fail at basic visual perception tasks that are trivial for humans, and has been covered by outlets including TechCrunch and ars Technica. Before Auburn, I completed my M.S. in Artificial Intelligence and Robotics at Ferdowsi University of Mashhad, advised by Ahad Harati and S.K. Ghiasi-Shirazi.
") does not match the recommended repository name for your site ("").
", so that your site can be accessed directly at "http://".
However, if the current repository name is intended, you can ignore this message by removing "{% include widgets/debug_repo_name.html %}" in index.html.
",
which does not match the baseurl ("") configured in _config.yml.
baseurl in _config.yml to "".
Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen
IEEE/CVF International Conference on Computer Vision (ICCV) 2025
We propose a 1-head Transformer Attention Bottleneck (TAB) layer inserted after standard self-attention to constrain how much visual information flows through a VLM, enabling interpretable, user-adjustable intervention and debugging, demonstrated on image-difference captioning across three datasets.
Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen
IEEE/CVF International Conference on Computer Vision (ICCV) 2025
We propose a 1-head Transformer Attention Bottleneck (TAB) layer inserted after standard self-attention to constrain how much visual information flows through a VLM, enabling interpretable, user-adjustable intervention and debugging, demonstrated on image-difference captioning across three datasets.
Pooyan Rahmanzadehgervi*, Logan Bolton*, Mohammad Reza Taesiri, Anh Totti Nguyen (* equal contribution)
Asian Conference on Computer Vision (ACCV) 2024 Oral
While VLMs like GPT-4o and Gemini 1.5 Pro score highly on many vision benchmarks, on our BlindTest suite of 7 simple low-level vision tasks (e.g., detecting overlapping circles or counting line intersections), four state-of-the-art VLMs average only 58.07% accuracy, far below human-level performance.
Pooyan Rahmanzadehgervi*, Logan Bolton*, Mohammad Reza Taesiri, Anh Totti Nguyen (* equal contribution)
Asian Conference on Computer Vision (ACCV) 2024 Oral
While VLMs like GPT-4o and Gemini 1.5 Pro score highly on many vision benchmarks, on our BlindTest suite of 7 simple low-level vision tasks (e.g., detecting overlapping circles or counting line intersections), four state-of-the-art VLMs average only 58.07% accuracy, far below human-level performance.