2026

Vision Language Models Cannot Reason About Physical Transformation

Dezhi Luo, Yijiang Li, Maijunxian Wang, Tianwei Zhao, Bingyang Wang, Siheng Wang, Pinyuan Feng, Pooyan Rahmanzadehgervi, Ziqiao Ma, Hokin Deng

arXiv preprint arXiv:2603.07109 2026

We introduce ConservationBench, testing whether VLMs maintain transformation-invariant understanding of physical quantities (number, length, volume, size) across 23,040 questions and 112 models; performance stays near chance and often worsens with real visual content compared to text-only priors.

Vision Language Models Cannot Reason About Physical Transformation

Dezhi Luo, Yijiang Li, Maijunxian Wang, Tianwei Zhao, Bingyang Wang, Siheng Wang, Pinyuan Feng, Pooyan Rahmanzadehgervi, Ziqiao Ma, Hokin Deng

arXiv preprint arXiv:2603.07109 2026

We introduce ConservationBench, testing whether VLMs maintain transformation-invariant understanding of physical quantities (number, length, volume, size) across 23,040 questions and 112 models; performance stays near chance and often worsens with real visual content compared to text-only priors.

2025

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen

IEEE/CVF International Conference on Computer Vision (ICCV) 2025

We propose a 1-head Transformer Attention Bottleneck (TAB) layer inserted after standard self-attention to constrain how much visual information flows through a VLM, enabling interpretable, user-adjustable intervention and debugging, demonstrated on image-difference captioning across three datasets.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen

IEEE/CVF International Conference on Computer Vision (ICCV) 2025

We propose a 1-head Transformer Attention Bottleneck (TAB) layer inserted after standard self-attention to constrain how much visual information flows through a VLM, enabling interpretable, user-adjustable intervention and debugging, demonstrated on image-difference captioning across three datasets.

Vision Language Models fail to translate detailed visual features into words

Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, Anh Totti Nguyen

eXCV Workshop, International Conference on Computer Vision (ICCV) 2025

Vision Language Models fail to translate detailed visual features into words

Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, Anh Totti Nguyen

eXCV Workshop, International Conference on Computer Vision (ICCV) 2025

Improving zero-shot object-level change detection by incorporating visual correspondence

Hung Huy Nguyen, Pooyan Rahmanzadehgervi, Long Mai, Anh Totti Nguyen

IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025

We introduce a change-detection method that leverages correspondences between before/after regions during training and at test time to reduce false positives, improving zero-shot object-level change detection accuracy across in-distribution and cross-domain benchmarks.

Improving zero-shot object-level change detection by incorporating visual correspondence

Hung Huy Nguyen, Pooyan Rahmanzadehgervi, Long Mai, Anh Totti Nguyen

IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025

We introduce a change-detection method that leverages correspondences between before/after regions during training and at test time to reduce false positives, improving zero-shot object-level change detection accuracy across in-distribution and cross-domain benchmarks.

2024

Vision language models are blind: Failing to translate detailed visual features into words

Pooyan Rahmanzadehgervi*, Logan Bolton*, Mohammad Reza Taesiri, Anh Totti Nguyen (* equal contribution)

Asian Conference on Computer Vision (ACCV) 2024 Oral

While VLMs like GPT-4o and Gemini 1.5 Pro score highly on many vision benchmarks, on our BlindTest suite of 7 simple low-level vision tasks (e.g., detecting overlapping circles or counting line intersections), four state-of-the-art VLMs average only 58.07% accuracy, far below human-level performance.

Vision language models are blind: Failing to translate detailed visual features into words

Pooyan Rahmanzadehgervi*, Logan Bolton*, Mohammad Reza Taesiri, Anh Totti Nguyen (* equal contribution)

Asian Conference on Computer Vision (ACCV) 2024 Oral

While VLMs like GPT-4o and Gemini 1.5 Pro score highly on many vision benchmarks, on our BlindTest suite of 7 simple low-level vision tasks (e.g., detecting overlapping circles or counting line intersections), four state-of-the-art VLMs average only 58.07% accuracy, far below human-level performance.

2021

Vision-based obstacle avoidance in drone navigation using deep reinforcement learning

Pooyan Rahmanzadehgervi, Ahad Harati, Sayed Kamaledin Ghiasi-Shirazi

International Conference on Computer Engineering and Knowledge (ICCKE) 2021

We propose a bio-inspired deep reinforcement learning architecture for autonomous obstacle avoidance in drone navigation, learning a generalizable collision-avoidance policy in simulation and comparing it against a human expert baseline.

Vision-based obstacle avoidance in drone navigation using deep reinforcement learning

Pooyan Rahmanzadehgervi, Ahad Harati, Sayed Kamaledin Ghiasi-Shirazi

International Conference on Computer Engineering and Knowledge (ICCKE) 2021

We propose a bio-inspired deep reinforcement learning architecture for autonomous obstacle avoidance in drone navigation, learning a generalizable collision-avoidance policy in simulation and comparing it against a human expert baseline.