Doctoral student in CSSE is lead author on research paper published in leading computational science and AI journal
Published: Sep 4, 2026 7:50 AM
By Rachel Wingard
Pooyan Rahmanzadehgervi is interning at Google DeepMinds in San Francisco.
How well can vision-language models (VLMs) actually see? Pooyan Rahmanzadehgervi wants to find out.
Rahmanzadehgervi, a doctoral student in the Department of Computer Science and Software Engineering (CSSE), not only answered the question but served as lead author the paper, "Vision-language models are blind: Failing to translate detailed visual features into words," accepted for publication in “Nature Communications AI & Computing,” a leading journal in computational science and artificial intelligence (AI).
Joining Rahmanzadehgervi were co-authors Logan Bolton, who earned his bachelor’s degree in computer science from Auburn this past May and is pursuing a master's degree at New York University, Anh Nguyen, associate professor of computer science, and machine learning scientist Mohammad Reza Taesiri.
The paper is an extension of the 2024 research, "Vision-language models are blind: Failing to translate detailed visual features into words," which inspired the sequel. Research shared in these papers has been featured by Ars Technica and Tech Crunch and covered in online AI lectures by industry lead scientists.
The 2026 paper focuses on creating BlindTest, a benchmark of seven simple vision tasks, which demonstrate the accuracy of VLMs’ vision capabilities. It found that eight state-of-the-art VLMs could reach only 46% average accuracy, showing how VLMs still have much room for improvement.
“We test skills ranging from simple spatial relationships, like spotting overlapping shapes, to complex tasks like following a route on a transit map,” Rahmanzadehgervi said. “A VLM must master these tasks for real-world deployment.”
An example of these applications includes robotics, where a robot in a warehouse will need to judge physical boundaries between objects to grasp them safely. Another example, Rahmanzadehgervi said, could be an AI model reviewing legal files. Finely tuned visual skills are required to recognize characters and accurately process a document's physical details.
Rahmanzadehgervi said that while VLMs have no trouble “seeing” visual details, they often run into trouble processing them.
“The models try to shove every visual detail from an image into their working memory, the language space, all at once just in case they are asked a question targeting it,” he said. “This creates a massive data bottleneck. As a result, the AI's visual sensors might perfectly capture a tiny detail, but that detail gets lost in the crush of information and never successfully reaches the AI's 'conscious' language space.”
At his current internship as a doctoral student researcher with Google DeepMinds in San Francisco — one of the top AI labs in the world — Rahmanzadehgervi is investigating ways to bridge this ‘consciousness gap’ in AI models. His goal is to lead research that helps AI acquire basic human cognitive skills and shape how AI applies them in the real world.
He said his coursework at Auburn helped him reach where he is now.
“Auburn’s computer science and software engineering AI courses prepared me perfectly for my career by giving me the freedom to explore bold ideas and translate them into real-world applications,” he said. “My advisor, Dr. Anh Nguyen, truly champions this kind of creative exploration. Under his guidance, I learned how to think independently and test out-of-the-box solutions, which is the exact mindset I rely on for my research today.”
Media Contact: , jem0040@auburn.edu, 334.844.3447
