Assistant professor of computer science and software engineering leads effort to improve AI video surveillance, earns $417K NSF award
Published: Aug 5, 2026 1:30 PM
By Joe McAdory
Pan He is developing new AI-powered video surveillance systems that can identify important events, explain their conclusions and help human operators make better-informed decisions.
Modern surveillance systems generate more video footage than humans can realistically monitor. That’s a problem.
Many artificial intelligence (AI) surveillance tools remain difficult to trust, understand or use effectively. That’s another problem.
Pan He has answers.
Through a recent $417K National Science Foundation award-winning project, "Toward Human-Centered, Generalizable, and Verifiable Automated Video Surveillance: A Verbalized Vision-Language Model Paradigm," the assistant professor of computer science and software engineering will work with his collaborator, Muchao Ye from the University of Iowa, to develop new AI-powered video surveillance systems that can identify important events, explain their conclusions and help human operators make better-informed decisions.
“Our project uses vision-language models to address these challenges,” said He, the project’s principle investigator. “The goal is to develop surveillance systems capable of reasoning in ways that are intuitive and meaningful to human operators, synthesize evidence across cameras and time and produce findings that are both explainable and verifiable. This would make the systems more practical, reliable and useful for real-world decision-making.
Current surveillance systems often issue alerts without clearly explaining their reasoning.
“An alert is far more useful when the system can explain what happened, why it matters, and what visual evidence supports its conclusion,” He said. “For example, rather than simply flagging an anomaly, the system might identify a vehicle traveling the wrong way, point to the relevant moment in the footage and explain which observations led to that finding.
“This is important because automated systems can make mistakes due to lighting, camera angles, crowds, shadows, or unfamiliar environments,” He continued. “Human operators should be able to review the evidence, question the system’s reasoning and make the final decision.”
The technology is also designed to make video searchable using natural language. Rather than manually combing through hours of footage, users could simply describe the event they are looking for, such as ‘find when a vehicle stopped at the intersection,’ and allow the system to locate relevant clips.
“The system would analyze the footage in short segments, identify the moments that best match the request and return only the most relevant clips,” He said. “It would also recognize different phrases, such as ‘a person fell’ and ‘a worker lost balance.’”
Despite advances in AI, He emphasizes that the technology is intended to support people rather than replace them.
“Human judgment remains essential because the meaning of an event often depends on context that an automated system may not fully understand.” He said. “The goal is human-AI collaboration: AI handles the time-consuming work of searching and organizing video, while people provide context, accountability and the final decision.
“It should also be evaluated not only by technical accuracy, but by whether it improves human decision-making without overwhelming users or encouraging misplaced trust. Privacy, responsible deployment, and continued human communication are therefore essential parts of the framework.”
He and researchers will evaluate the technology in workplace safety and traffic-monitoring environments, testing whether it can identify safety-related activities, unusual behavior and other events of interest while providing information that human operators can understand.
Beyond this project, He envisions “a new generation of intelligent video systems” that can reason, learn and interact with users in far more sophisticated ways.
“Our broader vision is to move from passive video analysis to agentic video intelligence,” He said. “Most current systems are designed for narrow, predefined tasks and remain fixed after deployment. Future video systems should be able to understand complex events, search hours of footage, connect information across time and multiple cameras, answer questions in natural language, explain the evidence behind their conclusions and adapt to new environments and operational needs.
“Most importantly, these systems should remain human-centered. Their role is not to replace human judgment, but to provide transparent, verifiable and domain-informed support.”
Media Contact: , jem0040@auburn.edu, 334.844.3447
