Change Detection Filtering with Visual Language Models
2026-01-7508
9/22/2026
- Content
- Intelligence, surveillance and reconnaissance (ISR) often require review of significant quantities of video. While machine vision is used to flag objects for human review, too many flags are generated. Integrating newer methods like Open-Vocabulary Object Detection (OVOD) that support zero shot detection, but with significantly lower accuracy only make the problem worse. This work addresses the utility of OVOD in ISR missions by focusing only what has changed between successive runs through an environment. A Vision Language Model (VLM) compares current observations against a registered “cleared” baseline to focus only on what has changed. Testing across three distinct environments, and using either monocular camera phones or RGB-D equipped vehicles, demonstrates that integrating change detection can automatically remove as much as 80% of unchanged objects without impacting recall.
- Citation
- Martinson, E. and Fishta, I., "Change Detection Filtering with Visual Language Models," 2026 NDIA Michigan Chapter Ground Vehicle Systems Engineering and Technology Symposium, Novi, Michigan, United States, August 11, 2026, https://doi.org/10.4271/2026-01-7508.