Change Detection Filtering with Visual Language Models

2026-01-7508

9/22/2026

Authors
Abstract
Content
Intelligence, surveillance and reconnaissance (ISR) often require review of significant quantities of video. While machine vision is used to flag objects for human review, too many flags are generated. Integrating newer methods like Open-Vocabulary Object Detection (OVOD) that support zero shot detection, but with significantly lower accuracy only make the problem worse. This work addresses the utility of OVOD in ISR missions by focusing only what has changed between successive runs through an environment. A Vision Language Model (VLM) compares current observations against a registered “cleared” baseline to focus only on what has changed. Testing across three distinct environments, and using either monocular camera phones or RGB-D equipped vehicles, demonstrates that integrating change detection can automatically remove as much as 80% of unchanged objects without impacting recall.
Meta TagsDetails
DOI
https://doi.org/10.4271/2026-01-7508
Citation
Martinson, E. and Fishta, I., "Change Detection Filtering with Visual Language Models," 2026 NDIA Michigan Chapter Ground Vehicle Systems Engineering and Technology Symposium, Novi, Michigan, United States, August 11, 2026, https://doi.org/10.4271/2026-01-7508.
Additional Details
Publisher
Published
Sep 22
Product Code
2026-01-7508
Content Type
Technical Paper
Language
English