Scene-Level Multi-View Instance Consistency Using Multi-Feature Vision Fusion for Synthetic Ground Vehicle Environments

2026-01-7502

9/22/2026

Authors
Abstract
Content
Maintaining consistent object identities across multiple camera viewpoints is a critical challenge in synthetic perception environments used for autonomous ground vehicle evaluation. This paper presents a scene-level multi-view instance consistency framework that integrates OpenUSD scene composition, Omniverse Replicator synthetic-data generation, and a multi-feature vision fusion pipeline. The proposed approach combines semantic embeddings from CLIP, patch-level descriptors from DINOv2, geometric correspondences from LoFTR, mask-derived shape invariants using Hu moments, and relative-position priors to associate object instances across views, including visually identical objects. A compact composite scoring function fuses these complementary cues to achieve robust cross-view identity assignment while preserving OpenUSD asset modularity through grouped-prim support. Synthetic experiments across 120 multi-camera scenes demonstrate improved Top-1 Match Accuracy and Identity Consistency Rate, with reduced ID-switch occurrences compared to single-cue baselines. The framework supports scalable, repeatable, and traceable digital engineering workflows for defense-oriented perception evaluation.
Meta TagsDetails
DOI
https://doi.org/10.4271/2026-01-7502
Citation
Bhattacharya, S. and Nakamoto, K., "Scene-Level Multi-View Instance Consistency Using Multi-Feature Vision Fusion for Synthetic Ground Vehicle Environments," 2026 NDIA Michigan Chapter Ground Vehicle Systems Engineering and Technology Symposium, Novi, Michigan, United States, August 11, 2026, https://doi.org/10.4271/2026-01-7502.
Additional Details
Publisher
Published
Sep 22
Product Code
2026-01-7502
Content Type
Technical Paper
Language
English