View Translation: Geometry-Guided Latent Diffusion for Cross-View Image Synthesis on Ground Vehicles
2026-01-7501
9/22/2026
- Content
- Synthesizing novel camera views is important for autonomous ground vehicles, with applications in surround-view monitoring, occlusion recovery, and training data augmentation. We present View Translation, a geometry-guided latent diffusion framework that generates a target camera view from a source image, relative camera pose, and an available target-view depth prior. The method combines three components: a Vector Quantized Variational Autoencoder for compact latent encoding, a depth-based warping module that projects the source image into the target view to provide geometric guidance, and a ControlNet-augmented denoising UNet conditioned on source appearance, relative pose, and an auxiliary Image-Depth fusion network. Evaluated on KITTI and a simulated off-road dataset, our method achieves competitive FID while improving LPIPS and PSNR over baseline approaches, supporting cross-view synthesis for ground vehicle perception.
- Citation
- Mayekar, O., Aiyetigbo, M., Salvi, A., Samak, T., et al., "View Translation: Geometry-Guided Latent Diffusion for Cross-View Image Synthesis on Ground Vehicles," 2026 NDIA Michigan Chapter Ground Vehicle Systems Engineering and Technology Symposium, Novi, Michigan, United States, August 11, 2026, https://doi.org/10.4271/2026-01-7501.