Accurate vehicle trajectory prediction remains critical for autonomous driving
systems. However, accurately representing interaction-rich and time-varying
behaviors is still challenging for many existing predictors, often resulting in
reduced predictive accuracy. To address these limitations, we propose a
kinematics-constrained Transformer with spatiotemporal repulsive force. First,
we model interactions between the target vehicle and its surroundings using a
spatiotemporal repulsive force model, enhancing the network’s environmental
perception. Next, a Transformer encoder extracts temporal features from the
observed motion sequence. A ST-RF attention module captures interaction
dynamics, while an adaptive gating mechanism fuses these cues with the encoded
features. The Transformer decoder outputs a sequence of control commands, which
are integrated through a vehicle kinematics network layers to generate
continuous, physically grounded trajectories. Furthermore, we introduce a
multi-objective physical constraint loss function enforcing kinematic and
dynamic constraints across all predicted agents. Extensive evaluations on the
HighD and NGSIM datasets demonstrate that our model achieves significant
reductions in RMSE across all prediction horizons, with average decreases of
1.63% and 7.69%, respectively. Additionally, in diverse challenging traffic
scenarios, our approach exhibits exceptional robustness, producing trajectories
that are both physically plausible and readily interpretable.