ScopeDrive: Text-Anchored Cross-Modal Calibration and Density-Aware Modulation for Autonomous Driving
Journal Publication ResearchOnline@JCUVision–language models are emerging as a unified paradigm for perception, prediction, and planning in autonomous driving. However, most existing approaches still rely on shallow fusion between visual and textual features, leading to weak cross-modal interaction, feature drift, and poor adaptability in complex scenes. This work present ScopeDrive, an end-to-end VLM framework that enables deep semantic alignment and adaptive reasoning through two novel components. The text-anchored calibrator transforms textual semantics into multiscale anchors that progressively calibrate visual features across layers, strengthening cross-modal correspondence. The density-aware agent modulator estimates scene complexity and dynamically adjusts attention distribution, allowing the model to focus on dense or dynamic regions when needed. Evaluated on the DriveLM, DriveBench, and NuScenes-QA benchmarks, ScopeDrive surpasses both lightweight and large-scale baselines, achieving a BLEU-4 of 53.27 and METEOR of 38.75 on DriveLM while maintaining only 328 M parameters. It also delivers state-of-the-art performance on perception and planning tasks under both clean and corrupted conditions. These results demonstrate that ScopeDrive effectively breaks the shallow-fusion barrier, offering a lightweight yet semantically aligned foundation for interpretable autonomous-driving intelligence.
IEEE transactions on industrial informatics
IEEE Transactions on Industrial Informatics
22
1941-0050
N/A
8
12
N/A
IEEE
N/A
N/A
N/A
N/A
N/A
N/A
10.1109/TII.2026.3683432
