ScopeDrive: Text-Anchored Cross-Modal Calibration and Density-Aware Modulation for Autonomous Driving

Journal Publication ResearchOnline@JCU
Hou, Minghui;Wang, Gang;Guan, Runwei;Liu, Jianan;Huang, Tao;Han, Qing Long
Abstract

Vision–language models are emerging as a unified paradigm for perception, prediction, and planning in autonomous driving. However, most existing approaches still rely on shallow fusion between visual and textual features, leading to weak cross-modal interaction, feature drift, and poor adaptability in complex scenes. This work present ScopeDrive, an end-to-end VLM framework that enables deep semantic alignment and adaptive reasoning through two novel components. The text-anchored calibrator transforms textual semantics into multiscale anchors that progressively calibrate visual features across layers, strengthening cross-modal correspondence. The density-aware agent modulator estimates scene complexity and dynamically adjusts attention distribution, allowing the model to focus on dense or dynamic regions when needed. Evaluated on the DriveLM, DriveBench, and NuScenes-QA benchmarks, ScopeDrive surpasses both lightweight and large-scale baselines, achieving a BLEU-4 of 53.27 and METEOR of 38.75 on DriveLM while maintaining only 328 M parameters. It also delivers state-of-the-art performance on perception and planning tasks under both clean and corrupted conditions. These results demonstrate that ScopeDrive effectively breaks the shallow-fusion barrier, offering a lightweight yet semantically aligned foundation for interpretable autonomous-driving intelligence.

Journal

IEEE transactions on industrial informatics

Publication Name

IEEE Transactions on Industrial Informatics

Volume

22

ISBN/ISSN

1941-0050

Edition

N/A

Issue

8

Pages Count

12

Location

N/A

Publisher

IEEE

Publisher Url

N/A

Publisher Location

N/A

Publish Date

N/A

Url

N/A

Date

N/A

EISSN

N/A

DOI

10.1109/TII.2026.3683432