CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
Abstract
Cellular vehicle-to-everything (C-V2X) enables cooperativeperception, prediction, and planning beyond the field of view of individualagents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dy-namics, heterogeneous sensing modalities, and scalability beyond a singlecooperative partner. In this paper, we introduce CooperScene, a high-fidelity cooperative autonomy dataset with real-world C-V2X communica-tion characterization. The dataset is organized into diverse scenes, includ-ing intersections, highway ramps, and parking lots. These scenes involvethree connected and autonomous vehicles (CAVs) and one infrastructureroadside unit (RSU), all equipped with multi-modal sensors and commer-cial off-the-shelf C-V2X communication radios. All scenes are annotatedwith globally consistent 3D labels at 10 Hz, totaling 344K objects across59K frames, underpinned by tight sensor- and agent-synchronization,centimeter-level localization and spatial alignment, precise cross-modalitycalibration, and 3GPP-standard-compliant C-V2X communication. Coop-erScene establishes a rigorous benchmark for evaluating multi-agentscaling and actual performance in real-world deployable settings. Projectwebsite for data and benchmark: https://cisl.ucr.edu/CooperScene.