Tech & Data

Autonomous Driving Open Datasets in 2026

Self-driving software learns from examples: millions of camera frames, lidar scans and recorded trips in which humans have marked every car, pedestrian and lane line. Collecting and labeling that data costs a fortune, so carmakers, robotaxi companies and universities publish some of it for free. Researchers get a shared test to compete on, and the publishers get better ideas and new hires in return. This is an unranked list of 22 real, downloadable datasets, grouped by what they're mainly used for. Each row shows the publisher's home country, the year it came out, its sensors, a scale fact and its license. Most are free for research only, so check the license column before building a product on one.

The KITTI Vision Benchmark Suite, recorded around Karlsruhe, Germany

Pictured: the KITTI benchmark from KIT and Toyota Technological Institute at Chicago, from the perception segment.

Dataset Publisher Segment Country Released Sensors Notable scale License
KITTIKarlsruhe Institute of Technology & Toyota Technological Institute at ChicagoPerceptionGermany2012Camera, lidar, GPS/IMUAbout 6 hours of driving in and around KarlsruheCC BY-NC-SA 3.0
nuScenesMotional (nuTonomy)PerceptionUSA2019Camera, lidar, radar1,000 scenes; 1.4M camera images; 1.4M 3D boxesCC BY-NC-SA 4.0
Waymo Open Dataset: PerceptionWaymoPerceptionUSA2019Camera, lidar2,030 twenty-second segmentsNon-commercial (Waymo license)
Argoverse 2: SensorArgo AIPerceptionUSA2021Camera, lidar1,000 3D-annotated scenariosCC BY-NC-SA 4.0
A2D2AudiPerceptionGermany2020Camera, lidar41,277 labeled frames from 6 cameras and 5 lidarsCC BY-ND 4.0 (commercial use allowed)
PandaSetHesai & Scale AIPerceptionChina / USA2020Camera, lidar103 scenes; 48,000 camera images; 16,000 lidar sweepsCC BY 4.0 (commercial use allowed)
ONCEHuaweiPerceptionChina2021Camera, lidar1M lidar scenes; 7M camera images; 144 driving hoursCC BY-NC-SA 4.0
Zenseact Open Dataset (ZOD)ZenseactPerceptionSweden2023Camera, lidar100,000 frames from 14 European countriesCC BY-SA 4.0 (commercial use allowed)
MAN TruckScenesMAN Truck & Bus & TU MunichPerceptionGermany2024Camera, lidar, radar747 truck scenes; first with 360° 4D radarCC BY-NC-SA 4.0
Waymo Open Motion DatasetWaymoMotion & planningUSA2021Object tracks, 3D maps103,354 segments; 570+ hoursNon-commercial (Waymo license)
Argoverse 2: Motion ForecastingArgo AIMotion & planningUSA2021Object tracks, HD maps250,000 eleven-second scenarios from 6 U.S. citiesCC BY-NC-SA 4.0
nuPlanMotionalMotion & planningUSA2021Camera, lidar, object tracks, HD mapsAbout 1,200 hours of driving in 4 citiesCC BY-NC-SA 4.0 (modified)
Lyft Level 5 PredictionLyft Level 5 (now Woven by Toyota)Motion & planningUSA2020Object tracks, HD map1,000+ hours; 170,000 twenty-five-second scenesCC BY-NC-SA 4.0
CityscapesDaimler, TU Darmstadt, MPI Informatics & TU DresdenSegmentationGermany2016Stereo camera5,000 finely + 20,000 coarsely labeled images; 50 citiesCustom, non-commercial
Mapillary VistasMapillarySegmentationSweden2017Camera (crowdsourced)25,000 street images from 6 continentsCC BY-NC-SA (research edition)
ApolloScapeBaiduSegmentationChina2018Camera, lidar140,000+ images with per-pixel labelsCustom, non-commercial
BDD100KUC Berkeley (Berkeley DeepDrive)Video & end-to-endUSA2018Dashcam video, GPS/IMU100,000 forty-second videos; 1,000+ hoursBSD 3-Clause
comma2k19comma.aiVideo & end-to-endUSA2018Camera, GNSS, IMU, CAN bus33+ hours of commuting on California's Highway 280MIT
OpenDV-YouTubeOpenDriveLab (Shanghai AI Lab)Video & end-to-endChina2024Camera (YouTube videos)1,700+ hours of driving videoVideo list (YouTube terms)
DriveLMOpenDriveLab (Shanghai AI Lab)Video & end-to-endChina2023Camera (nuScenes + CARLA simulator)Question-and-answer driving reasoning on nuScenes and CARLACC BY-NC-SA 4.0
Waymo Open Dataset: End-to-End DrivingWaymoVideo & end-to-endUSA2025Camera (8 views), routing4,021 rare “long-tail” driving segmentsNon-commercial (Waymo license)
PhysicalAI-Autonomous-VehiclesNVIDIAVideo & end-to-endUSA2025Camera, lidar, radar1,700 hours; 306,152 clips from 25 countriesNVIDIA license (commercial use allowed)

Highlights

Perception

KITTI

This is where it started. In 2012, researchers from Germany's Karlsruhe Institute of Technology and the Toyota Technological Institute at Chicago fitted a Volkswagen station wagon with stereo cameras, a Velodyne laser scanner and a precise GPS/IMU unit. They then recorded about six hours of traffic around Karlsruhe, on highways, country roads and city streets. KITTI's benchmarks for stereo vision, odometry, object detection and tracking were the standard tests for self-driving research through the 2010s.[1]

Segmentation

Cityscapes

Daimler's research arm built Cityscapes with TU Darmstadt, the Max Planck Institute for Informatics and TU Dresden, photographing street scenes in 50 cities, mostly in Germany. In 5,000 of its images every pixel is labeled as road, sidewalk, pedestrian, car, sky and so on, and another 20,000 have rougher outlines. That made it the standard test for teaching a car to understand everything in a street scene. Research use is free, but commercial use needs a separate license.[2]

Perception

nuScenes

nuTonomy, now part of Motional, released the full nuScenes in March 2019. It was the first public dataset recorded with a full self-driving car's sensors: six cameras, five radars and a lidar, all seeing 360 degrees. Its 1,000 twenty-second scenes from Boston and Singapore, two cities with dense and tricky traffic, carry about 1.4 million 3D boxes drawn around vehicles, people and other objects in 23 classes.[3]

Perception

Waymo Open Dataset

Alphabet's robotaxi company first shared its data in 2019. The Perception set now has 2,030 twenty-second segments of high-resolution camera and lidar data. In March 2021 Waymo added the Motion set: more than 100,000 segments (over 570 hours) showing how vehicles, pedestrians and cyclists actually moved in six U.S. cities, with matching 3D maps. It quickly became a popular benchmark for predicting what other road users will do next. An End-to-End Driving set of rare, tricky situations followed in 2025.[4]

Motion & planning

Argoverse 2

Argo AI, the self-driving startup backed by Ford and Volkswagen, published Argoverse 2 in 2021 as three datasets. One holds 1,000 fully annotated sensor scenarios and another 20,000 unlabeled lidar sequences. The third has 250,000 eleven-second motion-forecasting scenarios from Austin, Detroit, Miami, Palo Alto, Pittsburgh and Washington, D.C. Argo AI shut down in 2022, but its data is still online and widely used.[5]

Perception

Zenseact Open Dataset (ZOD)

Most driving datasets are for research only. Zenseact, a Swedish self-driving software company, released ZOD in 2023 under the CC BY-SA 4.0 license, which allows commercial use. Its 100,000 selected frames were collected over two years in 14 European countries, from snowy northern Sweden to sunny Italy. It also includes 1,473 twenty-second sequences and 29 longer drives with the full set of sensors.[6]

Video & end-to-end

BDD100K

UC Berkeley's Berkeley DeepDrive lab chose variety over a pricey sensor rig. BDD100K holds 100,000 forty-second dashcam videos, more than 1,000 hours in all, from over 50,000 rides in New York, the San Francisco Bay Area and other regions. The videos cover all kinds of weather and every time of day. First released in 2018, it now has labels for many tasks, from lane markings to the parts of the road a car can drive on.[7]

Video & end-to-end

NVIDIA PhysicalAI-Autonomous-Vehicles

This is the biggest recent addition. In late 2025 NVIDIA released 306,152 twenty-second clips, 1,700 hours in all, recorded in 25 countries and more than 2,500 cities. Every clip has seven-camera coverage, and most also include lidar, with radar on about half. The release takes up about 133 TB and came alongside NVIDIA's Alpamayo models for self-driving cars that can reason through a situation. It's offered for commercial as well as research use under NVIDIA's own license.[8]

Self-driving Datasets Technology

Sources

  1. “Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite” (CVPR 2012) and “Vision meets Robotics: The KITTI Dataset” (International Journal of Robotics Research, 2013); KITTI license terms (cvlibs.net).
  2. “The Cityscapes Dataset for Semantic Urban Scene Understanding” (CVPR 2016); Cityscapes terms and conditions (cityscapes-dataset.com).
  3. “nuScenes: A multimodal dataset for autonomous driving” (CVPR 2020); Aptiv's nuScenes release article; nuScenes terms of use.
  4. Waymo Open Dataset site (About, FAQ and Terms pages, waymo.com/open); The Robot Report coverage of the 2019 launch and the 2021 Motion dataset; “Large Scale Interactive Motion Forecasting for Autonomous Driving: The Waymo Open Motion Dataset” (ICCV 2021).
  5. “Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting” (NeurIPS 2021 Datasets and Benchmarks); argoverse.org.
  6. “Zenseact Open Dataset: A Large-Scale and Diverse Multimodal Dataset for Autonomous Driving” (ICCV 2023); zod.zenseact.com and the ZOD development kit on GitHub.
  7. “BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning” (CVPR 2020); Synced coverage of the 2018 release; the BDD100K GitHub repository (BSD 3-Clause license).
  8. NVIDIA's PhysicalAI-Autonomous-Vehicles dataset card on Hugging Face; NVIDIA's Alpamayo developer blog; coverage of NVIDIA GTC Washington, D.C. (October 2025).
  9. Table-only perception entries: “A2D2: Audi Autonomous Driving Dataset” (2020) and its AWS Open Data listing; TechCrunch on Scale AI and Hesai's PandaSet release (May 2020); “One Million Scenes for Autonomous Driving: ONCE Dataset” (NeurIPS 2021 Datasets and Benchmarks); “MAN TruckScenes: A multimodal dataset for autonomous trucking in diverse conditions” (NeurIPS 2024) and its AWS Open Data listing.
  10. Table-only motion and planning entries: Motional's nuPlan announcements, “nuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles” (2021) and the nuplan-devkit documentation; “One Thousand and One Hours: Self-driving Motion Prediction Dataset” (2020) and Lyft Level 5's prediction dataset announcement.
  11. Table-only segmentation and video entries: “The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes” (ICCV 2017) and mapillary.com; “The ApolloScape Dataset for Autonomous Driving” (CVPR 2018 workshops) and the ApolloScape user license; “A Commute in Data: The comma2k19 Dataset” (2018); “Generalized Predictive Model for Autonomous Driving” (GenAD/OpenDV-YouTube, CVPR 2024) and the OpenDriveLab DriveAGI repository; “DriveLM: Driving with Graph Visual Question Answering” (ECCV 2024); “WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios” (2025).