Source-linked AI summary

MIT Advanced Vehicle Technology Study: Large-Scale Naturalistic Driving Study of Driver Behavior and Interaction with Automation

Lex Fridman, Daniel E. Brown, Michael Glazer, William Angell, Spencer Dodd, Benedikt Jenik, Jack Terwilliger, Aleksandr Patsekin, Julia Kindelsberger, Li Ding, Sean Seaman, Alea Mehler, Andrew Sipperley, Anthony Pettinato, Bobbie Seppelt, Linda Angell, Bruce Mehler, Bryan Reimer

arXiv:1711.06976v4cs.CYcs.CVcs.HC

TL;DR

Human drivers are expected to remain part of the driving task, creating a need to understand shared autonomy in real-world conditions. MIT-AVT addresses this need through large-scale naturalistic data collection and automated analysis of human, vehicle, and environmental behavior. The resulting dataset currently includes 122 participants, 15,610 days, 511,638 miles, and 7.1 billion video frames.

  • Problem

    Shared autonomy requires understanding complex human-vehicle interaction across driver variability, perception edge cases, and environmental conditions.

  • Method

    The study combines long- and medium-term naturalistic driving data with instrumented vehicles, multimodal sensing, distributed processing, computer vision, and deep learning.

  • Results

    The dataset includes 122 participants, 15,610 days of participation, 511,638 miles, and 7.1 billion video frames.

  • Takeaways & Limitations

    MIT-AVT aims to provide real-world insight for designing shared-autonomy systems and informing vehicle, insurance, and government decisions about automation use.

Abstract

from arXiv · show

For the foreseeble future, human beings will likely remain an integral part of the driving task, monitoring the AI system as it performs anywhere from just over 0% to just under 100% of the driving. The governing objectives of the MIT Autonomous Vehicle Technology (MIT-AVT) study are to (1) undertake large-scale real-world driving data collection that includes high-definition video to fuel the development of deep learning based internal and external perception systems, (2) gain a holistic understanding of how human beings interact with vehicle automation technology by integrating video data with vehicle state data, driver characteristics, mental models, and self-reported experiences with technology, and (3) identify how technology and other factors related to automation adoption and use can be improved in ways that save lives. In pursuing these objectives, we have instrumented 23 Tesla Model S and Model X vehicles, 2 Volvo S90 vehicles, 2 Range Rover Evoque, and 2 Cadillac CT6 vehicles for both long-term (over a year per driver) and medium term (one month per driver) naturalistic driving data collection. Furthermore, we are continually developing new methods for analysis of the massive-scale dataset collected from the instrumented vehicle fleet. The recorded data streams include IMU, GPS, CAN messages, and high-definition video streams of the driver face, the driver cabin, the forward roadway, and the instrument cluster (on select vehicles). The study is on-going and growing. To date, we have 122 participants, 15,610 days of participation, 511,638 miles, and 7.1 billion video frames. This paper presents the design of the study, the data collection hardware, the processing of the data, and the computer vision algorithms currently being used to extract actionable knowledge from the data.

I. INTRODUCTION

The MIT-AVT study addresses the variability and complexity of human-in-the-loop driving by collecting large-scale naturalistic data on drivers, vehicles, environments, and automation use. It combines long-duration real-world observation with automated analysis to understand shared autonomy and support safer vehicle-system design.

  • Motivation: Human behavior remains integral to autonomous driving because driver variability, social interactions, and the need to retake control complicate shared autonomy.The study considers driver styles, experience, trust, human-machine interaction, and system failures requiring human intervention.
  • Motivation: Environmental conditions, scene-perception edge cases, sensor imperfections, software limitations, and underactuated control all constrain autonomous driving systems.These challenges affect perception, control, interaction dynamics, and the reliability of human-in-the-loop vehicle systems.
  • Study scope: The MIT-AVT study collects unconstrained, long-term video, audio, vehicle telemetry, and sensor data to characterize real-world interaction with autonomous driving technology.Naturalistic driving studies aim to capture behavior with minimal influence from structured experimental tasks or experimenters in the vehicle.
  • Study scope: The study spans automation from emergency braking to continuous semi-autonomous control, enabling analysis of driver behavior across levels of vehicle assistance.Its autonomy-at-all-levels principle covers systems that sense the environment or cabin and control the vehicle or communicate with the driver.
  • Analysis approach: The MIT-AVT pipeline extends analysis beyond crash epochs and manual annotation by processing billions of high-definition video frames with computer vision and deep learning.Traditional annotation and expert review remain useful, but the study also targets long-tail behavior such as glance allocation over extended autonomous driving.
  • Study scale: 122 drivers, 29 vehicles, 15,610 participant days, 511,638 miles, and 7.1 billion video frames define the study’s current scale.The reported measures cover the active study, recorded participation, fleet size, mileage, and processed video data.

A. Naturalistic Driving Studies

MIT-AVT extends earlier naturalistic driving studies beyond crash-focused epochs by analyzing entire trips and the long tail of human interaction with automation using automated methods.

  • Earlier naturalistic driving studies focused on crash and near-crash epochs detected through vehicle kinematics.
  • MIT-AVT targets large-scale computer-vision analysis of human behavior rather than relying on manual annotation of selected driving epochs.
  • The study analyzes entire trips and automation strategies, including when systems are activated, deactivated, and exchanged with human control.

B. Datasets for Application of Deep Learning

The paper situates MIT-AVT within deep-learning research that uses large annotated datasets to extract driver, scene, and vehicle-state information from naturalistic driving data.

  • Deep learning uses multilayer neural networks or hierarchical representation learning with minimal human specification of representations.
  • Learning-based driving systems benefit from large-scale data collection and annotation because models must generalize across real-world operating edge cases.
  • Large-scale annotated datasets are required to train deep neural networks for extracting human behavior from raw video.
  • Existing datasets support object detection, 3D driving perception, semantic scene understanding, and frame-wise video segmentation.
  • Automotive applications include face and gaze analysis, body-pose estimation, semantic scene perception, and driving-state prediction.

II. MIT-AVT STUDY STRUCTURE AND GOALS

MIT-AVT’s study structure combines continual hardware and software innovation with long- and medium-duration naturalistic driving protocols and standardized participant preparation.

  • The study’s governing principle is continual innovation while maintaining backward compatibility across hardware, software, and data processing.
  • The long-duration study uses subject-owned vehicles for over one year, whereas the medium-duration study uses MIT-owned vehicles for approximately one month.
  • Medium-duration participants must meet commuting, consent, background-check, driving-record, and training requirements before joining the MIT-owned fleet.
  • Participants receive vehicle and feature instruction before using systems including Adaptive Cruise Control, Pilot Assist, Super Cruise, and safety-assistance functions.
  • On-road training provides at least 30 minutes of highway exposure to the systems in a real-world setting.

C. Qualitative Approaches for One Month NDS

The one-month NDS combines unobtrusive self-reporting with instrumented, synchronized multimodal data collection to study driver experience and vehicle-system behavior.

  • Self-report methods capture participants’ vehicle experiences, technology perceptions, and barriers to adopting or discarding automation.
  • Questionnaires collect demographics, driving history, technology exposure, trust, and post-training perceptions across three stages.
  • A 30–60 minute end-of-study interview examines initial reactions, training effects, technology experiences, and driver perceptions.
  • RIDER records synchronized multimodal driving data, including high-definition video and vehicle sensors, as the study’s hardware and software backbone.
  • RIDER is installed in the trunk with secured hidden cabling, while its power system is designed for vehicle transfer and low battery drain.
  • The Knights of CANelot monitors CAN traffic and activates RIDER when a predefined vehicle signal indicates that the car is on.

B. Computing Platform and Sensors

The computing platform centers on a Banana Pi Pro with integrated vehicle, motion, positioning, and timing interfaces, supported by storage, communications, and CAN-controlled power hardware.

  • Computing Platform: A Banana Pi Pro provides the RIDER computing platform with expandable GPIO for IMU, GPS, and CAN integration.It includes a 1GHz ARM Cortex-A7 processor, 1GB of RAM, and a professionally manufactured sensor-integration daughter board.
  • Sensors and Interfaces: The platform includes onboard CAN control and a dedicated CAN transceiver for vehicle telemetry collection.The listed components include an ARM onboard CAN controller and a Texas Instruments SN65HVD230 CAN transceiver.
  • Sensors and Interfaces: A DS3231 real-time clock provides accurate timekeeping and timestamping with ±2 ppm accuracy.
  • Sensors and Interfaces: A nine-degree-of-freedom inertial measurement unit combines gyroscope, accelerometer, and compass sensing.The IMU uses the STMicro L3GD20H gyroscope and LSM303D accelerometer/compass.
  • Supporting Hardware: The system adds GPS, 4G LTE, powered USB expansion, and 1TB/2TB solid-state storage.The GPS unit is DGPS-capable and accurate within 5 meters; the USB hub is a powered USB 3.0 four-port hub.

C. Cameras

RIDER uses multiple synchronized cameras and vehicle sensors to capture interior, forward-roadway, and instrument-cluster data, while its software manages recording and startup operation.

  • Cameras: Three or four Logitech C920 webcams record 1280x720 video at 30 frames per second inside and around the vehicle.Two cameras accept CS-type lens mounts for adaptable face or body-pose placement; a standard webcam views the forward road, and an optional fourth records the instrument cluster.
  • Cameras: Camera microphones capture audio, and custom mounts support specialized placement within the vehicle.
  • Cameras: On-camera compression enables a RIDER installation to support up to 6 cameras despite limited single-board-computer processing capacity.The Logitech C920 off-loads compression from the compute platform.
  • Startup Scripts: The MIT-AVT software framework processes timestamped sensory data through cleaning, synchronization, annotation, interpretation, knowledge extraction, and aggregate analysis.The framework distributes processing across thousands of GPU-enabled compute cores.
  • Startup Scripts: RIDER software powers on with the vehicle, creates trip directories, records timestamped streams, transmits metadata, and powers down after the vehicle turns off.
  • Startup Scripts: A CAN-triggered microcontroller connects vehicle power to RIDER when a specified signal indicates startup and disconnects it when that signal ends.
  • Startup Scripts: Boot scripts synchronize the onboard clock, load communication modules, monitor shutdown and GSM connectivity, and start Dacman and Lighthouse.

D. Dacman

Dacman coordinates RIDER trip recording across cameras, CAN, audio, and metadata, while supporting programs create timestamped files and remote system monitoring.

  • Dacman: Dacman centrally manages RIDER data streams using device, RIDER, subject, vehicle, and study identifiers from its configuration.
  • Dacman: Each trip receives a dated, uniquely named directory containing configuration, subsystem CSV files, and specification metadata.The naming convention is rider-id_date_timestamp.
  • Dacman: Subsystem manager scripts invoke system calls for recording, while Dacman and companion C programs write timestamped CSV, RAW, and H264 data.Camera and audio frames receive accompanying microsecond timestamps.
  • Recording Programs: Cam2hd records camera streams through Video4Linux at 720p and writes raw H264 frames.
  • Recording Programs: Dump_can receives CAN data into CSV and supports listen-only CAN-controller operation for enhanced security.
  • Remote Monitoring: Lighthouse sends encrypted trip timing, GPS, power, temperature, and storage information to Homebase for remote monitoring.Homebase decrypts and stores the information, while Heartbeat displays RIDER system status and helps identify maintenance needs.

J. RIDER Database

The RIDER database stores trip, participant, vehicle, instrumentation, epoch, and system-health information, supporting processing from offload through synchronization.

  • J. RIDER Database: PostgreSQL stores incoming trip information and records trips offloaded to the storage server for later processing and querying.Processed trip information can be queried by trip, time, event, or condition.
  • J. RIDER Database: The database links instrumentation dates, vehicle IDs, subject and study IDs, RIDER identifiers, and vehicle attributes.
  • J. RIDER Database: Trip records include synchronization state, available camera and subsystem data, and metadata such as GPS frequency and technology use.
  • J. RIDER Database: Epoch tables identify trips and video-frame ranges associated with events such as Tesla Autopilot use.
  • J. RIDER Database: Homebase logs preserve streamed RIDER system-health and state information.
  • J. RIDER Database: Raw trips are inspected for inconsistencies, repaired when possible, and removed when failures such as an unplugged camera make recovery impossible.
  • J. RIDER Database: Camera timestamps are aligned into a 30-frame-per-second synchronized video, with repeated frames used when low light reduces recording to 15 frames per second.
  • J. RIDER Database: Decoded CAN values are aligned to synchronized frame IDs by timestamp, enabling combined visualizations of video streams and CAN information.

A. Trip Configuration Files

Trip configuration files organize subject, subsystem, diagnostic, and timing information, while trip data files contain synchronized-ready sensor, video, and audio recordings.

  • Configuration files: Trip dacman.json stores subject and systems information used to record each trip.
  • Configuration files: Trip diagnostics.log records environmental, hardware-temperature, power-usage, and free-disk-space diagnostics.
  • Configuration files: Trip specs.json records start and end timestamps for all subsystems.
  • Trip data files: Trip data files contain timestamped CSV streams alongside raw H264 video and RAW audio files.
  • Trip data files: Camera directories contain camera-specific H264, error, and frame-to-system-timestamp CSV files, while separate files store CAN, GPS, IMU, audio, and error data.

C. Cleaning Criteria

Cleaning preserves trips with recoverable file or metadata problems and removes trips that violate data-quality, consent, duration, movement, or participation criteria.

  • Recoverable errors: Recoverable errors are repaired through permission correction, backups, timestamp reconstruction, identifier correction, removal of failed nonessential files, or deletion of incomplete CSV lines.
  • Removal criteria: Trips are removed for nonconsenting drivers, requested removal, no vehicle movement, files smaller than 15MB, or duration shorter than 30 seconds.
  • Removal criteria: Trips are removed when essential camera, configuration, or CAN files are missing or when participation falls outside the volunteer range.
  • Removal criteria: Large error files for essential cameras or CAN data also trigger trip removal.
  • Removal criteria: Trips are removed when subsystem timestamps mismatch by at least one minute.

E. Synchronized Files

Synchronization aligns video and vehicle-state data at 30 frames per second, producing linked files for downstream analysis; the platform’s hardware limits expansion and onboard preprocessing.

  • Synchronization: 30 frames per second is the synchronization rate for aligning video frames and CAN messages.
  • Synchronized outputs: synced video.csv records each camera’s video-frame ID and timestamp at 30 frames per second.
  • Synchronized outputs: synced video camera-name.mp4 stores camera videos synchronized across streams at 30 FPS using H264 encoding.
  • Synchronized outputs: synced can.csv associates each synchronized video frame with the closest CAN values for every CAN message.
  • Hardware limitations: RIDER’s limited system memory can hinder expansion, while its dual-core ARM processor can fluctuate when preprocessing data onboard.
  • Study contribution: The study applies embedded systems, distributed computing, computer vision, and deep learning to analyze large-scale naturalistic driving data.
Loading 1711.06976v4…