Source-linked AI summary
Human-Like Decision Making for Autonomous Driving: A Noncooperative Game Theoretic Approach
Peng Hang, Chen Lv, Yang Xing, Chao Huang, Zhongxu Hu
TL;DR
The paper addresses how AVs can coexist with human drivers while offering personalized driving styles for passengers. It proposes a human-like framework combining driving-style modeling, noncooperative games, and potential-field MPC, and reports reasonable decisions in lane-change tests, with Stackelberg costs reduced by over 20% under normal driving.
Problem
AVs must coexist with human-driven vehicles and accommodate passengers’ different priorities regarding safety, comfort, and travel efficiency.
Method
The framework models aggressive, normal, and conservative styles, applies Nash and Stackelberg games to AV interactions, and combines potential fields with MPC for motion prediction and collision-avoidance planning.
Results
Both game-theoretic approaches provide reasonable human-like decisions, while Stackelberg costs under normal driving are reduced by 22%, 20%, and 22% for safety, comfort, and efficiency, respectively.
Takeaways & Limitations
The framework supports human-like AV decisions across different driving styles and reports better performance for Stackelberg equilibrium than Nash equilibrium in safety, comfort, and efficiency.
Takeaways & Limitations
The method has limited ability to handle varied complex scenarios, and increasing system complexity requires more computation resources and execution capability.
Abstract
from arXiv · showhide
Considering that human-driven vehicles and autonomous vehicles (AVs) will coexist on roads in the future for a long time, how to merge AVs into human drivers traffic ecology and minimize the effect of AVs and their misfit with human drivers, are issues worthy of consideration. Moreover, different passengers have different needs for AVs, thus, how to provide personalized choices for different passengers is another issue for AVs. Therefore, a human-like decision making framework is designed for AVs in this paper. Different driving styles and social interaction characteristics are formulated for AVs regarding driving safety, ride comfort and travel efficiency, which are considered in the modeling process of decision making. Then, Nash equilibrium and Stackelberg game theory are applied to the noncooperative decision making. In addition, potential field method and model predictive control (MPC) are combined to deal with the motion prediction and planning for AVs, which provides predicted motion information for the decision-making module. Finally, two typical testing scenarios of lane change, i.e., merging and overtaking, are carried out to evaluate the feasibility and effectiveness of the proposed decision-making framework considering different human-like behaviors. Testing results indicate that both the two game theoretic approaches can provide reasonable human-like decision making for AVs. Compared with the Nash equilibrium approach, under the normal driving style, the cost value of decision making using the Stackelberg game theoretic approach is reduced by over 20%.
I. INTRODUCTION
The paper addresses how AVs can behave compatibly with human drivers while accommodating passengers’ differing priorities. It proposes human-like driving styles and game-theoretic decision making, evaluated through simulated lane-change scenarios.
- AV decision making affects driving safety, ride comfort, travel efficiency, and energy consumption while bridging environment perception and motion control.
- Human-like AV behavior is motivated by prolonged coexistence with human-driven vehicles and the need for more predictable interaction in complex traffic.
- Prior game-theoretic studies address complex vehicle interactions, but studies incorporating different driving styles were rarely reported.
- Different passengers may prioritize safety and comfort or travel efficiency, motivating personalized driving-style options.
- The paper defines driving styles, formulates AV interactions with Nash and Stackelberg games, and validates the algorithms through simulations in various scenarios.
II. HUMAN-LIKE DECISION-MAKING FRAMEWORK
The framework embeds aggressive, normal, and conservative human-driving characteristics into AV modeling and decision making. Driving data, driver modeling, game theory, and motion prediction are combined to support human-like behavior.
- System Framework for Human-like Decision Making: The framework contains modeling, decision-making, and motion-planning modules that embed predefined driving styles into an integrated driver–vehicle model.
- Decision Making: Nash and Stackelberg games provide alternative noncooperative formulations for human-like AV decision making and interaction.
- Characteristic Analysis of Different Driving Styles: NGSIM I-80 and US-101 freeway data are used to analyze human driving styles across different traffic scenarios.
- Characteristic Analysis of Different Driving Styles: Velocity represents travel efficiency, acceleration represents ride comfort, and time headway represents driving safety.
- Characteristic Analysis of Different Driving Styles: Aggressive drivers tend toward higher velocity, larger acceleration, and smaller time headway, whereas conservative drivers show the opposite pattern; normal drivers lie between them.
- Driver Model: A single-point preview driver model predicts vehicle motion and forms a planned path from consecutive preview points by minimizing predicted-point and preview-point distance.
- Driver Model: The driver model reflects driving style through physical delay time, predicted time, and steering proportional gain.
C. Vehicle-Road Model
The vehicle-road model simplifies a four-wheel vehicle into a two-wheel bicycle model and combines its dynamics with global kinematics and the driver model.
- The four-wheel vehicle model is simplified as a two-wheel bicycle model to reduce controller-design complexity.
- Assuming a small front-wheel steering angle, the model derives simplified vehicle dynamics from longitudinal and lateral tire forces.
- The model defines front and rear tire forces, wheel bases, vehicle mass, yaw inertia, cornering stiffness, and tire slip angles.
- The vehicle parameters are listed in Table II, and the vehicle kinematic model is expressed in global coordinates.
- Combining the driver model with the vehicle-road model yields an integrated model with state vector x and preview-point control input u = Y_p.
IV. DECISION MAKING BASED ON NONCOOPERATIVE GAME THEORY
Lane-change decision making models safety, comfort, and travel efficiency under different driving styles, then applies noncooperative game theory to interactions between the ego and surrounding vehicles.
- A. Cost Function for Lane-change Decision Making: Obstacle-car modeling includes acceleration and deceleration behaviors but excludes lane-change behaviors.The cost function for the adjacent car has a similar expression, with lateral safety equal to the ego car's lateral safety cost.
- A. Cost Function for Lane-change Decision Making: The ego car chooses whether to decelerate behind a slow lead car or change to the left or right lane.The scenario places the ego car on the middle lane of a three-lane highway.
- A. Cost Function for Lane-change Decision Making: Driving safety cost combines longitudinal safety relative to the lead car with lateral safety relative to adjacent cars.Lane-change behavior is represented as changing left, maintaining the lane, or changing right.
- A. Cost Function for Lane-change Decision Making: Ride comfort is associated with longitudinal and lateral acceleration, while travel efficiency is associated with the ego car's longitudinal velocity.The adjacent-car cost assumes no lane change and includes only longitudinal acceleration for ride comfort.
- A. Cost Function for Lane-change Decision Making: The lane-change model considers aggressive, normal, and conservative driving styles through different weighting coefficients for safety, comfort, and travel efficiency.These three performance indexes are integrated into the ego car's decision-making cost function.
B. Noncooperative Decision Making Based on Nash Equilibrium
The Nash formulation treats the ego car and adjacent car as a two-player noncooperative game whose decisions minimize their respective costs subject to vehicle constraints.
- B. Noncooperative Decision Making Based on Nash Equilibrium: The lane-change interaction between the ego car and adjacent car is formulated as a two-player Nash-equilibrium game.The formulation jointly represents the ego car's longitudinal acceleration and lane-change behavior with the adjacent car's longitudinal acceleration.
- B. Noncooperative Decision Making Based on Nash Equilibrium: Nash equilibrium yields optimal longitudinal accelerations for both vehicles and an optimal lane-change behavior for the ego car.The formulation includes cases with adjacent cars on both the left and right lanes.
- B. Noncooperative Decision Making Based on Nash Equilibrium: The Nash decision is bounded by minimum and maximum constraints on longitudinal acceleration and velocity.These constraints apply to the participating vehicles' longitudinal states.
C. Noncooperative Decision Making Based on Stackelberg Equilibrium
The Stackelberg formulation makes the ego car the leader and the adjacent car the follower, so the follower's response affects the ego car's decision.
- C. Noncooperative Decision Making Based on Stackelberg Equilibrium: Unlike Nash equilibrium, Stackelberg equilibrium assigns unequal roles: the ego car leads and the adjacent car follows.In Nash equilibrium, both vehicles are equal and independent players; in Stackelberg equilibrium, the adjacent car's behavior affects the ego car's decision.
- C. Noncooperative Decision Making Based on Stackelberg Equilibrium: The Stackelberg lane-change formulation incorporates the adjacent car's follower behavior into the ego car's decision making.The formulation is extended to scenarios with adjacent cars on both the left and right lanes.
V. MOTION PREDICTION AND PLANNING BASED ON THE POTENTIAL FIELD MODEL METHOD AND MPC
Motion prediction and planning combine a potential field model for obstacles and roads with MPC to predict vehicle states and generate collision-avoidance paths.
- V. MOTION PREDICTION AND PLANNING BASED ON THE POTENTIAL FIELD MODEL METHOD AND MPC: The integrated potential field represents both obstacle cars and roads in the motion-planning model.Obstacle cars include lead vehicles and adjacent cars.
- V. MOTION PREDICTION AND PLANNING BASED ON THE POTENTIAL FIELD MODEL METHOD AND MPC: Combining potential fields with MPC provides predicted motion states and a collision-avoidance path for the decision-making framework.The potential field handles dynamic obstacles and roads while MPC is integrated for prediction and planning.
- V. MOTION PREDICTION AND PLANNING BASED ON THE POTENTIAL FIELD MODEL METHOD AND MPC: The obstacle-car potential field depends on the obstacle's position, heading angle, longitudinal velocity, and shape-related parameters.The model also uses a maximum potential-field value and convergence coefficients.
- V. MOTION PREDICTION AND PLANNING BASED ON THE POTENTIAL FIELD MODEL METHOD AND MPC: The road potential field uses distance to the lane line, a safety threshold, obstacle-car width, and a maximum road potential-field value.These quantities define the road-related potential function.
B. Motion Prediction Based on MPC
The MPC motion-prediction module transforms the integrated model into a time-varying linear system and optimizes constrained control inputs over prediction and control horizons.
- MPC transforms the integrated model into a time-varying linear system for motion prediction.
- The output vector includes potential-field value, lateral distance error, and yaw-angle error relative to the lane centerline.
- Prediction uses defined predictive and control horizons, with N_p ≥ N_c, alongside state, output, and control sequences.
- The motion-prediction cost function is defined for EC using output and control-variation weighting matrices Q and R.
- The optimization constrains both control inputs and their increments within minimum and maximum bounds.
- After optimization, the optimal input-increment sequence is obtained, the control output is calculated, and the state is updated at the next time step.
VI. TESTING RESULTS AND PERFORMANCE EVALUATION
Two MATLAB-Simulink scenarios evaluate human-like decision making for merging, showing that driving styles and game strategies produce substantially different behaviors and performance trade-offs.
- Two testing scenarios evaluate the proposed human-like decision-making framework on the MATLAB-Simulink platform.
- Scenario A: Scenario A is a highway merge requiring EC to react to an AC while lane-ending pressure and driving style affect decisions.
- Scenario A: Aggressive merging minimizes decision time and favors travel efficiency, whereas conservative driving decelerates to preserve safety at the cost of longer travel time.
- Scenario A: Aggressive EC behavior uses sudden large acceleration and an immediate lane change, while normal behavior accelerates progressively and waits for a sufficiently large gap.
- Game-theoretic comparison: 22%, 20% and 22%: under normal driving, SE reduces driving safety, ride comfort and travel efficiency cost values relative to NE.
- Scenario A: Travel Efficiency costs are reported as RMS values, with the table listing 114, 79, 275, 215, 394 and 329 across the displayed conditions.
B. Scenario B
Scenario B tests overtaking on a curved three-lane highway, where EC must choose between following a slow vehicle and changing lanes around surrounding vehicles.
- Scenario B is an overtaking maneuver on a curved three-lane highway involving slow LC and surrounding vehicles AC1 and AC2.
- EC must decide whether to slow down and follow LC or change lanes, and if changing lanes, whether to pass on the left or right.
- Testing results compare different driving styles and decision strategies using decision-making, velocity, and gap results.
- 31%, 34% and 27%: under normal driving, SE reduces driving safety, ride comfort and travel efficiency costs relative to NE.
- Conservative EC prefers deceleration and a larger safe distance, while normal EC takes longer to accelerate and complete overtaking.
C. Discussion of the Testing Results
Across the tests, driving styles create distinct decision patterns, while both game-theoretic approaches yield reasonable decisions and SE performs better on the evaluated driving criteria.
- Aggressive, conservative and normal modes prioritize travel efficiency, safety and comfort, or a trade-off among these objectives, respectively.
- Both noncooperative game-theoretic approaches simulate human-driver interaction and produce reasonable decisions for AVs.
- SE shows better performance than NE in driving safety, ride comfort and travel efficiency across the reported evaluations.
- The proposed method remains limited in handling various complex scenarios because models must differ substantially as traffic conditions change.
- Future work targets computation efficiency, adaptability to complex traffic scenarios, and real-time validation through hardware-in-the-loop experiments.