5th DriveX Workshop In conjunction with Intelligent Vehicles (IV) Symposium 2026

Foundation Models for Autonomous Driving

A premier forum uniting academic, industry, and standards communities to explore advances in Foundation Models and 3D Perception in Cooperative Autonomous Driving (CAD).

June 22, 2026 Detroit, MI, United States Room: Brule A (Level 5), Detroit Marriott at the Renaissance Center In conjunction with Intelligent Vehicles (IV) Symposium 2026
Curated keynote lineup from academia & industry
Focus on real-world V2X datasets & benchmarks
Safety, robustness, and trustworthy autonomy

Introduction

The 5th edition of the full-day DriveX workshop explores advances in Foundation Models and 3D Perception in Cooperative Autonomous Driving (CAD). This workshop brings together leading researchers and practitioners to discuss cutting-edge developments in large language models (LLMs), vision-language models (VLMs), vision-language action models (VLAs), and their applications to autonomous driving systems. Topics include 3D object detection, semantic segmentation, sensor fusion, V2X communication, cooperative perception, and real-world applications.

We explore methods to enhance scene understanding, perception accuracy, dataset curation, and novelty detection. By uniting experts across perception, V2X, and foundation model domains, this workshop aims to foster innovation in cooperative autonomous driving and intelligent transportation systems. The workshop addresses critical challenges in multi-modal sensor data fusion, vehicle-infrastructure coordination, and intelligent transportation systems that leverage both onboard and roadside sensing capabilities.

This year, we expanded our focus with the addition of V2X applications, exploring real-world vehicle-to-infrastructure connectivity that extends past collaborative perception. The workshop provides a platform for discussing V2X for localization, tolling, road safety, monitoring, and data analytics, bridging the gap between theoretical advances and practical deployment in intelligent transportation systems. Through keynote presentations, panel discussions, paper presentations, and challenge tracks, DriveX 2026 creates a comprehensive forum for advancing the state-of-the-art in foundation model-driven cooperative autonomous driving.

Topics of Interest

3D Perception

Cooperative Perception

Foundation Models

Applications in Real World

Schedule (Tentative)

Time Session
08:00 – 08:10 Introduction
08:10 – 08:30
Dr. Walter Zimmer
Opening Keynote Talk: Cooperative Perception Meets Foundation Models: Unlocking 3D Scene Understanding for Autonomous Driving Keynote

Dr. Walter Zimmer

UCLA & TUM

Abstract

Autonomous driving in urban environments is fundamentally limited by the range, occlusions, and failure modes of vehicle-only perception. This opening keynote shows research advances in cooperative roadside–vehicle perception by fusing multi-modal data from onboard sensors and intelligent roadside infrastructure via V2X communication to extend situational awareness beyond line of sight. The proposed methods improve real-time 3D object detection and tracking in dense traffic and are supported by new large-scale, multi-modal datasets for benchmarking cooperative perception in real-world urban settings. By integrating recent advances in foundation models, including vision-language models, the work further enables semantic understanding of complex traffic scenes, laying the groundwork for AI-driven urban digital twins and safer, more efficient intelligent transportation systems.

Speaker Bio

Dr. rer. nat. Walter Zimmer is a post-doctoral researcher at the University of California Los Angeles (UCLA) and guest researcher at the Technical University of Munich (TUM). He received his Ph.D. from the Technical University of Munich (TUM) in 2025. His research focuses on cooperative autonomous driving, 3D perception and 3D foundation models. He has authored over 40 publications at top venues such as CVPR, ICCV, ECCV, ICML, NeurIPS, and T-PAMI. Dr. Zimmer previously worked as an Autonomous Systems Engineer at the STTech startup and research assistant at Siemens AG. His work has earned multiple awards, including the IEEE ITSS Best Student Paper Award 2023 and IEEE ITSS Best Dissertation Award 2025.

08:30 – 08:50
Dr. Jim Misener
Keynote 1: Can V2X be an ADAS or ADS Sensor? Keynote

James (Jim) Misener

WSP

Abstract

V2X offers a compelling vision as an Advanced Driver Assistance System (ADAS) or Automated Driving System (ADS), based on the long-held promise of of cooperative perception, as V2X extends traditional sensors by enabling non-line-of-sight awareness, improving time-to-collision horizons, and addressing occlusions and intersection risks. However, major functional safety (FuSa) barriers limit its role as a true ADAS/ADS sensor. For V2V, ISO 26262 compliance is difficult, chipset support is limited, and OEMs are unlikely to rely on V2X for direct control actions. For V2I, challenges are even greater due to the lack of FuSa-compliant infrastructure in the US, although this might not necessarily be the case with Europe. Also, focusing on the air interface and not the legacy discussion, from a communications perspective, 5G NR V2X offers the performance needed for advanced applications, but spectrum constraints remain a critical issue. Given this, pragmatic path forward is phased: start with advisory use, then integrate into ADAS decision support, and only later pursue selective safety-critical roles—primarily in direct V2V use cases. In the end, V2X is a powerful enabler of perception and prediction, but in the foreseeable future it will act as a supporting sensor rather than a primary, safety-critical control input.

Speaker Bio

Jim Misener is Senior Vice President, Digital Infrastructure and Mobility Innovation at WSP in the U.S. As part of WSP's advisory and planning practice and executive leadership team, he develops and delivers strategies to harness innovation in digitizing transportation deployments, with focus on the synergies of radio communications, private sector and government to provide technical and enterprise transformation to enhance mobility and safety. Previously, he was Senior Director, Product Management and the Global V2X Ecosystem Lead for Qualcomm. He developed and executed V2X deployment strategy across all global regions, working with government, standards, automotive, road owner-operator and telecommunications partners. He also developed IoT solutions for transportation markets. Misener also established and led the automotive standards team at Qualcomm. Misener served as an ITS America Board member and was a 5GAA founder and Board member. He serves on the IEEE ITS Society Board of Governors and as an ITS California senior advisor. He is an Advisory Council member to Mobility 21-Traffic 21 led by Carnegie Mellon University and on the Technical Advisory Board to the Center for Connected and Automated Transportation at the University of Michigan. Misener was a pioneer in vehicle-highway automation and vehicle safety communication at the California Partners for Advanced Transit and Highways (PATH) at UC Berkeley. He has served as PATH Executive Director, Executive Advisor to Booz Allen Hamilton, and has been an independent consultant. He holds BS and MS degrees from UCLA and USC, respectively. He is an IEEE Fellow, with dozens of international patents and over 80 publications in the Connected and Automated Vehicle domain.

08:50 – 09:10
Prof. Jiaqi Ma
Keynote 2: Open-World Multi-Agent Mobility Keynote

Prof. Jiaqi Ma

UCLA

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Dr. Jiaqi Ma is Director of the FHWA/UCLA Center of Excellence on New Mobility and Automated Vehicles, Professor at the UCLA Samueli School of Engineering, Director of the UCLA Mobility Lab, and Associate Director of the UCLA Institute of Transportation Studies. He has led and managed numerous transportation research projects funded by the U.S. Department of Transportation, National Science Foundation, state Departments of Transportation, and other federal, state, and local agencies. His research spans automated driving, mobility digital twins, multimodal sensing, cooperative perception and decision-making, robotics, spatial data mining, simulation, and reasoning.

09:10 – 09:30
Prof. Ignacio Alvarez
Keynote 3 Keynote

Prof. Ignacio Alvarez

THI

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Dr. Alvarez is a professor at THI. He is now pioneering the next wave of human-centric AI to build a safer, more intelligent mobility future. Before joining THI, he worked at Intel Labs as a Senior Researcher and Principal Engineer, leading the development of intelligent systems from concept to production.

09:30 – 09:50
Lena Wild
Keynote 4: Rethinking Road Representations: From Self-Updating HD Mapping to Road-Level Reasoning Keynote

Lena Wild

TRATON Group/Scania & KTH Royal Institute of Technology, Sweden

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Lena Wild is an industrial PhD student at TRATON Group and KTH Royal Institute of Technology, where her research focuses on map learning for autonomous driving, road perception, and explainable end-to-end AI systems. She was also a visiting PhD student at Stanford University's Autonomous Systems Lab, and was working with Prof. Marco Pavone, and is affiliated with the Wallenberg AI, Autonomous Systems and Software Program (WASP). Before starting her PhD, she earned degrees in Technical Physics from TU Wien and in Philosophy and Classical Languages from the University of Vienna, reflecting a unique interdisciplinary background. Her research bridges deep learning, robotics, and map-centered reasoning to advance trustworthy and intelligent autonomous driving systems.

09:50 – 10:20 Morning Break & Networking Break
10:20 – 10:40
Dr. Sergei Avedisov
Keynote 5: Cooperative Driving Keynote

Dr. Sergei Avedisov

Toyota North America

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Dr. Sergei Avedisov is a Principal Researcher at Toyota Motor North America R&D, InfoTech Labs. He received his Ph.D. in Mechanical Engineering from the University of Michigan in 2019. His research interests include cooperative automated driving, cooperative perception, cooperative maneuvering, platooning, teleoperated driving, and V2X communications.

10:40 – 11:00
Prof. Vincent Fremont
Keynote 6: Foundation Models and 3D Semantic Occupancy Grids Keynote

Prof. Vincent Frémont

École Centrale de Nantes

Abstract

Semantic and panoptic occupancy prediction for road scene analysis provides a dense 3D representation of the ego vehicle's surroundings. Current camera-only approaches typically rely on costly dense 3D supervision or require training models on data from the target domain, limiting deployment in unseen environments. In this talk, I will describe FreeOcc, a training-free pipeline that leverages pretrained foundation models to predict 3D semantic and panoptic occupancy grids from multi-view images.

Speaker Bio

Prof. Vincent Frémont received his M.Sc. and PhD degrees in automatic control and computer science from the Ecole Centrale de Nantes, France. From 2005 to 2018, he was an Associate Professor at the Université de Technologie de Compiègne (UTC) within the Heudiasyc Lab. Since 2018, he is a Full Professor at Ecole Centrale de Nantes within the ARMEN team at the LS2N Lab. His research interests belong to perception systems and scene understanding for autonomous mobile robotics with an emphasis on computer vision, deep learning and multi-sensor fusion.

11:00 – 11:20
Prof. Ayesha Choudhary
Keynote 7 Keynote

Prof. Ayesha Choudhary

Jawaharlal Nehru Uni.

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Ayesha Choudhary is an Assistant Professor at the School of Computer & Systems Sciences, Jawaharlal Nehru University (JNU), New Delhi, where she has been a faculty member since 2013. Her research lies at the intersection of computer vision, machine learning, and digital image processing, with a strong focus on surveillance video analysis, distributed camera networks, and visual event understanding. She earned her Ph.D. from IIT Delhi, where her work focused on automated analysis of surveillance videos, and has since contributed extensively to both academic research and applied projects in visual computing. She brings over a decade of experience spanning academia and research labs, and has authored influential publications on multi-camera systems and video analytics.

11:20 – 11:40
Marion Neumeier
Keynote 8: Uncertainty-Aware Trajectory Prediction for Automated Driving Keynote

Marion Neumeier

THI

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Marion Neumeier is Technology Field Lead for AI in Mobility Applications at the Technische Hochschule Ingolstadt (THI) and a doctoral researcher in cooperation with the Technical University of Munich. Her research focuses on AI algorithms for automated driving, vehicle trajectory prediction, and uncertainty quantification.

11:40 – 12:00
11:40 – 12:30 Panel Discussion I & Group Picture Panel
12:30 – 13:30 Lunch
13:30 – 13:50
Prof. Arnaud de La Fortelle
Keynote 10: The Grey Zone: Engineering ODDs between Foundation Models and the Law Keynote

Dr. Arnaud de La Fortelle

Heex Technologies

Abstract

A significatif bottleneck in deploying Foundation-Model-based driving systems is the assignment of operating conditions to validated, excluded, or uncertain status. We formalise this as an ODD-engineering problem with two layers: a deterministic, trigger-based technical sub-ODD that a runtime supervisor can monitor in real time, and a semantic ODD legible to regulators and courts. The Grey Zone, i.e. conditions monitored but not yet certified is treated as a sampling target. This is lined to importance sampling over the ODD space. Combined with triggering on boundary events, this approach yields targeted datasets that each learning cycle converts into white (validated) or black (excluded). Smart-Data makes this loop operational as it helps build agile data pipeline at scale. It acts as the data-governance layer beneath the Foundation Model: the two are complementary by construction.

Speaker Bio

Dr. Arnaud de La Fortelle is co-founder and CTO of Heex Technologies. He was professor and director (2008–2021) of the Center for Robotics at MINES ParisTech (PSL University) and a Visiting Professor at UC Berkeley. He specializes in cooperative systems, intelligent transportation, autonomous driving, and smart-data platforms for edge and cloud AI systems.

13:50 – 14:10
Dr. Zhenzhen Liu
Keynote 11: Understanding and Mitigating Heterogeneity in Collaborative 3D Perception Keynote

Zhenzhen Liu

Cornell Uni.

Abstract

Collaborative perception addresses the occlusion and range limitations of single-vehicle systems by fusing information from connected vehicles and infrastructure. However, publicly available V2X datasets have historically offered limited diversity in sensing configurations and operating environments, with collaborating agents often sharing nearly identical LiDAR and perception stacks. As a result, many challenges encountered in real-world deployments remain underexplored. In this talk, I will first present Mixed Signals, a real-world V2X dataset featuring heterogeneous vehicle and infrastructure sensor setups collected in a left-hand-traffic environment, exposing perception and localization challenges that are often overlooked. I will then present DiffuBox, a domain-agnostic diffusion-based framework for refining 3D object detections. By conditioning a generative model on local point geometry, DiffuBox improves localization accuracy by correcting imperfect bounding box estimates without retraining. Together, these works highlight the importance of moving beyond homogeneous benchmarks and developing perception systems that better accommodate the diversity and uncertainty of real-world collaborative driving.

Speaker Bio

Zhenzhen Liu is a PhD candidate in Computer Science at Cornell University. Her research focuses on applied machine learning and computer vision, with particular interests in generative models, perception systems, and reliability and robustness, including topics such as out-of-distribution detection, domain adaptation and collaborative perception.

14:10 – 14:30
Dr. Zhong Cao
Keynote 12: Driving Foundations for Scaling Beyond the Driving Domain Keynote

Dr. Zhong Cao

Uni. of Michigan

Abstract

Abstract will be announced closer to the workshop date.

Speaker Bio

Dr. Zhong Cao is an Assistant Research Scientist in the Department of Civil and Environmental Engineering at the University of Michigan. He received his B.S. and Ph.D. in automotive engineering from Tsinghua University. His research interests include autonomous vehicles, trustworthy AI, reinforcement learning, and continual learning for self-driving systems.

14:30 – 14:50
Dr. Can Cui
Keynote 13: Foundation Models for Human-Autonomy Teaming in Autonomous Vehicles Keynote

Dr. Can Cui

BCAI

Abstract

Autonomous driving technology is experiencing a paradigm shift. While traditional systems have achieved impressive performance in perception and control, they remain "automated tools" that are black boxes that optimize for geometric safety but fail to understand human intent, communicate reasoning, or adapt to personal preferences. This presentation introduces a Human-Autonomy Teaming (HAT) framework designed to bridge this gap. By leveraging the emergent capabilities of Foundation Models (Large Language Models and Vision-Language Models), we propose a system where the vehicle functions not as a passive tool, but as an active, collaborative teammate.

Speaker Bio

Can Cui is currently a research scientist at the Bosch Center for Artificial Intelligence in Sunnyvale, CA. He received his Ph.D. in Autonomous Driving from Purdue University in 2025, under the supervision of Dr. Ziran Wang, and his M.S. in Electrical and Computer Engineering from the University of Michigan in 2022. Dr. Cui's research interests lie at the intersection of foundation models and autonomous systems, with a specific focus on Large Language Models (LLMs), Vision Language Models (VLMs), and human-autonomy teaming. He has authored papers in top-tier journals and conferences, including IEEE Transactions on Intelligent Vehicles, CVPR, EMNLP, and ICCV. He is the recipient of the 2026 IEEE ITSS Best Dissertation Award, the 2025 Purdue CEE Best Dissertation Award and the 2025 TRB Vehicle-Highway Automation Committee Best Paper Award. Dr. Cui actively serves the academic community as a Guest Editor for the SAE International Journal of Connected and Automated Vehicles and has organized multiple workshops on foundation models for autonomous driving at CVPR, ICCV, and WACV.

14:50 – 15:20 Afternoon Break: Demo by Nomadic AI (Curating Edge-Case Datasets for Self-Driving Vehicles) Break
15:20 – 15:40
Prof. Cathy Wu
Keynote 14: What Could Driving Behavior Prediction Do for Public Road Infrastructure? Keynote

Prof. Cathy Wu

MIT

Abstract

Trajectory prediction driving models, developed by and for autonomous driving, are usually evaluated on a finite set of logged road maps. Those maps cover only a small fraction of the roadway geometries a model may encounter: lanes can narrow in work zones, lane markings can be redrawn, traffic-calming features can shift a vehicle laterally, and intersections can be redesigned. We introduce CrossRoads, a benchmark for testing whether map-conditioned driving simulation agents remain stable under such roadway counterfactuals. Given a scenario from the Waymo Open Motion Dataset (WOMD), CrossRoads applies deterministic, validated edits to the lane graph and compares closed-loop rollouts before and after the edit using matched random seeds. The benchmark includes four classes of edits–boundary cues, lane-envelope changes, horizontal deflections, and route/topology changes–chosen because they stress different parts of driving behavior: perception of map semantics, lateral clearance, speed choice, and route consistency. We evaluate three runnable model families on 600 scenario selections, 3 operating regimes, and over 10800 paired rollout/evaluation units. Evaluated models are not uniformly robust: on validator-accepted edits, chicane-like horizontal deflections usually slow agents in free-flow scenes but have little effect in dense city scenes, while lane-line narrowing alone often fails to produce the speed reductions expected from roadway-design guidance. CrossRoads is released as code, patch manifests, validators, scenario/slice metadata, and reproduction scripts so that future models can be compared on the same counterfactual maps. Interactive demo: https://counterfactualroad.pages.dev

Speaker Bio

Cathy Wu is the Class of 1954 Career Development Professor at MIT, holding appointments in LIDS, CEE, and IDSS. She holds a Ph.D. in EECS from UC Berkeley, and B.S. and M.Eng. in EECS from MIT, and completed a Postdoc at Microsoft Research. Her research group advances data science for transportation systems, focusing on machine learning for optimization. Cathy is the recipient of the NSF CAREER (2023), the Ole Madsen Mentoring Award (2025), the IEEE ITS Best Dissertation Award (2019), and the CUTC Milton Pikarsky Memorial Award (2018). She serves on the Board of Governors for the IEEE ITSS, is an Associate Editor or Area Chair for ICML, NeurIPS, ICRA, Transportation Research Part C, and Operations Research, and served as Program Co-chair for RLC 2025. She is also the Chair and Co-founder of the REproducible Research In Transportation Engineering (RERITE) Working Group.

15:40 – 16:00
Prof. Torsten Schön
Keynote 15: OaK-Architecture for Autonomous Driving Keynote

Prof. Torsten Schön

THI

Abstract

In this talk, we will discuss a proposal for applying the principles of Rich Sutton’s OaK architecture to autonomous driving. This gives rise to new ideas and avenues of research for the use of foundation models in autonomous driving, and we will examine the various building blocks required for this. The talk does not present a ready-made solution, but rather a thought experiment intended to highlight an alternative direction to current developments.

Speaker Bio

Prof. Dr. Torsten Schön is Professor for Computer Vision for Intelligent Mobility Systems at the Technische Hochschule Ingolstadt (THI). Before joining THI in 2020, he was Senior Data Scientist for Artificial Intelligence at Audi AG. His research focuses on computer vision, machine learning, and AI for intelligent mobility systems.

16:00 – 16:20
Korbinian Moller
Keynote 16: Foundation Models as an Enabler for Scalable Scenario Generation and Analysis Keynote

Korbinian Moller

TUM

Abstract

Foundation models offer a promising way to scale scenario generation and analysis for autonomous driving by combining data-driven learning with structured knowledge about traffic scenes, rules, and system behavior. This talk discusses how model families such as LLMs, VLMs, diffusion models, and world models can support the generation of diverse and safety-critical scenarios as well as their evaluation. Based on recent research examples, it highlights current capabilities, limitations, and open challenges toward reliable, controllable, and usable scenario pipelines.

Speaker Bio

Korbinian Moller is a PhD researcher at the Technical University of Munich (TUM), Chair of Autonomous Vehicle Systems. His research focuses on motion planning, safety-critical decision-making, and real-time systems for autonomous driving, including occlusion-aware trajectory planning.

16:20 – 17:10 Panel Discussion II & Group Picture Panel
17:10 – 17:20
Björn Möller
Paper Oral 1: OpenViCA: Video Continuation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data Oral

Björn Möller

TU Braunschweig

Abstract

The presentation introduces OpenViCA, an open video-continuation system for automotive driving scenes. After a brief introduction to its discrete token-based generation pipeline, we explore how next-token sampling affects the system’s creativity when predicting future video continuations. Selected examples reveal a central challenge of balancing diverse, creative continuations and realistic, temporally consistent driving scenes.

Speaker Bio

Björn is a research assistant and PhD candidate at the Institute for Communications Technology at the TU Braunschweig. His research focuses on deep learning for computer vision, particularly the generation of video and LiDAR data for autonomous driving.

17:20 – 17:30
Santosh Vasa
Paper Oral 2: AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models Oral

Aditi Ramadwar

Mercedes-Benz Research & Development North America

Abstract

Training of autonomous driving systems requires extensive datasets with precise annotations to attain robust performance. Human annotations suffer from imperfections, and multiple iterations are often needed to produce high-quality datasets. However, manually reviewing large datasets is laborious and expensive. In this paper, we introduce AutoVDC (Automated Vision Data Cleaning) and investigate the utilization of Vision-Language Models (VLMs) to automatically identify erroneous annotations in vision datasets, thereby enabling users to eliminate these errors and enhance data quality.

Speaker Bio

Aditi Ramadwar is a Machine Learning Engineer at Mercedes-Benz Research & Development North America, where she develops perception and sensor fusion systems for autonomous driving, spanning deep learning, Vision-Language Models, and real-world deployment. She earned her Master's degree in Mechatronics, Robotics, and Automation Engineering from the University of Maryland, following earlier industry experience in robotics and computer vision. Her research and engineering interests include autonomous driving perception, multi-sensor fusion, explainable AI, and data-centric machine learning, and she has recently presented work on evaluating the faithfulness of Vision-Language Model explanations at NeurIPS.

17:30 – 17:40
17:40 – 17:50
17:50 – 18:00 Group Picture — All Organizers & Speakers

Final schedule, room allocation, and speaker order will be announced closer to the workshop date.

Paper Track

DriveX 2026 invites high-quality contributions on foundation models, V2X-based cooperative perception, large driving models, 3D perception, and related topics outlined above.

We welcome:

Submissions must follow the official IEEE IV 2026 style guidelines. Detailed submission instructions will be provided.

Accepted Papers

OpenViCA: Video Continuation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data

Björn Möller, Zhengyang Li, Malte Stelzer, Thomas Graave, Fabian Bettels, Muaaz Ataya and Tim Fingscheidt

AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models

Aditi Ramadwar, Santosh Vasa, Jnana Rama Krishna Darabattula, Md Zafar Anwar, Stanislaw Antol, Andrei Vatavu, Thomas Monninger, Sihao Ding

Paper Awards

DriveX Grand Challenge

The DriveX Grand Challenge fosters rigorous, reproducible benchmarking of cooperative perception and planning on real-world datasets. Tracks are designed in close collaboration with dataset creators and industry partners.

Feb 18 – May 1

TUMTraf-V2X

Vehicle-to-infrastructure cooperative 3D detection and tracking using the TUMTraf-V2X dataset. Teams fuse infrastructure-mounted LiDAR, radar, and cameras with on-vehicle sensors to tackle occlusion handling, long-range awareness, and reliability under real-world traffic conditions.

Go to challenge →
Feb 11 – May 20

doScenes

Natural-language-driven scene understanding and visual language navigation for autonomous vehicles. Participants build models that interpret free-form human instructions and ground them in driving scenes, advancing research on intuitive human-vehicle interaction.

Go to challenge →
Feb 20 – May 1

MDrive

End-to-end closed-loop cooperative driving across multiple agents. Using the MDrive benchmark, teams develop and evaluate planning, communication, and coordination strategies for safe multi-vehicle autonomy in complex traffic scenarios.

Go to challenge →

Organizers

Invited Program Committee

Chuheng Wei

University of California, Riverside

Xingcheng Zhou

Technical University of Munich

Haoxuan Ma

University of California, Los Angeles

Haibao Yu

The University of Hong Kong

Tzu-yun Tseng

The University of Sydney

Yaoqi Huang

The University of Sydney

Zenxing Ming

The University of Sydney

Xiangyu Chen

Waymo

Zhihao Zhao

University of California, Los Angeles

Sponsors

DriveX 2026 welcomes sponsorship from industry, startups, and institutions interested in foundation models, cooperative perception, simulation, and large-scale autonomous driving systems.

For sponsorship opportunities, please contact: wz@ucla.edu.