Wayformer: Motion Forecasting via Simple & Efficient Attention Networks

Abstract

Motion forecasting for autonomous driving is a challenging task because complex driving scenarios involve a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road geometry, lane connectivity, time-varying traffic light state, and history of a dynamic set of agents and their interactions into an effective encoding. To model this diverse set of input features, many approaches proposed to design an equally complex system with a diverse set of modality specific modules. This results in systems that are difficult to scale, extend, or tune in rigorous ways to trade off quality and efficiency. In this paper, we present Wayformer, a family of simple and homogeneous attention based architectures for motion forecasting. Wayformer offers a compact model description consisting of an attention based scene encoder and a decoder. In the scene encoder we study the choice of early, late and hierarchical fusion of input modalities. For each fusion type we explore strategies to trade off efficiency and quality via factorized attention or latent query attention. We show that early fusion, despite its simplicity, is not only modality agnostic but also achieves state-of-the-art results on both Waymo Open Motion Dataset (WOMD) and Argoverse leaderboards, demonstrating the effectiveness of our design philosophy.

References

Page 1

	Year	Citations
MizAR 60 for Mizar 50 DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)	2023	73.5K
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)	2018	45.3K
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows Ze Liu, Yutong Lin, Yue Cao, 2021 IEEE/CVF International Conference on Computer Vision (ICCV) Swin TransformerConvolutional Neural NetworkMachine VisionImage AnalysisMachine Learning	2021	27.9K
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, arXiv (Cornell University) Convolutional Neural NetworkEngineeringMachine LearningImage RetrievalMultimodal Llm	2020	21.2K
Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations (Short Paper) DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)	2023	14.1K
Decoupled Weight Decay Regularization Ilya Loshchilov, Frank Hutter arXiv (Cornell University) EngineeringMachine LearningWeight DecayAtomic DecompositionImage Analysis	2017	9K
Emerging Properties in Self-Supervised Vision Transformers Mathilde Caron, Hugo Touvron, Ishan Misra, 2021 IEEE/CVF International Conference on Computer Vision (ICCV) Self-supervised Vit FeaturesImage ClassificationConvolutional Neural NetworkImage AnalysisMachine Vision	2021	4.6K
Social LSTM: Human Trajectory Prediction in Crowded Spaces Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Artificial IntelligenceCrowd SimulationEngineeringMachine LearningDifferent Trajectories	2016	3.4K
Conformer: Convolution-augmented Transformer for Speech Recognition Anmol Gulati, James Qin, Chung‐Cheng Chiu, EngineeringMachine LearningGlobal DependenciesSpeech RecognitionNatural Language Processing	2020	2.5K
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, arXiv (Cornell University) EngineeringMachine LearningSpoken Language ProcessingSpeech RecognitionNatural Language Processing	2020	2.4K

Page 1