Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving

Domain:

mobility

Record type:

papermodel
Creator:
YanZhaCheLu,
Host:avatar
In this work, we reconceptualize autonomous driving as a generalized language and formulate the trajectory planning task as next waypoint prediction. We introduce Max-V1, a novel framework for one-stage end-to-end autonomous driving. Our framework presents a single-pass generation paradigm that aligns with the inherent sequentiality of driving. This approach leverages the generative capacity of the VLM (Vision-Language Model) to enable end-to-end trajectory prediction directly from front-view camera input. The efficacy of this method is underpinned by a principled supervision strategy derived from statistical modeling. This provides a well-defined learning objective, which makes the framework highly amenable to master complex driving policies through imitation learning from large-scale expert demonstrations. Empirically, our method achieves the state-of-the-art performance on the nuScenes dataset, delivers an overall improvement of over 30% compared to prior baselines. Furthermore, it exhibits superior generalization performance on cross-domain datasets acquired from diverse vehicles, demonstrating notable potential for cross-vehicle robustness and adaptability. Due to these empirical strengths, this work introduces a model enabling fundamental driving behaviors, laying the foundation for the development of more capable self-driving agents. Code will be available upon publication.

Visit

arxiv.org

Tasks

computer vision

Tags

Computer Vision and Pattern RecognitionArtificial IntelligenceRobotics

Similar

Metric-AI-Lab/less-is-more-embeddingsViDeBERTa: A powerful pre-trained language model for VietnameseIs Education More Benficial to the Less Able? Econometric Evidence from EthiopiaLess is More: Adapting Text Embeddings for Low-Resource Languages with Small Scale Noisy Synthetic DataGod Is More Powerful, But the Prophet is Closer: The Spiritual Security Framework for Understanding Seventh-Day Adventist Migration to Spirit Churches in Rural ZambiaBuilding a Strong Instruction Language Model for a Less-Resourced Language

Metric-AI-Lab/less-is-more-embeddings

Official repository for the paper "Less is More: Adapting Text Embeddings for Low-Resource Languages

ViDeBERTa: A powerful pre-trained language model for Vietnamese

This paper presents ViDeBERTa, a new pre-trained monolingual language model for Vietnamese, with thr

Is Education More Benficial to the Less Able? Econometric Evidence from Ethiopia

The paper investigates whether returns to schooling in Ethiopia vary according to the ability of ind

Less is More: Adapting Text Embeddings for Low-Resource Languages with Small Scale Noisy Synthetic Data

Low-resource languages (LRLs) often lack high-quality, large-scale datasets for training effective t

God Is More Powerful, But the Prophet is Closer: The Spiritual Security Framework for Understanding Seventh-Day Adventist Migration to Spirit Churches in Rural Zambia

Despite the remarkable growth of Christianity across sub-Saharan Africa, beliefs concerning witchcra

Building a Strong Instruction Language Model for a Less-Resourced Language

Large language models (LLMs) have become an essential tool for natural language processing and artif