Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

VITA: Vision-to-Action Flow Matching Policy

Record type:

papermodel
Creator:
GaoZhaLeeChu
Host:avatar
Conventional flow matching and diffusion-based policies sample through iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning mechanisms to incorporate visual information during the generative process, incurring substantial time and memory overhead. To reduce the complexity, we develop VITA(VIsion-To-Action policy), a noise-free and conditioning-free policy learning framework that directly maps visual representations to latent actions using flow matching. VITA treats latent visual representations as the source of the flow, thus eliminating the need of conditioning. As expected, bridging vision and action is challenging, because actions are lower-dimensional, less structured, and sparser than visual representations; moreover, flow matching requires the source and target to have the same dimensionality. To overcome this, we introduce an action autoencoder that maps raw actions into a structured latent space aligned with visual latents, trained jointly with flow matching. To further prevent latent space collapse, we propose flow latent decoding, which anchors the latent generation process by backpropagating the action reconstruction loss through the flow matching ODE (ordinary differential equations) solving steps. We evaluate VITA on 8 simulation and 2 real-world tasks from ALOHA and Robomimic. VITA outperforms or matches state-of-the-art generative policies, while achieving 1.5-2.3x faster inference compared to conventional methods with conditioning. Project page: ucd-dare.github.io Project page: ucd-dare.github.io Code: github.com

Visit

arxiv.org

Tasks

computer vision

Tags

Computer Vision and Pattern RecognitionArtificial IntelligenceRobotics

Similar

Jitubohh/VITAGenerative AI Against Poaching: Latent Composite Flow Matching for Wildlife ConservationGenerative AI Against Poaching: Latent Composite Flow Matching for Poaching PredictionMVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical MatchingLinking sanitation policy to service delivery in Rwanda and Uganda: From words to actionAI Policy Design and Implementation Toolkit: Twenty-five connected instruments for the ACTION policy lifecycle

Jitubohh/VITA

A three-layer healthcare system for Nigeria's informal workers: NFC health identity card, community

Generative AI Against Poaching: Latent Composite Flow Matching for Wildlife Conservation

Poaching poses significant threats to wildlife and biodiversity. A valuable step in reducing poachin

Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction

Poaching poses significant threats to wildlife and biodiversity. A valuable step in reducing poachin

MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Conse

Linking sanitation policy to service delivery in Rwanda and Uganda: From words to action

Abstract Motivation The gap between policy, implementation and outcome is neither new nor specifi

AI Policy Design and Implementation Toolkit: Twenty-five connected instruments for the ACTION policy lifecycle

A working instrument set for public officials designing, appraising, procuring, piloting and oversee