Overview of Multi-Modal Visual Data Compression for Embodied Intelligence
-
Abstract
Embodied artificial intelligence (AI) represents a paradigm shift from passive analysis to active interaction, serving as a critical pathway toward general AI. The emergence of the Vision-Language-Action (VLA) model has endowed robotic agents with unprecedented semantic understanding and generalization capabilities. However, the immense computational demands of these models necessitate an “Edge-Cloud Synergy” deployment architecture. This creates a major bottleneck: the massive data throughput generated by multimodal visual sensors far exceeds the limited and fluctuating bandwidth available in real-world scenarios. Traditional compression standards, designed for human-vision-oriented or machine-vision-oriented purposes, suffer from a fundamental objective mismatch in embodied contexts. Research indicates that artifacts introduced by these methods at low bitrates result in a significant performance degradation in task success rates and lead to dangerous error accumulation within the closed-loop control system. To address these urgent needs, this study provides a comprehensive survey of visual data compression and communication technologies tailored for embodied perception. First, we introduce the new paradigm of embodied-AI-oriented compression. We distinguish it from traditional paradigms by analyzing unique characteristics such as closed-loop error propagation, nonlinear performance responses, and the heterogeneous perceptual preference of downstream VLA models. Second, we systematically review state-of-the-art compression algorithms for planar vision, panoramic and fisheye vision, stereo vision, RGB-D data, and 3D point clouds, classified based on sensory modality. The survey places particular emphasis on Deep Neural Network-based innovations, including nonlinear transformation, end-to-end optimization, and cross-modal fusion, analyzing their adaptability to robotic manipulation and navigation tasks. Finally, we summarize existing challenges regarding latency and robustness and outline future research directions, including standardized evaluation benchmarks for embodied visual compression, task-driven dynamic compression, unified multimodal representation, end-to-end joint optimization of compression and control policies, and generalized compression for heterogeneous embodied models.
-
-