单曝光三维压缩成像:从单张压缩快照中重建三维场景

3D Snapshot Compressive Imaging: Reconstructing 3D Scenes from a Single Compressed Image

  • 摘要: 高维数据的捕获是信号处理和相关领域的长期挑战。单曝光压缩成像(Snapshot Compressive Imaging, SCI)使用二维探测器在单次曝光内积分捕捉高维数据,是一种低成本、高效率的计算成像技术。传统的SCI重建算法主要致力于从压缩测量中恢复二维视频帧,往往忽略了场景潜在的三维几何结构,导致重建结果在多视角观测时缺乏一致性,且受限于掩码的单一性。近年来,随着神经辐射场(Neural Radiance Fields, NeRF)和三维高斯泼溅(3D Gaussian Splatting, 3DGS)的三维场景表达技术的发展,使得从单次曝光的压缩图像中恢复三维场景成为可能。本文旨在系统综述单曝光三维压缩成像(3D Snapshot Compressive Imaging, SCI-3D)这一新兴领域的最新进展,它不仅能够高质量地重建多视角信息,更是完整地建模并重建了三维模型,使其可以实现高质量的新视角图像合成(Novel View Synthesis, NVS)。首先,本文建立了统一的数学模型,系统描述了从单张压缩测量到三维场景表征的逆问题求解框架;其次,深入梳理了从基于隐式神经表示的SCI-NeRF到基于显式点云表示的SCI-Gaussian与SCI-Splat的技术演进路径,重点分析了各阶段方法在应对监督信号高度退化、相机位姿未知及几何先验缺失等核心挑战时的关键技术突破;最后,本文对该领域的未来发展进行了展望,指出动态场景的四维时空重建以及基于通用大模型的前馈式网络将是实现高效、鲁棒端到端三维重建的重要方向。

     

    Abstract: The capture of high-dimensional data remains a long-standing challenge in signal processing and related fields. Snapshot compressive imaging (SCI), which integrally captures high-dimensional data using two-dimensional (2D) detectors, from within a single exposure, offers a low-cost and highly efficient computational imaging paradigm. Traditional SCI-reconstruction algorithms primarily focus on recovering 2D video frames from compressed data and often neglect the underlying three-dimensional (3D) geometric structure of the scene. This leads to a lack of consistency in multiview observations and limitations caused by mask singularity. The recent development of 3D scene representation technologies, such as neural radiance fields (NeRF) and 3D Gaussian splatting, has facilitated the recovery of 3D scenes from single-exposure compressed images. This paper provides a systematic review of the emerging field of 3D SCI (SCI-3D), which not only facilitates high-quality reconstruction of multiview information but also achieves complete 3D modeling, enabling high-fidelity novel view image synthesis. First, a unified mathematical model that systematically describes the framework for solving the inverse problem of mapping a single compressed measurement to a 3D scene representation is reviewed. Second, the technological evolution is thoroughly analyzed, tracing the path from the implicit neural-representation-based SCI-NeRF to the explicit point-cloud-representation-based SCI-Gaussian and SCI-Splat technologies. This analysis focuses on the key technical breakthroughs in addressing core challenges such as highly degraded supervisory signals, unknown camera poses, and absence of geometric priors. Finally, the paper offers an outlook on future research directions, identifying four-dimensional spatiotemporal reconstruction of dynamic scenes and feed-forward networks based on large foundation models as pivotal trends for achieving efficient, robust, and end-to-end 3D reconstruction.

     

/

返回文章
返回