基于深度信息约束的三维高斯泼溅实现高质量无纹理物体重建

High-Quality Textureless Object Reconstruction via Depth-Constrained 3D Gaussian Splatting

  • 摘要: 高质量的无纹理物体三维重建在增强现实、工业检测及机器人自主抓取等领域具有重要的应用价值。然而,由于表面缺乏显著的纹理特征,传统的基于多视图立体几何(Multi-View Stereo, MVS)及特征匹配的方法往往难以建立可靠的对应关系,导致重建失败。近年来,三维高斯泼溅(3D Gaussian Splatting, 3DGS)技术虽在渲染质量与速度上取得了突破,但其在处理无纹理表面时仍面临挑战,常导致重建模型出现空洞与伪影。为应对当前该领域标准化评测数据较少的现状,本文构建了一个包含11种典型无纹理物体的高质量RGB-D数据集,以促进相关研究的发展。基于此,本文提出了一个名为质量-深度增强三维高斯泼溅(Quality-Depth Enhanced 3D Gaussian Splatting, QDE-GS)的创新重建框架。首先,该框架集成Mask R-CNN实例分割网络,精确剥离背景干扰,聚焦于目标物体的几何恢复;其次,针对无纹理区域几何信息缺失的核心难题,本文提出了几何-光度融合网络(Geometric-Photometric Fusion Network,GP-FusionNet)。该网络创造性地将输入的实际深度图与RGB光度信息进行深层融合,并通过3D U-Net结构进行正则化处理,输出高保真的稠密深度图,为高斯点的初始化提供了强有力的几何先验约束。此外,为进一步抑制渲染伪影并提升模型对低质量输入的鲁棒性,本文引入了基于无参考图像质量评估的质量感知优化算法与失真网络;最后,为提高整体流程的效率,去除冗余的相机视角,本文还设计了基于相机位姿的视图优选算法。实验结果表明,QDE-GS在自建数据集和公开数据集上的新视角合成任务中均表现卓越。定量评估显示,该方法在峰值信噪比(Peak Signal-to-Noise Ratio, PSNR)、结构相似性指数(Structural Similarity Index Measure, SSIM)及学习感知图像块相似度(Learned Perceptual Image Patch Similarity, LPIPS)这些关键图像质量指标上均优于当前主流方法,显著提升了无纹理物体在复杂视角下的渲染保真度。

     

    Abstract: High-quality 3D reconstruction of textureless objects is of significant importance in applications such as Augmented Reality (AR), industrial inspection, and robotic autonomous grasping. However, owing to the lack of distinctive surface textures, traditional methods based on Multi-View Stereo (MVS) and feature matching often fail to establish reliable correspondences, leading to unsuccessful reconstructions. In recent years, 3D Gaussian Splatting (3DGS) has achieved notable advances in rendering quality and computational efficiency. Nevertheless, it continues to face inherent limitations when applied to textureless surfaces, frequently producing incomplete reconstructions characterized by holes and artifacts. To address the lack of standardized evaluation datasets in this domain, we construct a high-quality RGB-D dataset comprising 11 categories of representative textureless objects, thereby supporting further research in this area. Building upon this dataset, we propose a novel reconstruction framework, termed QDE-GS. This framework incorporates the Mask R-CNN instance segmentation network to effectively eliminate background interference and enable focused geometric reconstruction of the target object. To tackle the core challenge of missing geometric information in textureless regions, we further develop a dual-branch network, referred to as GP-FusionNet. The proposed network performs a deep fusion of the input raw depth maps and RGB photometric information, followed by regularization using a 3D U-Net architecture to produce a high-fidelity dense depth map. This output provides robust geometric priors for the initialization and optimization of 3D Gaussians. Furthermore, to suppress rendering artifacts and enhance robustness to low-quality inputs, we introduce a quality-aware optimization strategy along with a distortion network based on No-Reference Image Quality Assessment (NR-IQA). In addition, to improve overall pipeline efficiency and eliminate redundant camera views, we design a view selection algorithm guided by camera pose information. Experimental results demonstrate that QDE-GS consistently achieves superior performance on the Novel View Synthesis (NVS) task across both the proposed dataset and publicly available benchmarks. Quantitative evaluations further confirm that the proposed method surpasses existing state-of-the-art methods on key image quality metrics, including Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). In addition, the method substantially improves the rendering fidelity of textureless objects under challenging and complex viewing conditions.

     

/

返回文章
返回