双向光流引导的集成成像时空超分辨率重建

Bidirectional Optical Flow-Guided Spatio-Temporal Super-Resolution Reconstruction for Integral Imaging

  • 摘要: 集成成像作为目前最具发展潜力的裸眼三维视频显示技术,面临着帧率和分辨率之间相互制约的问题,导致在生成动态视频时,往往需要在空间清晰度和时域流畅度之间做出权衡,从而影响了最终的视觉效果和用户体验。为解决这一问题,本文提出了一种双向光流引导的时空分辨率协同增强方法,实现了集成成像高分辨率高帧率视频的动态三维显示。我们通过双向光流估计技术与双向光流引导的特征对齐模块精确捕捉帧间的运动信息,有效避免了传统方法中因视差跳变导致的边界伪影,实现了跨帧的精准融合,保持了空间细节和视差一致性,此外,时空超分辨率重建模块对图像进行高分辨率重建,显著提高了集成成像视频的时空分辨率,从而提升了视频的整体质量。实验结果表明,本文方法在峰值信噪比(Peak Signal-to-Noise Ratio,PSNR)上比现有二维和三维方法平均提高了8.5 dB,结构相似性指数(Structural Similarity Index Measure,SSIM)提升了15%,学习感知图像块相似度(Learned Perceptual Image Patch Similarity,LPIPS)减少了70%;同时,在表征三维结构保持能力的极平面图像上的峰值信噪比(Peak Signal-to-Noise Ratio on Epipolar Plane Image,EPI-PSNR)和极平面图像上的结构相似性指数(Structural Similarity Index Measure on Epipolar Plane Image,EPI-SSIM)指标上分别平均提升了12.77 dB和35.4%,说明所提方法能够更好地保持视差连续性和角度一致性。对于重建的视频序列,在基于运动的视频完整性评估(Motion-based Video Integrity Evaluation,MOVIE)指标上,本文方法平均得分为0.850,明显高于现有其他方法,显示出更接近真实图像的感知质量和更优的时域稳定性。这些结果表明,所提方法在提高视频重建质量、减少运动伪影和提升时域一致性方面具有显著优势,为集成成像动态3D显示的应用提供了一种有效的技术途径。

     

    Abstract: Integral imaging, recognized as one of the most promising autostereoscopic 3D video display technologies, continues to encounter a fundamental trade-off between frame rate and spatial resolution. This necessitates a trade-off between spatial clarity and temporal smoothness during dynamic video generation, thereby compromising the final visual effect and user experience. To address this limitation, this study introduces a bidirectional optical flow-guided spatio-temporal resolution enhancement framework, which enables dynamic 3D display of integral imaging with simultaneously high spatial resolution and frame rate. Inter-frame motion information is robustly estimated using bidirectional optical flow techniques, combined with a bidirectional optical flow-guided feature alignment module, thereby effectively mitigating boundary artifacts caused by parallax discontinuities in conventional methods, facilitating precise cross-frame fusion, and preserving fine-grained spatial details and parallax coherence. Additionally, the spatio-temporal super-resolution reconstruction module performs high-fidelity image reconstruction, substantially enhancing the spatio-temporal resolution of integral imaging videos, thereby improving overall visual quality. Experimental results demonstrate that the method significantly outperforms existing approaches, achieving an average improvement of 8.5 dB in Peak Signal-to-Noise Ratio (PSNR), a 15% increase in Structural Similarity Index Measure (SSIM), and a 70% reduction in Learned Perceptual Image Patch Similarity (LPIPS) compared with state-of-the-art 2D and 3D methods. Meanwhile, in terms of the PSNR on Epipolar Plane Image (EPI-PSNR) and the SSIM on Epipolar Plane Image (EPI-SSIM), which characterize the 3D structure preservation capacity, the proposed method achieves average improvements of 12.77 dB and 35.4%, respectively, thereby demonstrating its superior capability to preserve disparity continuity and angular consistency. For the reconstructed video sequences, under the Motion-based Video Integrity Evaluation (MOVIE) metric, the proposed method attains an average score of 0.850, which is substantially higher than that of other methods, indicating perceptual quality more closely approximating real-world imagery and enhanced temporal stability. These results collectively demonstrate that the proposed approach provides significant advantages in improving video reconstruction quality, mitigating motion artifacts, and strengthening temporal consistency, thereby offering a robust and effective technical framework for dynamic 3D display applications based on integral imaging.

     

/

返回文章
返回