Differentiable Diffusion for Dense Depth Estimation from Multi-view Images

Computer Vision and Pattern Recognition (CVPR) 2021

Numair Khan1, Min H. Kim2, James Tompkin1
1 2

Abstract

We present a method to estimate dense depth by optimizing a sparse set of points such that their diffusion into a depth map minimizes a multi-view reprojection error from RGB supervision. We optimize point positions, depths, and weights with respect to the loss by differential splatting that models points as Gaussians with analytic transmittance. Further, we develop an efficient optimization routine that can simultaneously optimize the 50k+ points required for complex scene reconstruction. We validate our routine using ground truth data and show high reconstruction quality. Then, we apply this to light field and wider baseline images via self supervision, and show improvements in both average and outlier error for depth maps diffused from inaccurate sparse points. Finally, we compare qualitative and quantitative results to image processing and deep learning methods.

Results

Reference RGB view of the dinosaur scene Our diffused dense depth for the dinosaur scene
Reference RGB view of the toy digger scene Our diffused dense depth for the toy digger scene

Reference RGB view (left) and our diffused dense depth (right) for two scenes.

Video

Download video (MP4, 20 MB)

Citation

@inproceedings{Khan_2021,
    author    = {Numair Khan and Min H. Kim and James Tompkin},
    title     = {Differentiable Diffusion for Dense Depth Estimation from Multi-view Images},
    booktitle = {2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    year      = {2021},
    pages     = {8908--8917},
    doi       = {10.1109/CVPR46437.2021.00880}
}

Acknowledgements

We thank the reviewers for their detailed feedback. Numair Khan thanks an Andy van Dam PhD Fellowship, and Min H. Kim acknowledges the support of Korea NRF grant (2019R1A2C3007229).