From IR Images to Point Clouds to Pose: Point Cloud-Based AR Glasses Pose Estimation

Ahmet Firintepe; Carolin Vey; Stylianos Asteriadis; Alain Pagani; Didier Stricker

doi:10.3390/jimaging7050080

From IR Images to Point Clouds to Pose: Point Cloud-Based AR Glasses Pose Estimation

Ahmet Firintepe^*, Carolin Vey, Stylianos Asteriadis, Alain Pagani, Didier Stricker

^*Corresponding author for this work

Research output: Contribution to journal › Article › Academic › peer-review

Abstract

In this paper, we propose two novel AR glasses pose estimation algorithms from single infrared images by using 3D point clouds as an intermediate representation. Our first approach "PointsToRotation" is based on a Deep Neural Network alone, whereas our second approach "PointsToPose" is a hybrid model combining Deep Learning and a voting-based mechanism. Our methods utilize a point cloud estimator, which we trained on multi-view infrared images in a semi-supervised manner, generating point clouds based on one image only. We generate a point cloud dataset with our point cloud estimator using the HMDPose dataset, consisting of multi-view infrared images of various AR glasses with the corresponding 6-DoF poses. In comparison to another point cloud-based 6-DoF pose estimation named CloudPose, we achieve an error reduction of around 50%. Compared to a state-of-the-art image-based method, we reduce the pose estimation error by around 96%.

Original language	English
Article number	80
Number of pages	18
Journal	Journal of Imaging
Volume	7
Issue number	5
DOIs	https://doi.org/10.3390/jimaging7050080
Publication status	Published - May 2021

Keywords

computer vision
augmented reality
object pose estimation
point clouds
deep learning

Access to Document

10.3390/jimaging7050080Licence: CC BY

Cite this

@article{364a3e3c8c0742be8095c61b8d5ec3f1,

title = "From IR Images to Point Clouds to Pose: Point Cloud-Based AR Glasses Pose Estimation",

abstract = "In this paper, we propose two novel AR glasses pose estimation algorithms from single infrared images by using 3D point clouds as an intermediate representation. Our first approach {"}PointsToRotation{"} is based on a Deep Neural Network alone, whereas our second approach {"}PointsToPose{"} is a hybrid model combining Deep Learning and a voting-based mechanism. Our methods utilize a point cloud estimator, which we trained on multi-view infrared images in a semi-supervised manner, generating point clouds based on one image only. We generate a point cloud dataset with our point cloud estimator using the HMDPose dataset, consisting of multi-view infrared images of various AR glasses with the corresponding 6-DoF poses. In comparison to another point cloud-based 6-DoF pose estimation named CloudPose, we achieve an error reduction of around 50%. Compared to a state-of-the-art image-based method, we reduce the pose estimation error by around 96%.",

keywords = "computer vision, augmented reality, object pose estimation, point clouds, deep learning",

author = "Ahmet Firintepe and Carolin Vey and Stylianos Asteriadis and Alain Pagani and Didier Stricker",

note = "Publisher Copyright: {\textcopyright} 2021 by the authors. Licensee MDPI, Basel, Switzerland.",

year = "2021",

month = may,

doi = "10.3390/jimaging7050080",

language = "English",

volume = "7",

journal = "Journal of Imaging",

publisher = "MDPI",

number = "5",

}

TY - JOUR

T1 - From IR Images to Point Clouds to Pose

T2 - Point Cloud-Based AR Glasses Pose Estimation

AU - Firintepe, Ahmet

AU - Vey, Carolin

AU - Asteriadis, Stylianos

AU - Pagani, Alain

AU - Stricker, Didier

PY - 2021/5

Y1 - 2021/5

N2 - In this paper, we propose two novel AR glasses pose estimation algorithms from single infrared images by using 3D point clouds as an intermediate representation. Our first approach "PointsToRotation" is based on a Deep Neural Network alone, whereas our second approach "PointsToPose" is a hybrid model combining Deep Learning and a voting-based mechanism. Our methods utilize a point cloud estimator, which we trained on multi-view infrared images in a semi-supervised manner, generating point clouds based on one image only. We generate a point cloud dataset with our point cloud estimator using the HMDPose dataset, consisting of multi-view infrared images of various AR glasses with the corresponding 6-DoF poses. In comparison to another point cloud-based 6-DoF pose estimation named CloudPose, we achieve an error reduction of around 50%. Compared to a state-of-the-art image-based method, we reduce the pose estimation error by around 96%.

AB - In this paper, we propose two novel AR glasses pose estimation algorithms from single infrared images by using 3D point clouds as an intermediate representation. Our first approach "PointsToRotation" is based on a Deep Neural Network alone, whereas our second approach "PointsToPose" is a hybrid model combining Deep Learning and a voting-based mechanism. Our methods utilize a point cloud estimator, which we trained on multi-view infrared images in a semi-supervised manner, generating point clouds based on one image only. We generate a point cloud dataset with our point cloud estimator using the HMDPose dataset, consisting of multi-view infrared images of various AR glasses with the corresponding 6-DoF poses. In comparison to another point cloud-based 6-DoF pose estimation named CloudPose, we achieve an error reduction of around 50%. Compared to a state-of-the-art image-based method, we reduce the pose estimation error by around 96%.

KW - computer vision

KW - augmented reality

KW - object pose estimation

KW - point clouds

KW - deep learning

U2 - 10.3390/jimaging7050080

DO - 10.3390/jimaging7050080

M3 - Article

C2 - 34460676

VL - 7

JO - Journal of Imaging

JF - Journal of Imaging

IS - 5

M1 - 80

ER -