This website stores cookies on your computer. These cookies are used to collect information about how you interact with our website and allow us to remember your browser. We use this information to improve and customize your browsing experience, for analytics and metrics about our visitors both on this website and other media, and for marketing purposes. By using this website, you accept and agree to be bound by UVic’s Terms of Use for web and social media privacy.  If you do not agree to the above, you can configure your browser’s setting to “do not track.”

Skip to main content

Abida Sultana

  • BSc (Rajshahi University of Engineering and Technology, 2022)

Notice of the Final Oral Examination for the Degree of Master of Applied Science

Topic

Object-wise Metric Distance Estimation from a Single RGB Image via Semantic and Geometric Reasoning

Department of Electrical and Computer Engineering

Date & location

  • Friday, May 1, 2026

  • 10:00 A.M.

  • Virtual Defence

Reviewers

Supervisory Committee

  • Dr. Hong-Chuan Yang, Department of Electrical and Computer Engineering, University of Victoria (Supervisor)

  • Dr. Kin Fun Li, Department of Electrical and Computer Engineering, UVic (Member) 

External Examiner

  • Dr. Kui Wu, Department of Computer Science, University of Victoria 

Chair of Oral Examination

  • Dr. Justin Albert, Department of Physics and Astronomy, UVic

     

Abstract

Estimating metric object distance from a single RGB image is challenging because monocular depth does not provide an absolute scale. Existing solutions either require active sensors such as LiDAR or stereo, rely on monocular depth that remains scale-ambiguous, or use implicit vision-language reasoning that can be unstable for precise measurement. This thesis proposes a semantic–geometric pipeline for recovering metric scale by combining open-vocabulary object grounding and segmentation, label normalization, monocular depth, and camera cues. Object-centric 3D points are reconstructed from the predicted depth, an oriented 3D bounding box is fitted to estimate object dimensions, and real-world size priors are used to compute a scale factor that converts relative depth into absolute distance. The proposed method is evaluated on HOT3D, ScanNet, ARKitScenes, and a custom iPhone dataset, achieving Multi-Threshold Relative Accuracy (MRA) values of 68.85%, 88.30%, 75.12%, and 89.85%, respectively, under the per-frame average mean-distance strategy. The results show that frame-level averaging improves stability by reducing the influence of instance-level outliers. The main limitations of the approach are its dependence on segmentation and depth quality, sensitivity to canonical size priors for categories with high size variation, possible instability under occlusion or truncation, and relatively high processing time. Future work includes more robust scale estimation, adaptive size priors, improved object fitting, the use of consecutive frames for temporal consistency, and pipeline optimization for lower latency.