Zhaoze Liu
- M.Sc. (Purdue University Northwest, 2021)
- B.Sc. (North University of China, 2019)
Topic
Statistical Learning and Inference from Imperfect Biological Data
Department of Mathematics and Statistics
Date & location
- Monday, August 24, 2026
- 8:00 A.M.
- Virtual Defence
Examining Committee
Supervisory Committee
- Dr. Xuekui Zhang, Department of Mathematics and Statistics, University of Victoria (Supervisor)
- Dr. John Taylor, Department of Biology, UVic (Co-Supervisor)
- Dr. Mary Lesperance, Department of Mathematics and Statistics, UVic (Member)
External Examiner
- Dr. You Liang, Department of Mathematics, Toronto Metropolitan University
Chair of Oral Examination
- Dr. Daniela Constantinescu, Department of Mechanical Engineering, UVic
Abstract
Biological, biomedical, and clinical studies often require inference or prediction from data that are incomplete, uncertain, or reduced representations of the scientific quantities of interest. Continuous outcomes may be available only through dichotomized or threshold-based summaries; time-to-event outcomes may be only partially observed because follow-up is incomplete; and genomic predictors may be sparse, noisy, or inferred from ultra-low-depth sequencing rather than directly observed with high confidence. These settings create a common statistical challenge: useful clinical and biological signal must be extracted from imperfect data while the observation process, target estimand, identifiability conditions, and uncertainty are made explicit.
This dissertation develops and applies statistical learning, inference, and genomic prediction methods for four distinct forms of imperfect data. The first two studies address imperfect phenotype and outcome information in clinical and biomedical settings. The final two studies address imperfect genotype information in sablefish (Anoplopoma fimbria) aquaculture. These studies are connected not by a single biological system or one common statistical model, but by a common inferential question: what information remains recoverable when observed data provide only an incomplete representation of the underlying process?
Chapter 2 considers an underlying continuous outcome that is observed only through binary or threshold-based summaries. A hierarchical binomial–probit framework is developed to estimate latent normal-distribution parameters from such reduced data. The framework considers fixed-mean and random-mean settings and evaluates likelihood-based, generalized linear-model, generalized linear mixed-model, and Bayesian approaches. The study clarifies when latent means, total variation, and within-study and between-study variance components are identifiable from threshold summaries, and when the available information is insufficient for the intended inferential target. The method is evaluated through simulation and a clinical biomarker application.
Chapter 3 addresses incomplete time-to-event follow-up under staggered recruitment. The chapter develops a landmark-indexed observed-support framework for interim assessment of crossing between original survival curves. The proposed method distinguishes crossing of the original survival curves from crossing of post-landmark conditional survival curves, which target different scientific quantities. It combines pre-landmark survival differences with post-landmark hazard information available at the interim analysis time and uses bootstrap stability, crossing-time confidence intervals, and observed-support diagnostics to quantify uncertainty. The framework is evaluated through simulation and a pseudo-prospective clinical-trial application.
Chapter 4 examines whether sparse genotype information can support useful association analysis and prediction in sablefish. A 300-locus SNP panel was used to study body-weight variation in independent harvest-age cohorts. The analysis recovered a strong sex-linked growth signal in the chromosome 14 sex-determining region, reflecting substantially faster female growth during ocean grow-out. Markers outside the sex-determining region retained additional predictive information, and an Elastic-Net model predicted body weight in an independent cohort with an R2 of approximately 0.49. This study demonstrates the practical utility of a low-density panel while showing its limited resolution for fine mapping, linkage-disequilibrium characterization, and broader trait discovery.
Chapter 5 extends the sablefish analysis from sparse markers to ultra-low-depth whole-genome sequencing with reference-panel-based genotype imputation. A reference panel of 65 fish sequenced at a mean depth of approximately 14× was used to impute genotypes in 1,380 farmed sablefish sequenced at a mean depth of approximately 0.4×. The resulting imputed marker resource supported genome-wide analyses of sex, body weight, and disease status while accounting for family structure. The analysis recovered the established chromosome 14 sex-determining region containing gsdf, supported within-cohort genomic prediction of sex, body weight, and disease status, and identified broad candidate regions for further study. However, moderate genotype concordance indicates that the imputed data are most appropriate for broad regional discovery and preliminary prediction rather than fine mapping or causal-variant inference.
Together, the four studies show that imperfect data can still support meaningful statistical inference and prediction when the source of information loss is acknowledged rather than ignored. Across thresholded outcomes, incomplete survival follow-up, sparse marker panels, and ultra-low-depth imputed genotypes, the central contribution of this dissertation is to demonstrate the importance of defining the target estimand, modelling the relationship between observed and underlying data, evaluating identifiability, and communicating uncertainty at a level consistent with the available information.