Modern statistical methods provide powerful tools for studying complex data, but theirsuccessful application often depends on how well uncertainty, structure, and data limitations
are addressed. In many real-world problems, meaningful inference requires both methodological
development and careful attention to the process by which the data are observed.
This dissertation presents two studies motivated by these challenges.
The first study focuses on gender-related differences in pay rates using survey data andan attention to the methodological challenges created by missing responses to wage-related
questions. Because pay rate is a sensitive topic, survey participants may be unwilling to disclose
this information, and such nonresponse can introduce substantial bias into the analysis
if it is not properly addressed. The goal of this study is therefore twofold: first, to examine
pay rate differences between female and male participants for statistical consultants while
accounting for important background characteristics, and second, to better understand the
missingness mechanism associated with pay-related survey question. To accomplish these
objectives, we designed and administered a survey, obtained Institutional Review Board
(IRB) approval, and collected responses from participants across relevant demographic and
employment-related characteristics. Exploratory data analysis was conducted to study the
structure of the observed data and the pattern of missing responses. Multiple imputation
techniques were then applied to address the missing-data problem before proceeding to the
main analysis. To estimate the effect of gender on pay rate while reducing the impact of confounding
variables, several causal inference methods based on matching were implemented
and compared. In addition, the study examines the likelihood that participants decline to
answer pay-related questions, providing further insight into response behavior in sensitive
survey settings. The results of this study contribute to both substantive and methodological
understanding. Substantively, they provide evidence regarding pay rate disparities in this
specific labor market. Methodologically, they illustrate how combining missing-data methods
with causal inference techniques can improve the validity of conclusions drawn from
survey data when sensitive questions are subject to nonresponse.
The second study analyzes distributions of ordinal patterns using null models for timeseries data. Permutation entropy has emerged as a widely used statistical measure for assessing
the complexity of time series, with applications spanning diverse fields such as biology,
economics, and physics. Recent work has applied divergence measures, including Kullback-
Leibler divergence and Jensen-Shannon divergence, to quantify the similarity between time
series based on ordinal patterns, and to measure the deviation from generative models. This
approach also allows for creating meaningful embeddings of the data. While early approaches
implicitly assumed a uniform null model for the underlying distribution of patterns, many
real-world applications involve more complex distributions and are better modeled with random
walks. In this study, we present an analysis of several null models for generating time
series data, from both theoretical and empirical perspectives. We successfully derive theorems
describing the behavior of random walks with uniformly and normally distributed
steps, as well as introducing a novel random walk null model based on transition matrices
inferred from real-world data. We show that this data-driven approach allows for a more
accurate representation of empirical distributions. We also demonstrate the practical utility
of these methods by applying them to real-world datasets from economics and other domains,
highlighting the effectiveness of divergence measures in capturing the complexity of
empirical time series data. We conclude by exploring a threshold-based method to capture
structure of subsets of the data with a strong suggestion of a 20% threshold, for aligning
the distribution of patterns below the threshold with those from the complete data. This
approach allows us to learn the ordinal pattern distributions, even in the presence of noise,
with a smaller data set and less computation cost when analyzing high-dimensional, large,
or more complex series.