Artificial intelligence increasingly influences decisions in domains where errors, latent bias, misunderstandings, and misplaced confidence can have serious ramifications. However, accountability in these systems remains difficult because it depends on more than model performance; accountability also depends on whether fairness is defined in a principled way, whether evaluation metrics are well-founded and remain reliable across groups and settings, and whether uncertainty is properly communicated. This dissertation examines accountability in artificial intelligence through four connected problems: how to define fairness, how model evaluation can become unreliable, when a metric is well-founded (measures the correct underlying thing and is axiomatically justified), and how uncertainty should constrain the conclusions drawn from metric values, group comparisons, and benchmark rankings.
Firstly, this dissertation develops the Objective Fairness Index (OFI), a fairness metric grounded in the legal principles of objective testing. OFI measures disparity through differences in marginal benefit (a numerical estimate of "what happened" versus "what should have happened"). This allows fairness assessments to distinguish algorithmic discrimination from broader social disparity and provides a practical diagnostic for determining whether a model or decision procedure apparently misallocates benefit. OFI is applied to binary-classification problems such as recidivism prediction, employment classification, and activity recognition in clinical datasets with synthetic data generation.
Next, this dissertation studies fairness metrics through the lens of social choice theory. By treating fairness metrics as aggregation rules over groups, we develop a taxonomy of parity-gap metrics, adapt social-choice-style axioms to binary classification, and identify important pathologies and trade-offs, including aggregation reversals and an impossibility result showing that Consistency and Participation cannot be jointly guaranteed across the taxonomy.
The dissertation then studies sample size-induced bias in classification metrics. Here, we discover that many confusion-matrix-based metrics become unstable or jagged for small sample sizes, which can distort group comparisons. To address this problem, we introduce the Metric Alignment Trial for Checking Homogeneity (MATCH) test for assessing whether an observed subgroup score is probabilistically consistent with a reference distribution, and Cross-Prior Smoothing (CPS) for improving metric stability and reliability.
In evaluating time-series similarity measures, this dissertation rigorously checks if the literature's default, leave-one-out 1-nearest-neighbor assessed by accuracy, is appropriate. We find that it generalizes poorly to other objectives and that the Matthews Correlation Coefficient often provides sharper statistical discrimination than accuracy. Additionally, our objective-informed benchmarking findings suggest that the best time-series similarity measure is objective-dependent, and the statistical uncertainties often provide multiple "best" options.
Taken together, these contributions advance accountable artificial intelligence. By introducing a new method for fairness assessment, formalizing desirable metric properties, improving the reliability of model evaluation, and revisiting time-series benchmarking approaches, this dissertation advances fairness formalizations, metric reliability, and interpretable uncertainty not only in artificial intelligence but also in its evaluations.