Uncertainty-Aware Classification via Neutral Zones: A Statistical Framework
Mengqi Yin
Doctor of Philosophy (PhD), Washington State University
2026
Files and links (1)
pdf
Dissertation_draft_3rd ETD_sub
Embargoed Access, Embargo ends: 01/22/2027
Abstract
The focus of this dissertation is incorporating stochasticity in the form of uncertainty in some of the common models used in Machine Learning (ML) and Artificial Intelligence (AI). Specifically, it looks at the role of statistics in AI through the Population (P), Question (Q), Representation (R), and Scrutiny (S) framework suggested by \citet{yu2018}. The overarching aim is trying to scrutinize the performance of the standard black-box models. We take a cursory look at Representation via simulation experiments with multiple factors to see how much of a part representation plays in how well the AI systems generalize beyond their training data. The results identified feature strength as the strongest factor with higher accuracy and lower false negative rates. Another factor that affected performance was whether training and testing data come from the same population. Size of training sample and class proportions had minor effects. Overall, these findings highlight that reliable classification depends less on the amount of data and more on whether the data properly represents the population. The theoretical contribution of this work lies in proposing neutral zones for post AI/ML classification. We propose methods that incorporate uncertainty in high-stakes decision settings that rely on AI-based classifiers where there is a human cost associated with misclassifications. Theoretically, we address the challenge by re-framing classification as a hypothesis testing problem and introduce a post hoc statistical scrutiny layer using neutral zones. We develop methods that allow us to incorporate the philosophy of hypothesis testing problem which is asymmetric by construction and applying it to a symmetric classification problem. Our methodological framework is built on three statistical concepts: cardinality, reliability, and calibration. These components are combined to develop a hybrid method for binary classification, allowing uncertainty to be incorporated without sacrificing overall accuracy. The method is then extended to multi-class problems using results from ranking and selection. For the multi-class problem, we look at the top two classes in terms of rank and probability. This then reduces the multi-class problem to a two-class problem. This allows us to use the hybrid method proposed for the two-class problem with a few modifications. We also simplified the problem further by looking the class with the highest probability and using the ideas of cardinality, reliability, and conformal prediction to propose a method that is computationally a lot easier to apply and can be easily used in situations with many classes. Performance is evaluated using Adeno-Associated Virus data \citep{Khan2022} and simulation studies. Results show that the hybrid method improves uncertainty handling in the binary setting, while ADEBO methods provide a better balance between accuracy and uncertainty in multi-class problems. In conclusion, our work builds a pathway for post hoc statistical scrutiny for improving the reliability, adaptability, and interpretability of AI systems.
Metrics
1 Record Views
Details
Title
Uncertainty-Aware Classification via Neutral Zones: A Statistical Framework
Creators
Mengqi Yin
Contributors
Nairanjana Dasgupta (Advisor)
Dean Johnson (Committee Member)
Yuan Wang (Committee Member)
Awarding Institution
Washington State University
Academic Unit
Department of Mathematics and Statistics
Theses and Dissertations
Doctor of Philosophy (PhD), Washington State University