Confirmatory body fluid identification via the detection of specific messenger RNA (mRNA) species is important in circumstances where a conventional test is not available for the questioned fluid type (such as vaginal material). In practice, samples submitted for testing often contain two or more body fluid types from different donors (i.e., mixtures). Molecular approaches such as endpoint reverse-transcription PCR (RT-PCR) or real-time quantitative reverse-transcription PCR (RT-qPCR) may be used for confirmatory body fluid identification.
The interpretation of mRNA profiling data has been a topic of recent interest, with two main approaches proposed: categorical and probabilistic. Categorical methods, which may use a scoring system or threshold-based approach, provide simple decision rules but often fail to utilise all relevant information. Probabilistic methods can incorporate known information and reflect uncertainty as part of classification, aligning more closely with biological reality. Various groups have developed probabilistic models for single-source and mixture profiles using endpoint RT-PCR and sequencing data. To our knowledge, our previous work represents the first application of such approaches to RT-qPCR data [1,2].
There remains a significant challenge in the interpretation of complex mRNA profiles, particularly for body fluids that are not represented in the reference dataset. Current methods are limited in their ability to generalise to unseen or underrepresented fluid types. Addressing these limitations is critical for improving forensic reporting.
We will discuss multi-class machine learning classifiers for predicting the components of body fluids in biological, not in-silico, mixed samples [CL2.1], extending prior work on single-source body fluids. We will also present methods for modelling and predicting an “unknown” category, composed of samples that do not exhibit specific features in the data, using two complementary approaches. The first approach utilises metric learning to identify samples that deviate from known class structures in feature space, while the second uses a leave-one-type-out simulation to explicitly model missing data scenarios.