Mixed STR profiles are common in forensic analysis, but their formation is influenced by the combined effects of multiple factors, including amplification randomness, variations in contributor proportions, and noise superposition. Traditional statistical models struggle to accurately capture their underlying generation mechanisms. To address this, this paper constructs a generative model incorporating an inference mechanism. It performs structured encoding of raw peak data into fixed-length sequences, mapping peak positions, peak heights, and noise information into multi-channel feature representations. Based on this, the model uses these feature sequences and prior knowledge of the number of contributors as conditional inputs to output the reconstructed peak height distribution at corresponding loci and the estimated contribution ratios of each contributor. The core of the network consists of a conditional generation module and a parameter inference module. The former employs an encoder-decoder architecture to learn peak shape distributions, while the latter uses attention weights to attribute different peak signals and constrains the update of contribution ratios. In modeling amplification fluctuations, random noise and systematic offset are parameterized separately; hierarchical variables are introduced to decompose peak height perturbations, and constraint optimization is performed via a joint loss function during training. Experimental results demonstrate that this method can stably reconstruct mixed STR spectra and reveal the correspondence between peak height variations and contributor composition, providing a more interpretable modeling approach for analyzing the formation mechanisms of mixed STR spectra.