Defense Date

2026

Document Type

Thesis

Degree Name

Master of Science

Department

Bioinformatics

First Advisor

Preetam Ghosh

Second Advisor

Michael Rosenberg

Third Advisor

LaMont Cannon

Abstract

Breast cancer is a heterogeneous disease driven by complex genomic and transcriptomic alterations, making accurate phenotype prediction a challenging task. Recent advances in genomic foundation models and representation learning provide new opportunities to extract biologically meaningful features directly from sequencing data. The objective of this study was to develop and evaluate a multimodal framework integrating whole genome sequence (WGS) and RNA sequencing (RNA-Seq) embeddings for breast cancer classification. DNABERT-2, a DNA foundation model, was evaluated for its ability to generate genomic representations from The Cancer Genome Atlas breast cancer (TCGA-BRCA) WGS data, and the impact of downstream fine-tuning on embedding performance was assessed. Multiple autoencoder architectures were investigated to generate low-dimensional RNA-Seq embeddings of paired TCGA-BRCA data, and principal component analysis (PCA) was evaluated as a baseline dimensionality reduction approach. WGS, RNA-Seq, and integrated multimodal embeddings were subsequently evaluated for binary healthy/disease classification and multi-class PAM50 molecular subtype classification. DNABERT-2-derived WGS embeddings demonstrated strong performance for binary disease classification but provided limited predictive information for PAM50 subtype classification. RNA-Seq embeddings consistently outperformed WGS embeddings for multi-class classification, while PCA achieved comparable or superior performance to more complex autoencoder architectures. Furthermore, integrating WGS and RNA-Seq embeddings resulted in only marginal improvements compared to RNA-Seq embeddings alone, suggesting that the current WGS representations provided limited complementary information for breast cancer subtype prediction. Collectively, these findings demonstrate that increased model complexity and multimodal integration do not inherently improve predictive performance and highlight the importance of selecting biologically informative representations for genomic phenotype prediction.

Rights

© Leiliani Clark

Is Part Of

VCU University Archives

Is Part Of

VCU Theses and Dissertations

Date of Submission

8-4-2026

Share

COinS