Bachelor dissertation day with Assoc. Prof. Tuyen Ngoc Truong at UMP
Ho Chi Minh City, Vietnam
I am currently a Master’s student following the Erasmus Mundus Joint Master’s ChEMoinformatics+ 2025-2027 (In Silico Drug Design, Paris-Milan-Paris) program. I am also a member of MedAI team.
My research interest lies in the application of deep learning to de novo drug design and discovery. My previous works focused on the implementation of molecular generative models, particularly variational autoencoders (VAEs). My bachelor’s thesis is a testament to this, as it focuses on the use of VAE models and Bayesian optimization, in conjunction with traditional drug discovery tools such as QSAR and Molecular docking, to identify new potential drug candidates with real-world applications. I seek to contribute to the advancement of this domain through novel approaches and applications.
Discovery of Vascular Endothelial Growth Factor Receptor 2 Inhibitors Employing Junction Tree Variational Autoencoder with Bayesian Optimization and Gradient Ascent
Gia-Bao Truong, Thanh-An Pham, Van-Thinh To, Hoang-Son Lai Le, Phuoc-Chung Van Nguyen, The-Chuong Trinh, Tieu-Long Phan, and Tuyen Ngoc Truong
In the development of anticancer medications, vascular endothelial growth factor receptor 2 (VEGFR-2), which belongs to the protein tyrosine kinase family, emerges as one of the most significant targets of interest. The ongoing Food and Drug Administration (FDA) approval of novel therapeutic medicines toward VEGFR-2 emphasizes the urgent need to discover sophisticated molecular structures that are capable of reliably limiting VEGFR-2 activity. Recognizing the huge potential of deep-learning-based molecular model advancements, we focused our study on exploring the chemical space to find small molecules potentially inhibiting VEGFR-2. To achieve this goal, we utilized the junction tree variational autoencoder in combination with two optimization approaches on the latent space: the local Bayesian optimization on the initial data set and the gradient ascent on nine FDA-approved drugs targeting VEGFR-2. The optimization results yielded a set of 493 uncharted small molecules. Quantitative structure–activity relationship (QSAR) models and molecular docking were used to assess the generated molecules for their inhibitory potential using their predicted pIC50 and binding affinity. The QSAR model constructed on RDK7 fingerprints using the CatBoost algorithm achieved remarkable coefficients of determination (R2) of 0.792 ± 0.075 and 0.859 with respect to internal and external validation. Molecular docking was implemented using the 4ASD complex with optimistic retrospective control results (the ROC-AUC value was 0.710 and the binding activity threshold was −7.90 kcal/mol). Newly generated molecules possessing acceptable results corresponding to both assessments were shortlisted and checked for interactions with the protein at the binding site on important residues, including Cys919, Asp1046, and Glu885.
@article{vegfr2,title={Discovery of Vascular Endothelial Growth Factor Receptor 2 Inhibitors Employing Junction Tree Variational Autoencoder with Bayesian Optimization and Gradient Ascent},author={Truong, Gia-Bao and Pham, Thanh-An and To, Van-Thinh and Lai Le, Hoang-Son and Van Nguyen, Phuoc-Chung and Trinh, The-Chuong and Phan, Tieu-Long and Truong, Tuyen Ngoc},keywords={Vascular endothelial growth factor receptor 2, junction tree variational autoencoder, bayesian optimization, gradient ascent},doi={10.1021/acsomega.4c07689},url={https://pubs.acs.org/doi/10.1021/acsomega.4c07689},journal={ACS Omega},year={2024},}
JCIM
KGG: Knowledge-Guided Graph Self-Supervised Learning to Enhance Molecular Property Predictions
Van-Thinh To, Phuoc-Chung Van Nguyen, Gia-Bao Truong, Tuyet-Minh Phan, Tieu-Long Phan, Rolf Fagerberg, Peter F. Stadler, and Tuyen Ngoc Truong
Journal of Chemical Information and Modeling, 2025
Molecular property prediction has become essential in accelerating advancements in drug discovery and materials science. Graph Neural Networks have recently demonstrated remarkable success in molecular representation learning; however, their broader adoption is impeded by two significant challenges: (1) data scarcity and constrained model generalization due to the expensive and time-consuming task of acquiring labeled data and (2) inadequate initial node and edge features that fail to incorporate comprehensive chemical domain knowledge, notably orbital information. To address these limitations, we introduce a Knowledge-Guided Graph (KGG) framework employing self-supervised learning to pretrain models using orbital-level features in order to mitigate reliance on extensive labeled data sets. In addition, we propose novel representations for atomic hybridization and bond types that explicitly consider orbital engagement. Our pretraining strategy is cost efficient, utilizing approximately 250,000 molecules from the ZINC15 data set, in contrast to contemporary approaches that typically require between two and ten million molecules, consequently reducing the risk of potential data contamination. Extensive evaluations on diverse downstream molecular property data sets demonstrate that our method significantly outperforms state-of-the-art baselines. Complementary analyses, including t-SNE visualizations and comparisons with traditional molecular fingerprints, further validate the effectiveness and robustness of our proposed KGG approach. The key advantages of KGG are its data efficiency and architectural versatility, driven by orbital-informed representations. By distilling essential chemical knowledge from modest corpora, it avoids extensive pretraining and excels in low-data fine-tuning, providing a robust and chemically meaningful foundation for diverse GNN architectures.
@article{kgg,author={To, Van-Thinh and Van Nguyen, Phuoc-Chung and Truong, Gia-Bao and Phan, Tuyet-Minh and Phan, Tieu-Long and Fagerberg, Rolf and Stadler, Peter F. and Truong, Tuyen Ngoc},title={KGG: Knowledge-Guided Graph Self-Supervised Learning to Enhance Molecular Property Predictions},journal={Journal of Chemical Information and Modeling},volume={65},number={18},pages={9443-9458},year={2025},doi={10.1021/acs.jcim.5c01068},note={PMID: 40916452},keywords={Drug discovery, graph neural networks, knowledge graph, self-supervised learning, orbital information},}
JCAMD
Synergy of advanced machine learning and deep neural networks with consensus molecular docking for virtual screening of anaplastic lymphoma kinase inhibitors
The-Chuong Trinh, Tieu-Long Phan, Van-Thinh To, Thanh-An Pham, Gia-Bao Truong, Lai Hoang Son Le, Xuan-Truc Dinh Tran, and Tuyen Ngoc Truong
Journal of Computer-Aided Molecular Design, Sep 2025
This study addresses the urgent need for an AI model to predict Anaplastic Lymphoma Kinase (ALK) inhibitors for Non-Small Cell Lung Cancer treatment, targeting the ALK-positive mutation. With only five Food and Drug Administration approved ALK inhibitors currently available, effective drugs remain in demand. Leveraging machine learning (ML) and deep learning (DL), our research accelerates the precise screening of novel ALK inhibitors using both ligand-based and structure-based approaches. In ligand-based approach, an ensemble voting model comprising three base learners to classify potential ALK inhibitors, achieving promising retrospective validation results. Notably, the ML-based XGBoost algorithm exhibited compelling results with external validation (EV)-f1 score of 0.921, EV-Average Precision (AP) of 0.961, cross-validation (CV)-f1 score of and CV-AP of . Besides, the DL-based Artificial Neural Network (ANN) model demonstrated comparative performance with EV-f1 score of 0.930, EV-AP of 0.955, CV-f1 score of and CV-AP of . For structure-based approach, an XGBoost consensus docking model utilized scores from three molecular docking programs (GNINA 1.0, Vina-GPU 2.0, and AutoDock-GPU) as features. Combining these two approaches, we virtually screened 120,571 compounds, identifying three promising ALK inhibitors, CHEMBL1689515, CHEMBL2380351, and CHEMBL102714, that bind to the protein’s pocket and establish hydrophobic contacts in the hinge region through their ketone groups, resembling Alectinib’s interaction. Comparative analysis revealed traditional ML models outperformed Graph Neural Networks (GNN), highlighting the critical role of feature engineering and dataset size importance. The study recommends further in vitro testing to validate the prospective screening performance of these models. A graphical user interface is available at https://huggingface.co/spaces/thechuongtrinh/ALK_inhibitors_classification.
@article{alk,author={Trinh, The-Chuong and Phan, Tieu-Long and To, Van-Thinh and Pham, Thanh-An and Truong, Gia-Bao and Le, Lai Hoang Son and Tran, Xuan-Truc Dinh and Truong, Tuyen Ngoc},title={Synergy of advanced machine learning and deep neural networks with consensus molecular docking for virtual screening of anaplastic lymphoma kinase inhibitors},journal={Journal of Computer-Aided Molecular Design},year={2025},month=sep,day={15},volume={39},number={1},pages={79},issn={1573-4951},doi={10.1007/s10822-025-00657-6},keywords={ALK, computer-aided drug design, artificial intelligence, machine learning, benchmarking, consensus molecular docking},}