STRUCTURE-BASED VALIDATION OF REINFORCEMENT-LEARNING-GENERATED JAK2 TYPE II INHIBITOR CANDIDATES: MOLECULAR DOCKING RESOLVES THE 1D REWARD-HACKING ARTIFACTS OF SINGLE-OBJECTIVE OPTIMIZATION

Authors

  • Xuelan Yang Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA, Shah Alam 40450, Malaysia https://orcid.org/0009-0006-0419-3913
  • Jasni Mohamad Zain Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA, Shah Alam 40450, Malaysia; Institute for Big Data Analytics and Artificial Intelligence, Universiti Teknologi MARA, Shah Alam 40450, Malaysia; Department of Informatics Engineering, Faculty of Computer Science, Universitas Brawijaya, Malang 65145, Indonesia https://orcid.org/0000-0003-2072-1510
  • Gembong Edhi Setyawan Department of Informatics Engineering, Faculty of Computer Science, Universitas Brawijaya, Jl. Veteran, Ketawanggede, Lowokwaru, Kota Malang, Jawa Timur 65145, Indonesia https://orcid.org/0000-0003-4989-8272
  • Diva Kurnianingtyas Department of Informatics Engineering, Faculty of Computer Science, Universitas Brawijaya, Jl. Veteran, Ketawanggede, Lowokwaru, Kota Malang, Jawa Timur 65145, Indonesia https://orcid.org/0000-0002-0865-7790

DOI:

https://doi.org/10.35631/JISTM.1144008

Keywords:

Chemical Space Exploration, De Novo Drug Design, Molecular Docking, Proximal Policy Optimization, Reinforcement Learning, Reward Hacking, StackRNN

Abstract

Reinforcement learning (RL) is proving to be a powerful way to tackle generative molecular design. Optimizing molecules using 1D surrogate reward functions, however, is frequently not sufficient to include the biophysical constraints that are necessary for successful drug discovery. Heuristic reward hacking is a problem with existing RL algorithms, which when fed with generated molecules can maximize the heuristic reward, but not produce a meaningful increase in pharmacological activity. This problem is addressed by Janus Kinase 2 (JAK2) Type II inhibitors which represent a challenging standard to meet in three-dimensional structure. We demonstrate that the same 1D surrogate can be used in two ways which are essentially different. Without any constraints, the unconstrained optimization using REINFORCE will lead to a “fragment trap”, which means molecules with low MW and lower than normal “drug-likeness” scores. Constrained optimization, on the other hand, with Proximal Policy Optimization (PPO) leads to molecules that are bigger and more hydrophobic than they need to be, taking advantage of the biases that are in the favor of molecular size and hydrophobicity. Both strategies do not result in the moderate complexity chemotypes essential for an effective JAK2 Type II inhibitor. Therefore, we propose as a orthogonal unsupervised validation step three-dimensional molecular docking as a structure-based surrogate independent of the surrogate reward. Docking re-centers molecular selection by giving preference to candidates in an intermediate molecular weight range (around 470-550 Da) typical of the deep DFG-out allosteric pocket, and systematically down-ranking undersized and oversized molecules. Empirically calculated docking scores are prone to bias based on the size and lipophilicity of the ligands so score differences should be taken with a pinch of salt and the docking is used as a confirmatory filter rather than an absolute affinity oracle. The results in this study show that the validation of molecular 3D structures should be added as another alignment tool in the generative molecular design process, and that the integration of multiple objectives in drug discovery with physically informed knowledge is desirable.

Downloads

Download data is not yet available.

References

Ai, C., Yang, H., Liu, X., Dong, R., Ding, Y., & Guo, F. (2024). MTMol-GPT: De novo multi-target molecular generation with transformer-based generative adversarial imitation learning. PLOS Computational Biology, 20(6), Article e1012229. https://doi.org/10.1371/journal.pcbi.1012229

Andraos, R., Qian, Z., Bonenfant, D., Rubert, J., Vangrevelinghe, E., Scheufler, C., Marque, F., Régnier, C. H., De Pover, A., Ryckelynck, H., Bhagwat, N., Koppikar, P., Goel, A., Wyder, L., Tavares, G., Baffert, F., Pissot-Soldermann, C., Manley, P. W., Gaul, C., ... Radimerski, T. (2012). Modulation of activation-loop phosphorylation by JAK inhibitors is binding mode dependent. Cancer Discovery, 2(6), 512–523. https://doi.org/10.1158/2159-8290.CD-11-0324

Bagal, V., Aggarwal, R., Vinod, P. K., & Priyakumar, U. D. (2022). MolGPT: Molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling, 62(9), 2064–2076. https://doi.org/10.1021/acs.jcim.1c00600

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Cho, K., van Merriënboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1724–1734). Association for Computational Linguistics. https://doi.org/10.3115/v1/D14-1179

De Cao, N., & Kipf, T. (2018). MolGAN: An implicit generative model for small molecular graphs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1805.11973

Gangwal, A., Ansari, A., Ahmad, I., Azad, A. K., Kumarasamy, V., Subramaniyan, V., & Wong, L. S. (2024). Generative artificial intelligence in drug discovery: Basic framework, recent advances, challenges, and opportunities. Frontiers in Pharmacology, 15, Article 1331062. https://doi.org/10.3389/fphar.2024.1331062

Gorantla, S. P., Oelschläger, L., Prince, G., Osius, J., Kolluri, S. B., Maluje, Y., Fähnrich, A., Ernst, N., Gulde, A. B., Ludwig, R. J., Gemoll, T., Fliedner, S., Walter, W., Haferlach, T., Gebauer, N., Busch, H., Duyster, J., & von Bubnoff, N. (2025). Ruxolitinib mediated paradoxical JAK2 hyperphosphorylation is due to the protection of activation loop tyrosines from phosphatases. Leukemia, 39(7), 1678–1691. https://doi.org/10.1038/s41375-025-02594-7

Hoogeboom, E., Satorras, V. G., Vignac, C., & Welling, M. (2022). Equivariant diffusion for molecule generation in 3D. In Proceedings of the 39th International Conference on Machine Learning (Vol. 162, pp. 8867–8887). PMLR. https://proceedings.mlr.press/v162/hoogeboom22a.html

Jayatunga, M. K. P., Xie, W., Ruder, L., Schulze, U., & Meier, C. (2022). AI in small-molecule drug discovery: A coming wave? Nature Reviews Drug Discovery, 21(3), 175–176. https://doi.org/10.1038/d41573-022-00025-1

Jing, B., Corso, G., Chang, J., Barzilay, R., & Jaakkola, T. (2022). Torsional diffusion for molecular conformer generation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2206.01729

Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In International Conference on Learning Representations. https://arxiv.org/abs/1412.6980

Kralovics, R., Passamonti, F., Buser, A. S., Teo, S. S., Tiedt, R., Passweg, J. R., Tichelli, A., Cazzola, M., & Skoda, R. C. (2005). A gain-of-function mutation of JAK2 in myeloproliferative disorders. The New England Journal of Medicine, 352(17), 1779–1790. https://doi.org/10.1056/NEJMoa051113

Krenn, M., Häse, F., Nigam, A. K., Friederich, P., & Aspuru-Guzik, A. (2020). Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4), Article 045024. https://doi.org/10.1088/2632-2153/aba947

Landrum, G. (2013). RDKit documentation (Release 2013.09.1) [Computer software documentation]. RDKit. https://www.rdkit.org/docs/

Lv, Y., Qi, J., Babon, J. J., Cao, L., Fan, G., Lang, J., Zhang, J., Mi, P., Kobe, B., & Wang, F. (2024). The JAK-STAT pathway: From structural biology to cytokine engineering. Signal Transduction and Targeted Therapy, 9, Article 221. https://doi.org/10.1038/s41392-024-01934-w

Mazuz, E., Shtar, G., Shapira, B., & Rokach, L. (2023). Molecule generation using transformers and policy gradient reinforcement learning. Scientific Reports, 13, Article 8799. https://doi.org/10.1038/s41598-023-35648-w

Meyer, S. C., Keller, M. D., Chiu, S., Koppikar, P., Guryanova, O. A., Rapaport, F., Xu, K., Manova, K., Pankov, D., O’Reilly, R. J., Kleppe, M., McKenney, A. S., Shih, A. H., Shank, K., Ahn, J., Papalexi, E., Spitzer, B., Socci, N., Viale, A., ... Levine, R. L. (2015). CHZ868, a type II JAK2 inhibitor, reverses type I JAK inhibitor persistence and demonstrates efficacy in myeloproliferative neoplasms. Cancer Cell, 28(1), 15–28. https://doi.org/10.1016/j.ccell.2015.06.006

Miao, Y., Virtanen, A., Zmajkovic, J., Hilpert, M., Skoda, R. C., Silvennoinen, O., & Haikarainen, T. (2024). Functional and structural characterization of clinical-stage Janus kinase 2 inhibitors identifies determinants for drug selectivity. Journal of Medicinal Chemistry, 67(12), 10012–10024. https://doi.org/10.1021/acs.jmedchem.4c00197

Mullard, A. (2017). The drug-maker’s guide to the galaxy. Nature, 549(7673), 445–447. https://doi.org/10.1038/549445a

Olivecrona, M., Blaschke, T., Engkvist, O., & Chen, H. (2017). Molecular de-novo design through deep reinforcement learning. Journal of Cheminformatics, 9, Article 48. https://doi.org/10.1186/s13321-017-0235-x

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., VanderPlas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. https://jmlr.org/papers/v12/pedregosa11a.html

Pereira, T., Abbasi, M., Ribeiro, B., & Arrais, J. P. (2021). Diversity oriented deep reinforcement learning for targeted molecule generation. Journal of Cheminformatics, 13, Article 21. https://doi.org/10.1186/s13321-021-00498-z

Philips, R. L., Wang, Y., Cheon, H., Kanno, Y., Gadina, M., Sartorelli, V., Horvath, C. M., Darnell, J. E., Jr., Stark, G. R., & O’Shea, J. J. (2022). The JAK-STAT pathway at 30: Much learned, much more to do. Cell, 185(21), 3857–3876. https://doi.org/10.1016/j.cell.2022.09.023

Popova, M., Isayev, O., & Tropsha, A. (2018). Deep reinforcement learning for de novo drug design. Science Advances, 4(7), Article eaap7885. https://doi.org/10.1126/sciadv.aap7885

Schulman, J., Levine, S., Moritz, P., Jordan, M. I., & Abbeel, P. (2015). Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning (Vol. 37, pp. 1889–1897). PMLR. https://proceedings.mlr.press/v37/schulman15.html

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1707.06347

Shi, C., Xu, M., Zhu, Z., Zhang, W., Zhang, M., & Tang, J. (2020). GraphAF: A flow-based autoregressive model for molecular graph generation. In International Conference on Learning Representations. https://openreview.net/forum?id=S1esMkHYPr

Ståhl, N., Falkman, G., Karlsson, A., Mathiason, G., & Boström, J. (2019). Deep reinforcement learning for multiparameter optimization in de novo drug design. Journal of Chemical Information and Modeling, 59(7), 3166–3176. https://doi.org/10.1021/acs.jcim.9b00325

Verstovsek, S., Mesa, R. A., Gotlib, J., Levy, R. S., Gupta, V., DiPersio, J. F., Catalano, J. V., Deininger, M., Miller, C., Silver, R. T., Talpaz, M., Winton, E. F., Harvey, J. H., Arcasoy, M. O., Hexner, E., Lyons, R. M., Paquette, R., Raza, A., Vaddi, K., ... Kantarjian, H. M. (2012). A double-blind, placebo-controlled trial of ruxolitinib for myelofibrosis. The New England Journal of Medicine, 366(9), 799–807. https://doi.org/10.1056/NEJMoa1110557

Weininger, D. (1988). SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28(1), 31–36. https://doi.org/10.1021/ci00057a005

You, J., Liu, B., Ying, R., Pande, V., & Leskovec, J. (2018). Graph convolutional policy network for goal-directed molecular graph generation. In Advances in Neural Information Processing Systems (Vol. 31, pp. 6410–6421). Curran Associates, Inc.

Zdrazil, B., Felix, E., Hunter, F. M. I., Manners, E. J., Blackshaw, J., Corbett, S., de Veij, M., Ioannidis, H., Mendez Lopez, D., Mosquera, J. F., Magarinos, M. P., Bosc, N., Arcila, R., Kizilören, T., Gaulton, A., Bento, A. P., Adasme, M. F., Monecke, P., Landrum, G. A., & Leach, A. R. (2024). The ChEMBL Database in 2023: A drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Research, 52(D1), D1180–D1192. https://doi.org/10.1093/nar/gkad1004

Zhou, Z., Kearnes, S., Li, L., Zare, R. N., & Riley, P. (2019). Optimization of molecules via deep reinforcement learning. Scientific Reports, 9, Article 10752. https://doi.org/10.1038/s41598-019-47148-x

Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., & Irving, G. (2019). Fine-tuning language models from human preferences [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1909.08593

Downloads

Published

2026-09-07

How to Cite

Xuelan , Y., Zain, J. M., Setyawan, G. E., & Kurnianingtyas, D. (2026). STRUCTURE-BASED VALIDATION OF REINFORCEMENT-LEARNING-GENERATED JAK2 TYPE II INHIBITOR CANDIDATES: MOLECULAR DOCKING RESOLVES THE 1D REWARD-HACKING ARTIFACTS OF SINGLE-OBJECTIVE OPTIMIZATION. JOURNAL INFORMATION AND TECHNOLOGY MANAGEMENT (JISTM), 11(44), 115–133. https://doi.org/10.35631/JISTM.1144008