Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Algorithmic Data Science for Computational Drug Discovery

Projekt: Forschungsförderung

Projektdetails

Abstract

Drug discovery is a lengthy and expensive process that suffers from a small and decreasing success rate. Finding a drug in the search space of possible molecules, which is estimated to contain up to 10^60 structures, is often compared to "finding a needle in a haystack". Although it is possible to perform millions of pharmacological tests in a reasonable time using automated high-throughput screening devices, only a negligible fraction of the chemical space can be synthesized and tested experimentally. Data science is of major importance to exploit the available information in order to direct the search for a new drug candidate to the relevant chemical subspace. The amount of available data in the life sciences grows rapidly, including the outcome of high-throughput experiments as well as data on target proteins, their structure and interaction. On the one hand, this poses new algorithmic challenges regarding the scalability of data mining and machine learning methods to very large molecular data sets. On the other hand, there is well justified hope to obtain more realistic models from the analysis of very large quantities of data with methods tailored to this domain. If this is the case, such methods will allow for a greater automation of the drug discovery process reducing the required time and costs drastically.

The goal of the project is to achieve this through the following work plan. First, we will develop efficient methods for the graph based similarity search in large molecular databases using index data structures. Our methods will support graphs with complex vertex and edge annotations to incorporate conformational features regarding the shape and flexibility. On this basis we will develop algorithms for mining relevant substructures, either preserving a certain biological property or having an extraordinarily strong effect on the property. These substructures will be of interest for direct inspection by domain specialist, but also for the machine learning approaches we will develop. For the former we will provide visual analysis modules extending the KNIME analytics platform. We will analyze existing machine learning approaches for graphs regarding their ability to learn successfully in scenarios with equivalent substructures and will identify their weaknesses. Based on this analysis we will develop new machine learning approaches, which will overcome these issues, e.g., by providing this information explicitly as part of the input or by learning structural equivalences in an end-to-end fashion. Finally, we will study generative models for molecular graphs, which allow to construct new compounds with desired properties from structural building blocks starting with their core structure. The developed methods will be released as KNIME extensions for use in chemical data mining workflows and evaluated in real-world pharmacoinformatics tasks.

The project is funded by WWTF (Vienna Science and Technology Fund).
StatusLaufend
Tatsächlicher Beginn/ -es Ende1/05/2030/11/28
  • A Computational Community Blind Challenge on Pan-Coronavirus Drug Discovery Data

    MacDermott-Opeskin, H., Scheen, J., Wognum, C., Horton, J. T., West, D., Payne, A. M., Castellanos, M. A., Colby, S., Griffen, E., Cousins, D., Stacey, J., Reid, L., Aschenbrenner, J. C., Fearon, D., Balcomb, B., Marples, P., Tomlinson, C. W. E., Lithgo, R., Godoy, A. S. & Winokan, M. &106 mehr, Barr, H., Lahav, N., Lavi, M., Duberstein, S., Cohen, G., Fate, G., Lefker, B., Robinson, R., Szommer, T., Lynch, N., Minh, D. D. L., La, V. N. T., Kang, L., Huddleston, K., Renslow, R., Tollefson, M., Walters, W. P., Xu, C., Hsu, J., St-Laurent, J., Etsmoberg, H., Zhu, L., Quirke, A., Abdul Haleem, M. I., Alibay, I., Baid, G., Birnbaum, B., Bishop, K. P., Bohorquez, H., Bose, A., Brown, C. J., Burns, J., Cai, L., Cedeno, R., de Cesco, S., Chupakhin, V., Clark, F., Cole, D. J., Corbi-Verge, C., Danial, M., Davi, A., Dehaen, W., Doering, N. P., Dougha, A., Dréanic, M. P., Eakin, B., Ehrlich, A., Elijosius, R., Fülöp, J., Gitter, A., Goossens, K., Gu, Y., Head-Gordon, T., Hoffer, L., Hofmans, J., Jiang, E., Kaminow, B., Khosravi, S., Khoualdi, A. F., Lenselink, E. B., Liu, Z., Liu, Y., Liu, S., Ma, Y., Maher, P., Mayer, I., Mendez-Lucio, O., Mey, A. S. J. S., Michel, J., Montanari, F., Niu, T., Ogino, R., Palaniappan, A., Pan, X., Patnaik, A., Pham, L. H., Pinto, L., Purnomo, J., Rich, A., Schaaf, L., Schran, C., Singh, R. K., Srilakshmi, M., Srivastava, S. P., Sun, K., Sun, Z., Talagayev, V., Thirukonda Subramanian Balakrishnan, B., Titus, I., Tkatchenko, A., Treyde, W., Tricarico, G., Tripp, A., Vithayapalert, N., Wang, Y., Wasi, A. T., Wedig, S., Wolber, G., Xu, B., Zhou, W., von Delft, F., Lee, A., Kirkegaard, K., Sjö, P., Fraser, J. S. & Chodera, J. D., 26 Feb. 2026, in: Journal of Chemical Information and Modeling. 66, 6, S. 3129-3149 21 S.

    Veröffentlichungen: Beitrag in FachzeitschriftArtikelPeer Reviewed

  • Chemical Similarity and Substructure Searches

    Kriege, N. M., Seidel, T., Humbeck, L. & Lessel, U., 1 Jan. 2025, Encyclopedia of Bioinformatics and Computational Biology. 2 Aufl. Elsevier, Band 3. S. 707-719

    Veröffentlichungen: Beitrag in BuchBeitrag in Buch/SammelbandPeer Reviewed

  • Crossfire: An Elastic Defense Framework for Graph Neural Networks Under Bit Flip Attacks

    Kummer, L., Moustafa, S., Gansterer, W. & Kriege, N. M., 11 Apr. 2025, Proceedings of the 39th Annual AAAI Conference on Artificial Intelligence. AAAI, S. 17990-17998 9 S. (Proceedings of the ... National Conference on Artificial Intelligence; Nr. 17, Band 39).

    Veröffentlichungen: Beitrag in BuchBeitrag in KonferenzbandPeer Reviewed

    Open Access