Development of integrative computational tools for refinement and modelling in structural biology
Not stated
- Funding
- Funded PhD Project (UK Students Only)
- Application deadline
- 16 October 2026
About the project
About the Project For more information and to apply, please visit the institution website. Entry point: 1 February 2027 Knowledge of protonation states and H atom positions in macromolecules can be critical in helping to formulate functional hypotheses and, generally, in providing a complete description of the biological processes under investigation. H atoms are typically responsible for the reversible protonation of active site residues involved in enzymatic reactions and are also necessary for the formation of H-bonds that can stabilise macromolecular structures, contributing to the establishment of biological interfaces. Additionally, H atoms are often involved in determining specificities in protein-ligand recognition processes, thus their correct identification and localisation can be critical for the development and design of new therapeutics. Previous work has led to the development of novel computational tools to enhance the refinement of macromolecular models utilising neutron crystallography. These novel refinement protocols have been implemented in the macromolecular refinement software Refmac 5, one of the flagship packages within the CCP4 suite and can be seamlessly integrated within its forward compatible replacement servalcat . Neutron crystallography represents a powerful approach for the identification of H atoms in macromolecules. However, this technique is quite niche owing to several limitations related to both sample requirements and instrumentation. Novel developments in electron diffraction (ED) data collection methods as well as single particle transmission cryogenic electron microscopy (EM) can, in theory, also help to identify H atoms at relatively modest resolution. Whilst single particle cryo-EM data of sufficiently high quality are gradually becoming more available, processing of macromolecular ED data still requires improvements. Much of the available prior chemical knowledge used in macromolecular crystallographic refinement derives from high-resolution small molecule X- ray diffraction experiments. Detecting H atoms using X-ray diffraction methods is particularly challenging at low-medium resolution (2.0-2.5 Å) where they are invisible in the electron density maps. However, it is possible to locate H atoms in small molecules using X-rays, due the availability of high-quality and high- resolution low-temperature data. This entails refining their positions without any imposed constraints or restraints. H atoms are often involved in covalent bonds with other elements like C, N, and O. Due to the overlap of electron density in covalent X–H (where X is a non-H ‘parent’ atom) bonds, the mean (centroid) of the electron distribution associated with the H atom is shifted toward their parent atoms. Consequently, the X–H bond distances are, on average, approximately 0.1–0.2 Å shorter than their ‘true’ values. In neutron crystallography, neutrons do not interact with the electron shell but with the nucleus of atoms, and the scattering power does not linearly increase with the atomic number, unlike X-rays. The distances between H nuclei and their parent atoms are longer than those between electron clouds. The most recent Refmac library takes the above issues in consideration and refinement H atoms using X-ray and neutron diffraction data is relatively straightforward. However, in ED, both electrons and nuclei contribute to scattering. This results in a shift of H density peaks beyond the position of the hydrogen nucleus and that varies depending on the ADP and resolution cut-off. Thus, for EM and ED data new methods need to be developed and implemented. These should account for two positions of H atoms. As H changes its position electrons will also need to be shifted. This property will need to be learned from QM data and then applied for atomic model refinement against EM/ED data. Analysis of chemical environment will be carried out by categorising the geometrical pattern of H-bond donors and acceptors also taking into consideration the analysis of a suitable selection of the small molecule and PDB database as well as QM optimised datasets. Such selection will be carried out considering resolution criteria that will depend on the experimental technique employed. Identification, fitting and refinement of H atom will use all available data and chemical prior knowledge accounting for the source of data (electron, X-ray and neutron scattering). H refinement and modelling will be performed using experiment-dependent restraints with one set of monomer library for all. Overall, this project seeks to develop tools and extend what is currently possible by developing tools aimed at the identification and refinement of H atoms from multiple experimental techniques. This will allow to take maximum advantage of future advancements in data processing and sample quality. Supervisors: Professor Roberto Steiner Professor Jon Agirre