Monte Carlo Methods for Constrained and Aggregated Data.
Not stated
- Location
- Birmingham, United Kingdom
- Funding
- Funded PhD Project (UK Students Only)
- Application deadline
- Year-round applications
About the project
About the Project In many applications, the quantities of interest are never observed directly, only through aggregates. Disease counts are reported by region, official statistics are published as tables of totals, and network traffic is measured on links rather than routes. Bayesian inference then requires sampling latent variables that satisfy exact constraints. Standard Markov chain Monte Carlo (MCMC) methods often mix poorly in this setting, and there is little theory to guide how they should be designed or tuned. This project will develop and analyse Monte Carlo methods for such problems. It has two complementary strands. Fusion methods under exact constraints. Monte Carlo Fusion and its sequential Monte Carlo extensions combine separate posterior inferences, for example from different data sources, into a single inference. Existing versions that enforce exact aggregate constraints are too slow for practical use, while approaches that impose constraints only approximately leave a tolerance error. The project will develop constructions that enforce the constraints exactly while remaining efficient, with theoretical guarantees on their accuracy and computational cost. Optimal scaling in constrained domains. Optimal scaling theory, such as the well-known 0.234 acceptance rule, tells practitioners how to tune MCMC algorithms in high dimensions. It has been developed mainly for unconstrained targets, however, and can break down near the boundary of a constrained space. The project will derive diffusion limits and optimal tuning rules for samplers such as random-walk Metropolis and MALA on constrained domains, starting with the probability simplex in growing dimension. The two strands feed into each other: efficient fusion schemes rely on well-tuned MCMC moves on constrained spaces, and scaling theory tells us how to tune them. The balance between theory, methodology and computation can be shaped around the student's interests, and there is scope to bring in optimisation-based or learned proposals where they help. Who should apply Applicants should have, or expect to obtain, a 2:1 four-year degree (or equivalent) in a mathematics-related subject. The key requirements are strong mathematical ability and a genuine interest in probability and statistics. Applicants whose background is mainly in pure or applied mathematics are encouraged to apply: prior study of statistics, stochastic processes or Monte Carlo methods is helpful but not essential. Some programming experience, for example in Python or R, is an advantage, and applicants should be willing to develop it during the PhD. How to apply Informal enquiries are welcome before applying. Please email Dr Shenggang Hu ( s.hu@bham.ac.uk ) with a CV, your transcripts and a short statement (no more than one page) saying which strand interests you and why. Formal applications should be made through the University of Birmingham online application system for the Statistics PhD, naming Dr Hu as the proposed supervisor and quoting the project title. Applications are considered on a rolling basis, and the position will close once a suitable candidate is found, so early applications are strongly encouraged. The intended start date is September 2027. Funding notes: The scholarship will cover Home tuition fees, training support and a stipend at standard rates for 3–3.5 years, for an applicant with Home (UK) fee status. Self-funded students worldwide are welcome to apply, as are applicants with external scholarships such as the China Scholarship Council (CSC) scheme.