Carbon- and Cost-Aware Data Engineering for Lakehouse Platforms
Not stated
- Funding
- Self-Funded PhD Students Only
- Application deadline
- Year-round applications
About the project
About the Project Modern analytics and artificial intelligence increasingly rely on lakehouse platforms such as Databricks and Snowflake, which combine large-scale data storage, transformation and analytics within a unified architecture. As organisations migrate more workloads to these platforms, the energy, carbon and financial cost of the underlying data-engineering pipelines is becoming a primary concern. Data centres already account for a significant and rapidly growing share of global electricity demand, driven in large part by data- and AI-intensive workloads. At the same time, increasing regulatory, organisational and environmental pressures are making carbon-aware computing an expectation rather than an aspiration. Despite these trends, most data-engineering practice continues to prioritise performance and developer productivity, with sustainability often treated as an afterthought. This PhD will develop novel methods and tools to make data pipelines on lakehouse platforms measurably more sustainable, cost-efficient and performant. The research focuses specifically on the data and analytics layer—including data ingestion, transformation, orchestration, storage optimisation and query execution—rather than the application or microservices layer, providing a clear and distinctive research focus. Indicative research objectives include: Develop methods to measure and attribute energy consumption, carbon emissions and operational costs to individual pipeline stages and query workloads on lakehouse platforms. Investigate carbon-aware scheduling techniques for batch and ETL/ELT workloads, enabling non-time-critical processing to be shifted to periods of lower grid carbon intensity while maintaining data freshness and service-level requirements. Explore storage and query optimisation strategies—including file sizing, partitioning, clustering, materialised views, caching and workload optimisation—to understand and quantify the trade-offs between sustainability, operational cost and analytical performance. Design practical policy-as-code governance approaches that enable organisations to embed sustainability objectives into data-engineering workflows and validate these approaches using realistic case studies, such as routinely collected healthcare or public-sector datasets. The project sits at the intersection of sustainable software engineering, cloud data platforms, FinOps, GreenOps and large-scale data engineering. It offers opportunities to develop novel measurement techniques, optimisation algorithms and governance approaches while producing practical research outputs, including benchmarks, software prototypes, datasets and implementation guidance with direct value to industry and the public sector. The successful candidate will work with modern cloud data technologies, including platforms such as Databricks and Snowflake, alongside SQL, Python and distributed data-processing frameworks. Experience with cloud data platforms or Apache Spark would be advantageous but is not essential, as appropriate training will be provided throughout the project. Applicants should have a strong background in Computer Science, Data Science, Software Engineering or a closely related discipline, together with programming experience in Python and familiarity with SQL. An interest in data-intensive systems, cloud computing, optimisation or sustainable computing would be highly beneficial. This project offers an excellent opportunity to contribute to one of the fastest-growing areas of modern computing while developing expertise that is in high demand across academia, industry and the public sector. The successful candidate will join a supportive and research-active supervisory team with strong expertise in data-intensive systems, cloud computing and applied AI, and will be encouraged to publish in leading international conferences and journals while engaging with industrial collaborators.