Senior Site Reliability Engineer - Python
Help build and operate reliable, scalable and secure digital services that support businesses across the UK. The Department for Business and Trade (DBT), in partnership with Inspire People, is seeking a Senior Site Reliability Engineer with strong software engineering capability (either experience of Python development or a desire to develop Python expertise), experience building applications hosted on cloud platforms, expertise within a modern cloud engineering environment. and a passion for automation and continuous improvement. Based from London, Cardiff, Darlington, Edinburgh, Belfast, Birmingham or Salford, this permanent opportunity offers hybrid working, flexible working patterns and a National salary of £63,824 to £80,158 (£67,547 to £83,778 for London) plus 29% Pension and Civil Service Benefits. Salary awarded based on technical skills assessed at interview. Shape the Future of Cloud Engineering at DBT The Department for Business and Trade has a clear mission: to grow the economy by helping businesses invest, grow and export, creating jobs and opportunities across the UK. DBT's Digital, Data and Technology directorate develops and operates the digital platforms and services that support this mission. As a Senior Site Reliability Engineer, you'll join a highly skilled engineering team responsible for developing, operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior Site Reliability Engineer you will: Build and maintain shared service products enabling developers to work more efficiently. Write clean, maintainable and well-tested code to support platform and tooling development. Work closely with development teams to provide and improve platform tooling, including monitoring, logging, metrics, dashboards and CI/CD pipelines. Build, maintain and improve reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error bu...