Senior Cloud Engineer

THE POSITION

The Senior Cloud Engineer, ML Platforms is responsible for building, maintaining and evolving the cloud infrastructure that enables machine learning activities within Computational Biology. Working closely with researchers, ML engineers and platform teams, you will ensure that AWS-based environments remain secure, scalable and reliable while supporting the full machine learning lifecycle, from experimentation and model training through deployment and production operations.

 

As part of this role, you will work closely with Boehringer Ingelheim's newly established AI Accelerator in London, supporting the cloud and platform capabilities behind its AI and machine learning initiatives. This is an opportunity to contribute to the infrastructure that enables cutting-edge research teams to develop, scale and deploy advanced AI solutions across a wide range of biomedical challenges.

 

Tasks and Responsibilities

  • Design, maintain and continuously improve AWS-based infrastructure supporting machine learning workloads, including SageMaker, networking, IAM, storage, compute resources and model endpoints.
  • Manage cloud environments through Infrastructure as Code, ensuring consistency, scalability and compliance with enterprise architecture, security and governance standards.
  • Monitor platform performance, availability, security findings and resource utilization, proactively identifying and resolving operational issues.
  • Plan and manage cloud capacity, including CPU, GPU, storage and networking resources, balancing business needs, platform performance and cost efficiency.
  • Build and support infrastructure for MLOps processes, including CI/CD pipelines, experiment tracking, model registries, automated workflows and model deployment.
  • Develop reusable automation and platform capabilities that simplify onboarding, reduce manual work and improve the user experience for researchers and ML teams.
  • Enable and maintain integrations between AWS services and supporting technologies such as Databricks, MLflow, Jenkins, Bitbucket, OpenShift and related platforms.
  • Act as the primary technical contact for stakeholders, translating business and research requirements into effective cloud and platform solutions.
  • Create and maintain technical documentation, support onboarding activities and contribute to the evaluation of new cloud and MLOps technologies.

 

Requirements

  • Hands-on experience designing, implementing and supporting cloud infrastructure in AWS environments.
  • Strong knowledge of AWS services including SageMaker, IAM, networking, storage, compute services and container technologies.
  • Experience with Infrastructure as Code and cloud automation practices.
  • Understanding of cloud security, governance, compliance and access management principles.
  • Experience supporting machine learning, data science or MLOps platforms.
  • Knowledge of CI/CD practices and tools used for software and machine learning delivery.
  • Experience working with technologies such as Databricks, MLflow, Jenkins, Bitbucket, OpenShift or comparable platforms.
  • Ability to troubleshoot complex technical issues and continuously improve platform reliability, performance and efficiency.
  • Strong stakeholder management and communication skills, with the ability to work effectively across international and cross-functional teams.
  • Degree or equivalent qualification in Information Technology, Computer Science or a related field.

 

This is a hybrid role with approximately 3 days a week in the office.