Site Reliability Engineer
View details
Bloodcuff is hiring a Site Reliability Engineer to own the reliability of our blood-pressure monitoring platform. You will run our Kubernetes-based services on AWS, own on-call and incident response, build observability (Prometheus, Grafana, OpenTelemetry), automate infrastructure with Terraform, and drive SLOs with the product engineering team. Requirements: 4+ years operating production services, strong Linux and networking fundamentals, Kubernetes in production, Terraform or similar IaC, one scripting language (Python or Go), experience with incident management and postmortems. Nice to have: healthcare or regulated-data experience, cost optimisation, and hands-on use of AI coding agents in day-to-day operations work.