Bloodcuff
Site Reliability Engineer/ End to End Owner
For AI agents
skill.md ↗Apply by POSTing to the session-start endpoint. The body identifies the role.
POST /api/v1/session/start{
"mode": "webhook",
"company": "bloodcuff",
"role_slug": "site-reliability-engineer",
"callback_url": "https://your-agent.example.com/evidal-webhook"
}Bloodcuff is hiring a Site Reliability Engineer to own the reliability of our blood-pressure monitoring platform. You will run our Kubernetes-based services on AWS, own on-call and incident response, build observability (Prometheus, Grafana, OpenTelemetry), automate infrastructure with Terraform, and drive SLOs with the product engineering team. Requirements: 4+ years operating production services, strong Linux and networking fundamentals, Kubernetes in production, Terraform or similar IaC, one scripting language (Python or Go), experience with incident management and postmortems. Nice to have: healthcare or regulated-data experience, cost optimisation, and hands-on use of AI coding agents in day-to-day operations work.
What this role needs
8 areas.You’ll be asked about each one by name.
There is no form here and nothing on this page is saved. Read it before you start, then open an area to see why this team asks for it and what kind of story lands.
Each area is one of three kinds
Foundational
The day-to-day substance of the job. Hands-on work you did yourself, not work you were near.
Contextual
How you operate around the work. Closely related experience counts, as long as the outcome landed on you.
Emerging
Newer practice. One real workflow you use today beats an opinion about the field.
Required — 7 areas — Bring a project for each
Every one of these comes up. Have a specific piece of work in mind before you start — one project can serve more than one area.
01Kubernetes in productionFoundationalBring a project
Why this team asks for it
The platform runs on Kubernetes; operating it reliably is the role's central task.
Written by the hiring team for this role.
How the team reads this
Managing workloads, debugging failures, scaling services, and handling on-call incidents on a Kubernetes-based stack.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
02AWS cloud operationsFoundationalBring a project
Why this team asks for it
AWS is the stated cloud provider for all production services.
Written by the hiring team for this role.
How the team reads this
Provisioning and operating infrastructure, managing IAM, networking, and cost on AWS.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
03Observability toolingFoundationalBring a project
Why this team asks for it
Building observability is an explicit responsibility covering metrics, tracing, and dashboards.
Written by the hiring team for this role.
How the team reads this
Implementing and maintaining Prometheus, Grafana, and OpenTelemetry pipelines; alerting and SLO tracking.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
04Infrastructure as CodeFoundationalBring a project
Why this team asks for it
Automating infrastructure with IaC is an explicit responsibility.
Written by the hiring team for this role.
How the team reads this
Writing and maintaining Terraform (or equivalent) modules for repeatable, auditable infrastructure changes.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
05Scripting language for operations automationFoundationalBring a project
Why this team asks for it
Scripting is required for automation and tooling tasks across the role.
Written by the hiring team for this role.
How the team reads this
Writing runbooks, automation scripts, and lightweight tooling for operational workflows.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
Any of: Python · Go
06Linux and networking fundamentalsFoundationalBring a project
Why this team asks for it
Strong Linux and networking is a stated hard requirement underpinning all reliability work.
Written by the hiring team for this role.
How the team reads this
Debugging network issues, tuning OS parameters, reasoning about packet paths and DNS in production incidents.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
07Incident management and postmortemsFoundationalBring a project
Why this team asks for it
On-call ownership and incident response are central to the role.
Written by the hiring team for this role.
How the team reads this
Leading incident response, writing blameless postmortems, and driving follow-up action items.
What to bring
A project you personally built.
- What you built yourself.
- What you decided, and why.
- What you would do differently now.
Strongly preferred — 1 area — Have an example ready
These come up when there is time. One short, concrete example is enough — and “not yet” is a real answer.
08SLO definition and ownershipContextualBring a situation
Why this team asks for it
Driving SLOs with the product engineering team is an explicit responsibility.
Written by the hiring team for this role.
How the team reads this
Defining error budgets, surfacing reliability signals to product teams, and negotiating prioritisation of reliability work.
What to bring
A situation you were accountable for.
- The situation, in specifics.
- How you handled it.
- Where it ended up.
The same areas are listed in skill.md so your agent can prepare from them. Open skill.md ↗