Sign in

Bloodcuff

Site Reliability Engineer/ End to End Owner

For AI agents

skill.md ↗

Apply by POSTing to the session-start endpoint. The body identifies the role.

POST /api/v1/session/start
{
  "mode": "webhook",
  "company": "bloodcuff",
  "role_slug": "site-reliability-engineer",
  "callback_url": "https://your-agent.example.com/evidal-webhook"
}

Bloodcuff is hiring a Site Reliability Engineer to own the reliability of our blood-pressure monitoring platform. You will run our Kubernetes-based services on AWS, own on-call and incident response, build observability (Prometheus, Grafana, OpenTelemetry), automate infrastructure with Terraform, and drive SLOs with the product engineering team. Requirements: 4+ years operating production services, strong Linux and networking fundamentals, Kubernetes in production, Terraform or similar IaC, one scripting language (Python or Go), experience with incident management and postmortems. Nice to have: healthcare or regulated-data experience, cost optimisation, and hands-on use of AI coding agents in day-to-day operations work.

What this role needs

8 areas.You’ll be asked about each one by name.

There is no form here and nothing on this page is saved. Read it before you start, then open an area to see why this team asks for it and what kind of story lands.

Each area is one of three kinds

Foundational

The day-to-day substance of the job. Hands-on work you did yourself, not work you were near.

Contextual

How you operate around the work. Closely related experience counts, as long as the outcome landed on you.

Emerging

Newer practice. One real workflow you use today beats an opinion about the field.

Required — 7 areas — Bring a project for each

Every one of these comes up. Have a specific piece of work in mind before you start — one project can serve more than one area.

01Kubernetes in productionFoundationalBring a project

Why this team asks for it

The platform runs on Kubernetes; operating it reliably is the role's central task.

Written by the hiring team for this role.

How the team reads this

Managing workloads, debugging failures, scaling services, and handling on-call incidents on a Kubernetes-based stack.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.
02AWS cloud operationsFoundationalBring a project

Why this team asks for it

AWS is the stated cloud provider for all production services.

Written by the hiring team for this role.

How the team reads this

Provisioning and operating infrastructure, managing IAM, networking, and cost on AWS.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.
03Observability toolingFoundationalBring a project

Why this team asks for it

Building observability is an explicit responsibility covering metrics, tracing, and dashboards.

Written by the hiring team for this role.

How the team reads this

Implementing and maintaining Prometheus, Grafana, and OpenTelemetry pipelines; alerting and SLO tracking.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.
04Infrastructure as CodeFoundationalBring a project

Why this team asks for it

Automating infrastructure with IaC is an explicit responsibility.

Written by the hiring team for this role.

How the team reads this

Writing and maintaining Terraform (or equivalent) modules for repeatable, auditable infrastructure changes.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.
05Scripting language for operations automationFoundationalBring a project

Why this team asks for it

Scripting is required for automation and tooling tasks across the role.

Written by the hiring team for this role.

How the team reads this

Writing runbooks, automation scripts, and lightweight tooling for operational workflows.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.

Any of: Python · Go

06Linux and networking fundamentalsFoundationalBring a project

Why this team asks for it

Strong Linux and networking is a stated hard requirement underpinning all reliability work.

Written by the hiring team for this role.

How the team reads this

Debugging network issues, tuning OS parameters, reasoning about packet paths and DNS in production incidents.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.
07Incident management and postmortemsFoundationalBring a project

Why this team asks for it

On-call ownership and incident response are central to the role.

Written by the hiring team for this role.

How the team reads this

Leading incident response, writing blameless postmortems, and driving follow-up action items.

What to bring

A project you personally built.

  • What you built yourself.
  • What you decided, and why.
  • What you would do differently now.

Strongly preferred — 1 area — Have an example ready

These come up when there is time. One short, concrete example is enough — and “not yet” is a real answer.

08SLO definition and ownershipContextualBring a situation

Why this team asks for it

Driving SLOs with the product engineering team is an explicit responsibility.

Written by the hiring team for this role.

How the team reads this

Defining error budgets, surfacing reliability signals to product teams, and negotiating prioritisation of reliability work.

What to bring

A situation you were accountable for.

  • The situation, in specifics.
  • How you handled it.
  • Where it ended up.

The same areas are listed in skill.md so your agent can prepare from them. Open skill.md ↗

How to prepare

  • Pick one piece of work per required area. The same project can serve several.
  • Know the boundary of what you did — what you built, what your team built, what you inherited.
  • Bring the decision, not only the outcome: what you chose, and what you gave up.
  • Say plainly where something did not work. Honest limits read better than a clean story.
  • Nothing on this page is a form, and nothing here is saved.