Skip to main content

DevOps / SRE Lead

Own 99.99% uptime SLAs across GPU cloud, Yobibyte platform, and customer-facing APIs. Terraform, Prometheus, Grafana.

Platform Remote full time Remote-eligibleCompetitive

About the role

We carry uptime commitments to customers, which means availability is a contractual obligation rather than an aspiration. You will own the discipline that keeps us honest about it: service level objectives, error budgets, incident command, blameless post-mortems and a paging policy that wakes people only when it should.

This is a lead role covering the GPU cloud, the Yobibyte platform and the customer-facing APIs. You will set the reliability standard and bring the rest of engineering along with it, which is as much an influence and culture problem as a technical one. Expect to spend real time on observability quality, because most reliability problems are visibility problems first.

What success looks like

  • Every critical service has meaningful SLOs, and error budgets actually influence what ships.
  • Alerts are trusted, because the noisy ones have been deleted or fixed.
  • Incidents produce durable fixes and better runbooks rather than repeat visits.

Responsibilities

  • Define and maintain service level objectives and error budgets across the GPU cloud, platform and public APIs.
  • Own the observability stack end to end: metrics, logs, traces, dashboards and the signal-to-noise quality of alerting.
  • Run incident management, including on-call rotation design, incident command, escalation paths and blameless post-mortems.
  • Drive reliability improvements to completion, tracking post-mortem actions rather than filing and forgetting them.
  • Own infrastructure as code standards and change management, so risky changes are staged and reversible.
  • Build and maintain runbooks and operational documentation, and keep them current as systems change.
  • Coach engineers across teams on operability, so reliability is designed in rather than bolted on afterwards.

Requirements

  • Substantial experience running multi-region production systems with genuine availability commitments.
  • Depth in modern observability tooling, for example Prometheus, Grafana, Loki, OpenTelemetry or equivalents.
  • Practical SLO and error-budget experience, including the organisational conversations they force.
  • Strong infrastructure-as-code background, such as Terraform or Pulumi, with review discipline to match.
  • Proven incident-command capability under real pressure, and the temperament that goes with it.
  • Automation-first instincts, with the coding ability in Python or Go to act on them.
  • Willingness to lead through influence across teams that do not report to you.

Nice to have

  • Experience operating GPU or other specialised hardware fleets.
  • Background in capacity planning and cost optimisation for expensive compute.
  • Exposure to chaos engineering or systematic resilience testing.
  • Security or compliance experience relevant to enterprise customers.

Role reference: YJ-00007