About the role
We carry uptime commitments to customers, which means availability is a contractual obligation rather than an aspiration. You will own the discipline that keeps us honest about it: service level objectives, error budgets, incident command, blameless post-mortems and a paging policy that wakes people only when it should.
This is a lead role covering the GPU cloud, the Yobibyte platform and the customer-facing APIs. You will set the reliability standard and bring the rest of engineering along with it, which is as much an influence and culture problem as a technical one. Expect to spend real time on observability quality, because most reliability problems are visibility problems first.
What success looks like
- Every critical service has meaningful SLOs, and error budgets actually influence what ships.
- Alerts are trusted, because the noisy ones have been deleted or fixed.
- Incidents produce durable fixes and better runbooks rather than repeat visits.
Responsibilities
- Define and maintain service level objectives and error budgets across the GPU cloud, platform and public APIs.
- Own the observability stack end to end: metrics, logs, traces, dashboards and the signal-to-noise quality of alerting.
- Run incident management, including on-call rotation design, incident command, escalation paths and blameless post-mortems.
- Drive reliability improvements to completion, tracking post-mortem actions rather than filing and forgetting them.
- Own infrastructure as code standards and change management, so risky changes are staged and reversible.
- Build and maintain runbooks and operational documentation, and keep them current as systems change.
- Coach engineers across teams on operability, so reliability is designed in rather than bolted on afterwards.
Requirements
- Substantial experience running multi-region production systems with genuine availability commitments.
- Depth in modern observability tooling, for example Prometheus, Grafana, Loki, OpenTelemetry or equivalents.
- Practical SLO and error-budget experience, including the organisational conversations they force.
- Strong infrastructure-as-code background, such as Terraform or Pulumi, with review discipline to match.
- Proven incident-command capability under real pressure, and the temperament that goes with it.
- Automation-first instincts, with the coding ability in Python or Go to act on them.
- Willingness to lead through influence across teams that do not report to you.
Nice to have
- Experience operating GPU or other specialised hardware fleets.
- Background in capacity planning and cost optimisation for expensive compute.
- Exposure to chaos engineering or systematic resilience testing.
- Security or compliance experience relevant to enterprise customers.
Role reference: YJ-00007