About the role
The Yobibyte platform runs on a multi-cluster Kubernetes fabric spanning regions and providers. You will own the control plane that schedules, scales, isolates and gracefully drains GPU workloads across that fabric, including the parts that decide what happens when capacity runs short and something has to give.
This is a staff-level role with genuine architectural latitude. You will set direction on how the fleet is managed rather than only implementing someone else's design, and you will be expected to make the trade-offs legible to the rest of engineering. Multi-tenancy is a first-class concern here, since customers share expensive hardware and must never be able to affect each other.
What success looks like
- Cluster upgrades and fleet-wide changes become routine and low-drama instead of events.
- GPU allocation and eviction behaviour is predictable and explainable under contention.
- Tenant isolation holds under adversarial and noisy-neighbour conditions.
Responsibilities
- Own the multi-cluster control plane: fleet lifecycle, upgrade choreography, drift detection and configuration consistency across regions and providers.
- Design and build custom resources and operators covering GPU allocation, sharing, reclamation and eviction.
- Enforce multi-tenant isolation across compute, network and storage so that noisy or hostile neighbours stay contained.
- Own scheduling behaviour for GPU workloads, including queueing, prioritisation, preemption and graceful drain.
- Drive the GitOps and change-management model for infrastructure, so changes are reviewable, auditable and reversible.
- Build the cost-allocation and utilisation signals that let the business see where capacity actually goes.
- Set technical direction for the platform group and mentor engineers through design review rather than by decree.
Requirements
- Several years operating Kubernetes in production at scale, including having been responsible when it went wrong.
- Experience authoring controllers or operators, with a working understanding of the reconciliation model rather than only using off-the-shelf charts.
- Strong Go, or equivalent depth in another systems language plus willingness to work primarily in Go.
- Real multi-tenancy experience covering namespaces, network policy, resource governance and the isolation boundaries that matter.
- Fluency with infrastructure as code and GitOps-style delivery.
- Sound judgement about operational blast radius, staged rollout and rollback.
- Ability to write a design document that a sceptical reviewer can disagree with productively.
Nice to have
- Direct experience scheduling GPU or other accelerator workloads in Kubernetes.
- Familiarity with device plugins, GPU operators or topology-aware scheduling.
- Background in multi-cloud or hybrid estates spanning owned hardware and public cloud.
- Experience with service mesh, advanced networking or eBPF-based tooling.
Role reference: YJ-00002