About the role
Yobitel builds and operates the AI-native compute layer that our customers' training and inference workloads run on. We own the hardware rather than reselling someone else's: H100 fleets in production today, B200 capacity landing next. When a rack degrades at 3am, it is our pager and our problem.
We are hiring a senior systems engineer to keep that fleet fast, healthy and full. This is a deep systems role that sits close to the metal, spanning firmware and driver stacks, NUMA and topology tuning, RDMA fabric behaviour, and the performance regressions that only ever show up under real multi-node load. You will work alongside the platform and ML infrastructure teams, and your work sets the ceiling on what every workload above you can achieve.
What success looks like
- In your first 90 days, you have commissioned a rack end to end and written the runbook that lets someone else repeat it.
- Node-level performance regressions get caught by automated checks before customers feel them.
- Fleet utilisation and time-to-recovery both improve, and you can show the numbers.
Responsibilities
- Commission new GPU racks end to end: burn-in, hardware validation, firmware and driver baselines, topology verification, acceptance into the fleet.
- Diagnose failures at the hardware boundary, covering GPU faults, NVLink and NVSwitch degradation, PCIe and RDMA fabric issues, thermal and power events.
- Tune node performance for real workloads: NUMA placement, CPU and interrupt affinity, huge pages, GPUDirect and RDMA paths.
- Build and maintain automated health checking and burn-in tooling so a bad node is quarantined before a customer schedules onto it.
- Own driver, CUDA toolkit and firmware version policy across the fleet, including safe staged rollout and rollback.
- Investigate multi-node training and inference performance regressions with the ML infrastructure team, and drive them to root cause.
- Carry the on-call rotation for fleet health, and write the runbook the first time you solve something by hand.
Requirements
- Substantial production experience operating GPU or HPC compute at fleet scale, rather than single workstations.
- Strong Linux systems fundamentals: kernel and userspace boundaries, cgroups, scheduling, memory, storage and network stacks.
- Hands-on familiarity with the NVIDIA stack: driver and CUDA toolkit management, NVML, DCGM, MIG, and GPU telemetry.
- Practical experience debugging high-performance networking, whether InfiniBand, RoCE or RDMA, including fabric-level diagnosis.
- Comfort automating operational work in Python, Go or Bash, and treating tooling as a first-class deliverable.
- Experience carrying production on-call for infrastructure with real availability commitments.
- Clear written communication, since much of this role is documenting what you learned so the next person does not relearn it.
Nice to have
- Experience bringing up a brand-new GPU generation, including the surprises that only appear on new silicon.
- Familiarity with Kubernetes device plugins, GPU operators, or scheduling GPUs in a containerised environment.
- Background in data-centre-side concerns such as rack power budgeting, liquid cooling or high-density thermal design.
- Contributions to relevant open-source projects, or published benchmarking work.
Role reference: YJ-00001