About the role
GPU racks need power, cooling, networking and competent hands at any hour. You will own the operations floor at our Hyderabad site, covering everything physical that keeps the fleet running and everything procedural that keeps it running safely.
This is a hands-on management role rather than a purely administrative one. You will run the team, own vendor and supplier relationships, and be accountable for capacity, uptime and safety on the floor. Modern AI racks run at power and thermal densities that punish improvisation, so process discipline and clear documentation matter as much as technical capability.
What success looks like
- Installs and hardware replacements happen predictably against a plan rather than reactively.
- Power and thermal headroom is known and tracked, not discovered during an incident.
- The floor runs safely, and the team can explain why it is safe.
Responsibilities
- Run day-to-day data centre floor operations, covering installations, decommissions, hardware replacement and spares management.
- Own capacity planning across rack space, power and cooling, and forecast constraints before they bite.
- Track power and thermal budgets against rack design limits, and manage the headroom deliberately.
- Manage the site team including scheduling, training, safety practice and on-site escalation cover.
- Own supplier and vendor relationships across hardware, networking and connectivity, including warranty and replacement processes.
- Maintain accurate asset, inventory and change records so that what is documented matches what is installed.
- Coordinate with remote engineering teams on hardware faults, and provide reliable hands during incidents.
Requirements
- Several years running data centre operations at tier-3 standard or better, with accountability for the floor.
- Hands-on experience with high-density rack deployments, including modern air and liquid cooling approaches.
- Practical understanding of data centre power, covering distribution, redundancy, UPS and generator arrangements.
- Server hardware competence spanning installation, diagnosis and component-level replacement.
- Team management experience including scheduling, training and safety accountability.
- Working knowledge of structured cabling and data centre networking practice.
- Calm, methodical judgement during incidents, and the communication to match.
Nice to have
- Direct experience with GPU or other high-density compute deployments.
- Familiarity with data centre infrastructure management tooling.
- Relevant certification in data centre design or operations.
- Experience commissioning a new site or a significant expansion.
Role reference: YJ-00009