What are the responsibilities and job description for the Data Center Operations Engineer position at Intraedge?
Job Description:
As our Senior Data Center Operations Engineer, your main mission is to own day-to-day operations, ensure maximum uptime, and drive infrastructure scalability for enterprise and high-density, GPU-based AI platforms. You will serve as the primary technical lead for 24x7 colocation facility operations, reporting to infrastructure leadership and partnering cross-functionally to manage expansions, capacity planning, and vendor SLAs.
As our Senior Data Center Operations Engineer, your main mission is to own day-to-day operations, ensure maximum uptime, and drive infrastructure scalability for enterprise and high-density, GPU-based AI platforms. You will serve as the primary technical lead for 24x7 colocation facility operations, reporting to infrastructure leadership and partnering cross-functionally to manage expansions, capacity planning, and vendor SLAs.
Core Responsibilities
- Enforce Standards: Establish, maintain, and audit data center operational procedures, disaster recovery runbooks, and strict change management policies.
- Lead Deployments: Perform and oversee rack/stack installations, structured cabling, configuration, troubleshooting, and decommissioning of enterprise server, network, and storage hardware.
- Manage Capacity: Track and forecast rack utilization, power consumption, cooling capacity, and floor space while maintaining precise DCIM inventory records and rack elevation diagrams.
- Support 24x7 Operations: Participate in a 24x7 on-call rotation to provide rapid incident response, remote/smart hands support, and physical hardware remediation.
- Drive AI Innovation: Evaluate emerging technologies to optimize high-density NVIDIA GPU platforms, cooling strategies, and automated data center workflows.
- Mentor Teams: Create training documentation and provide technical guidance to junior engineers to scale operational readiness.
Required Qualifications
- Experience: 10 years in enterprise data center operations, with 5 years specifically supporting large-scale colocation facilities and complex Fortune 500 environments.
- Availability & Travel: Ability to travel up to 35% and fully participate in a mandatory 24x7 on-call operational support rotation.
- Education: Bachelor’s degree in IT, Computer Science, Engineering, Data Center Operations, or a related technical field (or an equivalent combination of education, technical certifications, military experience, and relevant professional experience).
- Infrastructure Expertise: Deep hands-on experience with server hardware, storage systems, Cisco Nexus switches, enterprise routing, VMware, Linux, and Windows Server environments.
- Facility Systems Knowledge: Comprehensive understanding of UPS systems, generators, HVAC, fire suppression, power distribution, and DCIM/ticketing platforms (ServiceNow preferred).
- Project Leadership: Proven track record of successfully leading data center build-outs, migrations, capacity modeling, and hardware lifecycle management.
Preferred Certifications
- Data Center: CDCDP, CDCS, or equivalent data center operations certifications.
- Networking: Cisco CCNA, CCNP, or higher.
- Systems & Process: VMware certifications, ITIL Foundation (or higher), and relevant cloud or infrastructure certifications.