Required Skills: Data Center Operations, Linux administration, Windows Administration, Python scripting, Shell Scripting, Ansible Scripting, DCIM , TCP/IP, DNS, NFS, SSL, GPU & Server Hardware Support, HPC, Slurm, BCM
Job Description
What you’ll be doing:
·Collaborate closely with engineering teams, including system architects, hardware/software engineers, QA, and more, to craft, develop, debug, and release next-generation products.
·Manage and maintain a high-performing Compute Farm of builders, packagers, testers, and core infrastructure.
·Ensure availability targets are consistently met and lead system recovery efforts.
·Support hardware and software teams with hardware and software issues.
·Gather critical metrics and build Standard Operating Procedures (SOPs) documentation.
·Problem solve Linux/Windows, hardware, and infrastructure issues alongside engineers and platform operations teams.
·Implement efficiency improvements to improve availability, throughput, and test accuracy while meeting SLAs and important metrics.
What We Need to See:
·Associate’s or bachelor’s degree in engineering/Technical Major (or equivalent experience).
·5+ years with data center technology or large engineering labs.
·Proficiency in DCIM (Nautobot, etc.) and scripting (shell, Python, Ansible).
·Working knowledge of protocols/services like TCP/IP, DNS, NFS, SSL, etc.
·Experience with Windows, Linux, and Mac operating systems.
·Hands-on experience with PCBs, GPUs, and system deployments.
·Outstanding communication, both written and verbal.
·Ability to explain technical concepts to non-technical audiences.
·Strong problem-solving skills and a collaborative spirit.