Senior Linux Administration
  • Soft Snippets Inc.
16 Days Ago
NA
C2C
Santa Clara-CA
7-10 Years
Required Skills: Python, data center, cloud, AI, HPC environments
Job Description
ENGAGEMENT SUMMARY
The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
 
WHAT THIS CANDIDATE WILL BE DOING
  • Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.

  • Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.

  • Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.

  • Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.

  • Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.

  • Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.

  • Automate repeatable administration and remediation tasks with Bash and Python.

  • Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.

WHAT WE NEED TO SEE
  • 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.

  • Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.

  • Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.

  • Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.

  • Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.

  • Strong shell scripting and Python-based automation capability.

  • Working knowledge of storage and network dependencies affecting Linux host health.

  • Ability to operate independently in ambiguous, high-severity production situations.

PREFERRED EXPERIENCE
  • Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.

  • Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.

  • Experience supporting validation labs or pre-production cluster certification.

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates