Required Skills: AWS, Azure, GCP, Kubernetes, CSI Storage integrations, Grafana, Prometheus, Splunk
Job Description
The Storage Engineering team is seeking a Senior / Lead Infrastructure Engineer with deep expertise in enterprise storage platforms, data protection ecosystems, and automation frameworks to lead validation and testing efforts. This role will drive automation strategy, testing frameworks, and large-scale validation to ensure the reliability, performance, and resilience of storage systems, with a primary focus on accelerating patching across the environment.
In this role, you will:
- Lead automation frameworks (Python, Playwright/Selenium) for validating storage and data workflows supporting accelerated patching.
- Design and execute validation for:
- High-velocity patch distribution and deployment workflows
- Storage performance during patching (IOPS, throughput, latency)
- API and distributed system integration for patch orchestration
- Define and execute validation strategies across:
- NetApp FAS (multi-protocol access for distributed repositories)
- StorageGRID (S3 object storage for artifacts, staging, and distribution)
- Cohesity (backup, recovery, and cyber resilience)
- Public Cloud Storage
- Validate data protection, replication, and disaster recovery for storage systems supporting accelerated patching
- Build reusable automation frameworks and integrate into CI/CD pipelines for continuous validation
- Use Terraform (IaC) to provision and validate storage environments optimized for rapid patching
- Drive testing for:
- Data integrity and consistency of patch artifacts
- High-volume data movement across environments and endpoints
- Lead deep troubleshooting and root cause analysis across storage and patch deployment workflows
Platform Focus Areas
- Cohesity: Backup/restore, cyber recovery, and protection of repositories and configurations
- PowerMax / PowerFlex: Low-latency, high-throughput storage to support rapid patch deployment
- NetApp FAS: Shared access (NFS/SMB/iSCSI/FC) for staging, repositories, and lifecycle management
- StorageGRID: S3 object storage for artifact storage, ingestion, and large-scale distribution
Basic Qualifications
- Bachelor's degree or equivalent experience
- 5+ years in storage, automation, or infrastructure engineering
- Strong Python and automation experience
- Knowledge of block, file, object storage, and data protection
Minimum Qualifications
- 5+ years with enterprise storage platforms (Cohesity, PowerMax/PowerFlex, NetApp, StorageGRID)
- Experience leading automation and validation for distributed infrastructure systems and patch deployment processes
- Expertise in:
- High-volume data delivery and validation
- Backup/restore, replication, and DR at scale
- Data integrity and large artifact handling
- Terraform (IaC) and CI/CD integration
- Advanced troubleshooting across storage, compute, and distributed systems
Desired Qualifications
- Direct experience supporting accelerated patching initiatives or large-scale update delivery environments
- Multi-cloud storage architectures (AWS, Azure, GCP)
- Kubernetes/CSI storage integrations
- Performance benchmarking for high-throughput systems
- Observability tools (Grafana, Prometheus, Splunk)
Job Expectations
- Act as SME for storage and data infrastructure with a focus on accelerating patching across storage systems
- Lead cross-functional initiatives across Storage, DevOps, and Platform Engineering teams
- Deliver scalable, production-grade automation for validation and patch delivery workflows
- Drive improvements in system performance, scalability, resiliency, and patch deployment velocity
- Mentor engineers and provide clear technical insights and recommendations