Production Operations Engineer, ProdOps Engineer
  • Metrix IT Solutions INC.
7 Hours Ago
NA
C2C
Las Vegas-NV
8-10 Years
Required Skills: Production Support, DevOps, SRE, Platform Engineering, GitLab CI/CD, PostgreSQL, Kubernetes
Job Description
Must-Have
  • 5+ years of experience in Production Support, DevOps, SRE, or Platform Engineering.
  • Strong hands-on experience with GitLab CI/CD, pipeline creation/optimization, GitLab Runners, and deployment automation.
  • Strong Kubernetes experience, including cluster operations, Helm Charts, and troubleshooting Pods, Services, Ingress, and cluster-level issues.
  • Strong PostgreSQL administration and database deployment experience, including schema migrations, backup/recovery, replication, and high availability.
  • Experience with Dynatrace/APM, monitoring, observability, log analysis, and alerting.
  • Strong application troubleshooting/debugging, infrastructure issue diagnosis, and performance analysis.
  • Experience with production incident management, RCA, problem management, change management, and release coordination.
  • Must have a development background and experience with operational automation/scripting.
 
Preferred / Nice-to-Have
  • Bachelor’s degree in Computer Science, Information Technology, or related field.
  • Experience with Linux administration and shell scripting.
  • Experience with AWS, Azure, or GCP.
  • Experience with Infrastructure as Code (Terraform preferred).
  • Understanding of microservices architecture and containerization.
 
Responsibilities
  • Provide production and non-production operational support and participate in 24x7 on-call rotation.
  • Design, maintain, and optimize GitLab CI/CD pipelines and automate deployment, testing, and release processes.
  • Deploy and manage containerized applications on Kubernetes and troubleshoot cluster/workload issues.
  • Plan and execute database deployments, schema migrations, releases, and rollback procedures.
  • Support and maintain PostgreSQL, including performance, backup/recovery, replication, and high availability.
  • Monitor application and infrastructure health using Dynatrace and develop proactive alerting/observability strategies.
  • Troubleshoot production incidents, perform Root Cause Analysis, and implement preventive measures.
  • Develop automation scripts, maintain runbooks/SOPs, and drive platform reliability and operational improvements.

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates