Required Skills: ArgoCD & Karpenter. Azure, Oracle Cloud, Python, Go, Bash, Linux, Datadog
Job Description
Responsibilities:
Reliability Engineering & Operations
· Own and improve service reliability through SLO/SLI definition, error budgets, and operational best practices.
· Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection.
· Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions.
· Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios.
Infrastructure & Cloud Architecture:
· Design and operate highly available, fault-tolerant Cloud architectures (multi-AZ and, where required, multi-region).
· Implement resilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication).
· Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams.
Infrastructure as Code (IaC) & DevOps Enablement:
· Build and maintain IaC modules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery.
· Develop, standardize, and optimize CI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline).
· Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns.
· Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation.
BCP/DR, RTO/RPO Definition & Testing:
· Collaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements.
· Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation).
· Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes.
· Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gaps identified during testing.
· Produce metrics and reporting on DR readiness, test results, and continuous improvement actions.
Qualifications:
· 7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments.
· Strong proficiency with observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft, etc).
· Strong hands-on AWS experience building and operating production systems.
· Proven expertise with Infrastructure as Code (Terraform and/or CloudFormation/CDK).
· Strong CI/CD and automation background (pipeline design, deployment strategies, testing automation).
· Experience defining and validating RTO/RPO, and implementing BCP/DR plans with structured testing.
· Experience with Kubernetes and auto-scaling container platforms (EKS, ECS, or Kubernetes on-prem).
· Strong Linux fundamentals, networking concepts (DNS, TCP/IP, load balancing), and troubleshooting skills.
· Proficiency in at least one scripting/programming language (Python, Go, Bash, or similar).
· Ability to write clear operational documentation, runbooks, and post-incident reports.
· Ability to work effectively in a fast-paced, dynamic and high-intensity environment including open-floor plan if applicable to the position, with timely responsiveness and the ability to work beyond normal business hours when required.
Preferred Qualifications:
· Familiarity with Azure and/or Oracle Cloud (OCI).
· Familiarity with Service Mesh, API Gateways, and distributed tracing tooling.
· Familiarity with OpenTelemetry, client instrumentations and collector configurations.
· Security and compliance familiarity in cloud environments (IAM design, secrets management, audit logging).
· Experience implementing progressive delivery (blue/green, canary), feature flags, and automated rollback.
· Relevant certifications (AWS Solutions Architect/DevOps Engineer, Kubernetes CKA/CKAD).
· Experience with ArgoCD & Karpenter.