Required Skills: AWS EKS, Azure AKS, Google GKE, Helm, Argo CD, Flux, GitOps, Jenkins, GitHub Actions, GitLab, CI/CD, Azure DevOps, Kubernetes networking, ingress controllers, RBAC, secrets, storage, cloud-native security, DevSecOps practices, Ansible, Istio, Linkerd, distributed systems, microservices architectures, incident management, PagerDuty, ServiceNow, databases, caching systems, messaging platforms, API infrastructure
Job Description
We are looking for an experienced SRE / DevOps Engineer with strong hands-on expertise in Python, Terraform, Kubernetes, cloud infrastructure, and CI/CD. The ideal candidate will be responsible for building and maintaining highly scalable, reliable, secure, and automated infrastructure and deployment platforms.
The candidate should have strong experience with Infrastructure as Code (IaC), Kubernetes administration, cloud platforms, observability, automation, and production incident management.
Key Responsibilities
- Design, implement, and maintain highly available and scalable cloud infrastructure.
- Develop automation and operational tools using Python.
- Build and manage infrastructure using Terraform and Infrastructure as Code best practices.
- Deploy, configure, and manage applications and services on Kubernetes.
- Develop and maintain CI/CD pipelines for automated application build, testing, and deployment.
- Manage containerized workloads using Docker and Kubernetes.
- Implement monitoring, logging, alerting, and observability solutions.
- Define and improve SLIs, SLOs, and SLAs to measure system reliability and performance.
- Troubleshoot complex infrastructure, application, networking, and production issues.
- Participate in on-call rotations and lead incident response, troubleshooting, and root-cause analysis.
- Automate repetitive operational tasks to improve reliability and engineering efficiency.
- Implement security, access controls, secrets management, and compliance best practices.
- Perform capacity planning, performance optimization, and infrastructure scaling.
- Collaborate with software engineering, QA, security, and architecture teams.
- Develop and maintain runbooks, technical documentation, and operational procedures.
- Conduct post-incident reviews and implement preventive measures to reduce recurring issues.
Required Skills
- 7+ years of experience in SRE, DevOps, Cloud Engineering, or Infrastructure Engineering.
- Strong programming and automation experience with Python.
- Strong hands-on experience with Terraform and Infrastructure as Code.
- Extensive experience with Kubernetes and container orchestration.
- Strong experience with Docker and containerized applications.
- Experience designing and managing CI/CD pipelines.
- Strong knowledge of Linux/Unix systems and shell scripting.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Strong understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK/EFK, Datadog, Splunk, or OpenTelemetry.
- Experience with Git and modern software development workflows.
- Strong understanding of high availability, scalability, disaster recovery, and fault-tolerant architectures.
Preferred Skills
- Experience with AWS EKS, Azure AKS, or Google GKE.
- Experience with Helm, Argo CD, Flux, or GitOps.
- Experience with Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Knowledge of Kubernetes networking, ingress controllers, RBAC, secrets, and storage.
- Experience with cloud-native security and DevSecOps practices.
- Experience with Ansible or other configuration management tools.
- Knowledge of service meshes such as Istio or Linkerd.
- Experience with distributed systems and microservices architectures.
- Experience with incident management and tools such as PagerDuty or ServiceNow.
- Knowledge of databases, caching systems, messaging platforms, and API infrastructure.
SRE Responsibilities
-
Establish and maintain reliability standards across production environments.
-
Define and monitor SLIs, SLOs, and error budgets.
-
Improve system availability, latency, scalability, and resilience.
-
Identify system bottlenecks and implement proactive improvements.
-
Automate infrastructure provisioning, deployments, monitoring, and remediation.
-
Reduce manual operational work through engineering and automation.
-
Lead production incident response and perform detailed Root Cause Analysis (RCA).
-
Continuously improve operational processes and reliability engineering practices.