Required Skills: Monitoring & Observability, Grafana, Splunk, Cloud monitoring, Monitoring dashboards, Proactive monitoring, Automated alerts and notifications, AWS, Lambda, SNS, SQS, DevOps, CI/CD, Deployment services, CI/CD pipelines, Version control, Code deployment, data deployment, Release support, Containers, Docker, Kubernetes, Automation, IaC, Ansible, Chef, Puppet, GitLab, Terraform, CloudFormation, SRE, Production Operations, Root Cause Analysis, RCA, Production troubleshooting, Distributed systems, Application performance, Availability
Job Description
Senior Site Reliability Engineer (SRE) – AWS
Must Have Technical/Functional Skills
• Experience in building monitoring dashboards in Grafana, Splunk, cloud monitoring, etc.
• Create dashboards for proactive monitoring.
• Automate alerts, notifications, and daemon processes.
• Good root cause analysis and communication skills.
• Experience in AWS services such as Lambda, SNS, and SQS.
• Experience with deployment services, CI/CD pipelines, version control tools, and monitoring tools.
• Experience with Docker and Kubernetes.
• Knowledge of automation tools such as Ansible, Chef, Puppet, GitLab, Terraform, and CloudFormation.
• Ensure the availability, performance, and scalability of a website or application.
• Provide solution support and configuration changes across all environments.
• Deep understanding of distributed systems to troubleshoot and optimize them.
• Code/data deployment and release support.
• Ability to solve problems quickly and effectively.
• Willingness to work in shifts and on weekends.
Roles & Responsibilities
• Identify the root cause and the application causing the issue.
• Identify whether issues are caused by code or configuration and determine the application causing the bottleneck.
• Work with different stakeholders to fix the issue.
• Troubleshoot and debug production issues.
Generic Managerial Skills, If any
Excellent verbal and written communication.
Education
Bachelor's degree