Required Skills: Java Spring boot, Apache Kafka, DevOps, CI/CD automation
Job Description
Site Reliability Engineer (SRE) – Production Services
Automation & Efficiency
-
Automate the top 5 high-volume support and request types
-
Build self-service and agent-driven solutions to reduce manual work
-
Harden operational workflows for consistency, auditability, and resilience
-
Implement auto-retry and backoff for recurring failure patterns
Reliability Engineering
-
Define and manage Service Level Objectives (SLOs) for critical services and batch processes
-
Apply error budget concepts to guide reliability and release decisions
-
Improve batch reliability through standardized recovery patterns and monitoring
Observability & Metrics
-
Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
-
Improve operational reporting and visibility across incidents, problems, and changes
Runbooks & Self-Service
-
Develop and expand runbooks for key production scenarios
-
Convert runbooks into automated remediation workflows
-
Enable self-service for repeat operational requests
-
Drive conversion of repeat incidents into permanent fixes and known problems
Self-Healing & Intelligent Operations
-
Implement self-healing capabilities to minimize manual intervention
-
Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
-
Leverage automation and AI to resolve recurring issues with minimal human involvement