Site Reliability Engineer
  • Galent
2 Days Ago
NA
W2
Pittsburgh-PA
8-10 Years
Required Skills: Java Spring boot, Apache Kafka, DevOps, CI/CD automation
Job Description
 
Automation & Efficiency
  • Automate the top 5 high-volume support and request types
  • Build self-service and agent-driven solutions to reduce manual work
  • Harden operational workflows for consistency, auditability, and resilience
  • Implement auto-retry and backoff for recurring failure patterns
Reliability Engineering
  • Define and manage Service Level Objectives (SLOs) for critical services and batch processes
  • Apply error budget concepts to guide reliability and release decisions
  • Improve batch reliability through standardized recovery patterns and monitoring
Observability & Metrics
  • Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
  • Improve operational reporting and visibility across incidents, problems, and changes
Runbooks & Self-Service
  • Develop and expand runbooks for key production scenarios
  • Convert runbooks into automated remediation workflows
  • Enable self-service for repeat operational requests
  • Drive conversion of repeat incidents into permanent fixes and known problems
Self-Healing & Intelligent Operations
  • Implement self-healing capabilities to minimize manual intervention
  • Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
  • Leverage automation and AI to resolve recurring issues with minimal human involvement

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates