Required Skills: AIOps, Python , AWS, Azure, GCP; strong Linux, networking, containers, Kubernetes, IaC
Job Description
Required Skills and Attributes
-
AIOps & SRE fundamentals: 5+ years in SRE/production operations; apply SLO/SLI, error budgets, incident management, and automated remediation patterns.
-
Observability toolset: 3+ years building dashboards/alerts and troubleshooting with tools such as Dynatrace and Splunk (log/metric/trace correlation, alert tuning, noise reduction).
-
Cloud & platform engineering: 3+ years operating services on AWS/Azure/GCP; strong Linux, networking, containers/Kubernetes, and IaC (e.g., Terraform/CloudFormation) for reliable, repeatable environments.
-
Automation & agentic ops mindset: Proficient in Python and/or Bash plus CI/CD; build runbooks, self-healing workflows, and safe change automation; strong troubleshooting, ownership, and collaboration in on-call rotations.
-
AI-enabled SRE / intelligent ops: 1–2+ years applying AI-assisted incident response (auto summarization, auto triage, pattern detection) and predictive alerting/anomaly detection in production ops.
-
Security/DevSecOps collaboration: Working knowledge of vulnerability management, secrets/cert governance, and secure CI/CD gates; ability to partner with cyber teams and vendors.
Preferred Skills and Attributes
-
Release safety & resiliency engineering: Experience with progressive delivery (canary/blue-green), chaos testing/DR drills, and performance/capacity engineering; familiarity with well-architected reviews (Azure WARA/Azure Advisor or equivalent).
-
ITSM & on-call tooling integration: Hands-on integrating monitoring signals to ServiceNow, PagerDuty (or equivalent), and building low-friction escalation/war-room workflows.
-
Portfolio communication & stakeholder management: Able to translate reliability risks into exec-ready updates (KPIs like MTTD/MTTR, error budget burn, top recurring toil) and drive cross-team alignment.
Primary Responsibilities
-
Own reliability outcomes for critical services: Define/track SLIs & SLOs, manage error budgets, and drive continuous improvement to availability, latency, and resiliency.
-
Operate and mature the observability/AIOps platform: Build and tune monitoring, dashboards, alerting, and correlation (logs/metrics/traces) to reduce noise and speed detection and diagnosis.
-
Lead incident response & problem management: Run on-call/war rooms, perform root-cause analysis, publish post-incident reviews, and ensure corrective/preventive actions are delivered.
-
Automate toil and enable self-healing: Create runbooks, scripts, and workflows for automated remediation, safe changes, and guardrails; improve MTTR through automation.
-
Partner with engineering and stakeholders: Consult on architecture/release readiness, capacity planning, and operational standards; communicate reliability risk and KPI trends to leadership.
Prior Experience & Domain Expertise Summary
-
5+ years in SRE/production operations
-
3+ years in observability dashboards building
-
3+ years in Cloud Engineering
-
1+ year of Agentic SRE tooling – build & operationalize
AI Skills & Expectations
All contractor resources are expected to demonstrate baseline proficiency in enterprise-approved AI tools as part of their day-to-day responsibilities. This includes, but is not limited to:
-
Consistent Use: Maintain a minimum of 90% weekly usage of AI tools such as GitHub Copilot, Microsoft 365 Copilot, and other GenAI platforms approved by the enterprise.
-
Applied Productivity: Leverage AI tools to enhance coding, documentation, data analysis, and decision-making workflows.
-
Continuous Learning: Stay current with evolving AI capabilities and features, and apply them to improve delivery quality and velocity.