AI-driven Agentic SRE
  • Xchange Software
1 Days Ago
NA
C2C
Minneapolis-MN
5-10 Years
Required Skills: AIOps, Python , AWS, Azure, GCP; strong Linux, networking, containers, Kubernetes, IaC
Job Description
Required Skills and Attributes
  • AIOps & SRE fundamentals: 5+ years in SRE/production operations; apply SLO/SLI, error budgets, incident management, and automated remediation patterns.
  • Observability toolset: 3+ years building dashboards/alerts and troubleshooting with tools such as Dynatrace and Splunk (log/metric/trace correlation, alert tuning, noise reduction).
  • Cloud & platform engineering: 3+ years operating services on AWS/Azure/GCP; strong Linux, networking, containers/Kubernetes, and IaC (e.g., Terraform/CloudFormation) for reliable, repeatable environments.
  • Automation & agentic ops mindset: Proficient in Python and/or Bash plus CI/CD; build runbooks, self-healing workflows, and safe change automation; strong troubleshooting, ownership, and collaboration in on-call rotations.
  • AI-enabled SRE / intelligent ops: 1–2+ years applying AI-assisted incident response (auto summarization, auto triage, pattern detection) and predictive alerting/anomaly detection in production ops.
  • Security/DevSecOps collaboration: Working knowledge of vulnerability management, secrets/cert governance, and secure CI/CD gates; ability to partner with cyber teams and vendors.
 
Preferred Skills and Attributes
  • Release safety & resiliency engineering: Experience with progressive delivery (canary/blue-green), chaos testing/DR drills, and performance/capacity engineering; familiarity with well-architected reviews (Azure WARA/Azure Advisor or equivalent).
  • ITSM & on-call tooling integration: Hands-on integrating monitoring signals to ServiceNow, PagerDuty (or equivalent), and building low-friction escalation/war-room workflows.
  • Portfolio communication & stakeholder management: Able to translate reliability risks into exec-ready updates (KPIs like MTTD/MTTR, error budget burn, top recurring toil) and drive cross-team alignment.
 
Primary Responsibilities
  1. Own reliability outcomes for critical services: Define/track SLIs & SLOs, manage error budgets, and drive continuous improvement to availability, latency, and resiliency.
  2. Operate and mature the observability/AIOps platform: Build and tune monitoring, dashboards, alerting, and correlation (logs/metrics/traces) to reduce noise and speed detection and diagnosis.
  3. Lead incident response & problem management: Run on-call/war rooms, perform root-cause analysis, publish post-incident reviews, and ensure corrective/preventive actions are delivered.
  4. Automate toil and enable self-healing: Create runbooks, scripts, and workflows for automated remediation, safe changes, and guardrails; improve MTTR through automation.
  5. Partner with engineering and stakeholders: Consult on architecture/release readiness, capacity planning, and operational standards; communicate reliability risk and KPI trends to leadership.
Prior Experience & Domain Expertise Summary
  • 5+ years in SRE/production operations
  • 3+ years in observability dashboards building
  • 3+ years in Cloud Engineering
  • 1+ year of Agentic SRE tooling – build & operationalize
 
AI Skills & Expectations
All contractor resources are expected to demonstrate baseline proficiency in enterprise-approved AI tools as part of their day-to-day responsibilities. This includes, but is not limited to:
  • Consistent Use: Maintain a minimum of 90% weekly usage of AI tools such as GitHub Copilot, Microsoft 365 Copilot, and other GenAI platforms approved by the enterprise.
  • Applied Productivity: Leverage AI tools to enhance coding, documentation, data analysis, and decision-making workflows.
  • Continuous Learning: Stay current with evolving AI capabilities and features, and apply them to improve delivery quality and velocity.

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates