Required Skills: AWS Data Platform, Production Support, Databricks, Workflows, Clusters, Notebooks, Delta Tables, SQL, SQL & Data Pipeline Troubleshooting, Incident Management, Monitoring & Production Troubleshooting
Job Description
Sr. DataOps AWS Support Engineer
We are looking for a Senior DataOps AWS Support Engineer responsible for the operational stability, monitoring, troubleshooting, and continuous improvement of AWS-based data workloads.
Responsibilities:
· Provide production support for AWS DataOps workloads across Databricks StoreOps and APIHUB
· Monitor Databricks jobs, workflows, clusters, notebooks, Delta tables, Unity Catalog assets, and data-access controls
· Support AWS data-lake services, including S3, IAM, CloudWatch, networking dependencies, and account-level access patterns
· Troubleshoot failed jobs, delayed loads, data freshness issues, compute problems, pipeline interruptions, schema changes, and environment-specific deployment failures
· Support APIHUB services, including FastAPI or REST/SOAP endpoints, API dependencies, OpenSearch indexes, PostgreSQL or equivalent stores, and service-health monitoring
· Investigate ingestion and integration issues involving Kafka, Confluent Cloud, GoldenGate, GCP Pub/Sub, S3 sinks, Databricks Auto Loader, and cross-cloud data flows
· Perform first-response triage, impact assessment, incident coordination, stakeholder communication, and escalation for high-priority issues
· Use logs, job history, SQL analysis, Databricks run details, CloudWatch, Datadog, OpenSearch, Kafka tooling, and service metrics to isolate failures
· Validate data completeness, freshness, record counts, duplicates, schema compatibility, and downstream availability
· Support access requests involving AWS accounts, Databricks workspaces, Unity Catalog, S3 paths, AD groups, service accounts, secrets, and governed datasets
· This position description identifies the responsibilities and tasks typically associated with the performance of the position. Other relevant essential functions may be required.
Requirements:
· 5-9 years of experience
· Strong experience supporting production AWS data platforms, analytics environments, or cloud-native applications
· Hands-on experience with Databricks, including workflows, clusters, notebooks, Delta tables, SQL, and job troubleshooting
· Working knowledge of AWS services such as S3, IAM, CloudWatch, and account or environment access controls
· Strong SQL skills and the ability to investigate data issues directly in analytical platforms
· Experience troubleshooting batch and streaming data pipelines in production
· Experience with incident management, operational triage, root-cause analysis, and post-incident actions
· Working knowledge of Python or another scripting language for diagnostics and automation
· Experience with monitoring and observability tools such as Datadog, CloudWatch, or equivalent platforms
· Experience supporting CI/CD workflows and infrastructure-as-code practices, preferably GitLab and Terraform
· Ability to document technical procedures clearly and communicate effectively with engineering, business, and vendor teams
· Ability to work in a managed-services environment with defined response expectations, ticket queues, escalation paths, and on-call coverage