Data Engineer
  • Apetan Consulting LLC
2 Days Ago
NA
C2C
Phoenix-AZ
7-10 Years
Required Skills: Python, Apache Spark, AWS, Azure
Job Description
Responsibilities
  • Design and build scalable, fault-tolerant data pipelines using Apache Spark and Python.
  • Establish and enforce data governance frameworks, including data quality rules, lineage, cataloging, and access controls.
  • Implement data reliability practices such as validation checks, anomaly detection, SLAs/SLOs, and automated alerting.
  • Develop and maintain automated data workflows using orchestration tools such as Apache Airflow.
  • Partner with stakeholders to define data contracts, schemas, and standards across domains.
  • Enable end-to-end observability of data pipelines, including freshness, completeness, and accuracy.
  • Support machine learning pipelines, including feature engineering, training data validation, and monitoring model input/output data quality.
  • Document data assets, lineage, and governance policies to improve discoverability and trust.
 
Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Strong programming skills in Python and experience with Apache Spark.
  • Proven experience building reliable ETL/ELT pipelines in production environments.
  • Hands-on experience with data quality frameworks.
  • Experience implementing data governance principles, including cataloging, lineage, metadata management, and access control.
  • Familiarity with cloud data platforms such as AWS, Azure, or GCP.
  • Strong understanding of data modeling, schema evolution, and data lifecycle management.
  • Experience supporting ML pipelines, with a focus on data validation and feature stores.

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates