-
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
-
Strong programming skills in Python and experience with Apache Spark.
-
Proven experience building reliable ETL/ELT pipelines in production environments.
-
Hands-on experience with data quality frameworks.
-
Experience implementing data governance principles, including cataloging, lineage, metadata management, and access control.
-
Familiarity with cloud data platforms such as AWS, Azure, or GCP.
-
Strong understanding of data modeling, schema evolution, and data lifecycle management.
-
Experience supporting ML pipelines, with a focus on data validation and feature stores.