Data Engineer
  • SoniTalent Corp
2 Days Ago
NA
C2C
Dallas-TX
7-10 Years
Required Skills: SQL, Python, AWS,
Job Description

This is a hands-on data engineering role on a small team. We run a Databricks Lakehouse using medallion architecture and Unity Catalog for governance. We have a mix of daily pipelines feeding partners data through Delta Sharing, and a growing catalog of public, HIPAA, and third-party datasets.

We are still actively building out our standards and conventions. You will join while a lot of that is being figured out, which means you will not just work within the platform, you will help shape how it is structured, governed, and documented as it matures.

Day to day, the weight of the role sits in two places: building and maintaining data pipelines, and helping steward the Databricks and Unity Catalog platform that everything else depends on. We are hiring this primarily as a mid-level role, with room to grow into more ownership over time.

What youʼll do:
  • Data pipeline engineering (core):
  • Own and build batch pipelines that ingest, clean, and transform diverse datasets into the lakehouse.
  • Land and standardize raw sources into bronze under a consistent, source-named schema pattern, then build the silver and gold transformations that analysts, dashboards, and applications depend on.
  • Operate and monitor scheduled jobs and workflows, building in data quality checks, validation, and alerting so the platform runs with minimal manual intervention.
  • Maintain and extend the Delta Sharing feeds that distribute daily data to partner organizations.
  • Platform and Unity Catalog governance (core):
  • Act as a steward of the Databricks workspace and Unity Catalog: manage catalogs and schemas across the medallion structure and the sandbox and governance catalogs.
  • Administer access control: group-level grants and revocations, identity and SCIM group management, and tiered access for staff, contractors, and external engineering partners.
  • Help shape, document, and enforce governance decisions, naming conventions, and data-access policies so the environment stays legible as it grows.
  • Steward platform cost, performance, and reliability, including compute and cluster configuration, query optimization, and storage layout.
  • Keep governed datasets discoverable and appropriately access-controlled for internal and external users, in coordination with our CKAN data portal.
  • You may also:
  • Manage external data acquisition end to end, including preparing and submitting open records requests and maintaining dependable data-flow relationships with public agencies and partners and/or supporting others doing this work.
  • Support analytics engineering and data modeling as project needs arise.
  • Shape the data models behind our data applications (R Shiny, React) and support their delivery.
  • Coordinate with external engineering partners, including scoping and reviewing new pipelines.
 
Must have :
  • Solid data engineering experience building and maintaining production ETL/ELT pipelines.
  • Hands-on experience with Databricks, or a strong and transferable modern lakehouse / Spark background with a clear ability to ramp quickly: Delta Lake, jobs and workflows, and SQL/PySpark transformations.
  • Strong SQL and Python.
  • A working grasp of data governance and access control concepts (cataloging, permissions, principals and groups); direct Unity Catalog experience is a strong plus.
  • Comfort integrating messy, large, heterogeneous external data and turning it into reliable, documented tables.
  • Version control (Git) and collaborative development practices.
  • Clear communication skills, with the ability to translate technical work for non-technical stakeholders and partners.
  • Ability to juggle multiple projects and shifting priorities.
 
Nice to have :
  • Unity Catalog administration, Delta Sharing, and Databricks workspace or metastore admin experience.
  • Experience with R alongside Python.
  • Cloud platform familiarity (AWS preferred given our Databricks footprint; other clouds welcome).
  • Workflow orchestration and CI/CD practices for data.
  • CKAN or other data-portal / open-data platform experience.
  • GIS and geospatial data experience, including coordinate reference systems.
  • Experience preparing open records or public records requests.
  • Awareness of data privacy and PII handling in a compliance-conscious environment.
  • Background in nonprofit, public-sector, or social-science data.

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates