EngineerJobs.io
← Back to all jobs

Job Description

Protective is hiring a Lead Data Engineer to define the technical direction for a delivery pod building data products on Voyager, its Databricks lakehouse on Azure. This is a hands-on leadership role focused on end-to-end design across the Bronze/Raw, Silver/Prep, and Gold/Prod medallion layers, with clear accountability for standards, quality, contract integrity, and consumer compatibility.

As the pod’s technical lead, you will shape how data is ingested, cleaned and conformed, modeled, and published, including dimensional design and the Gold layer that downstream consumers query and depend on.

Responsibilities

  • Lead end-to-end design of data products, including ingestion, cleansing and conformance, modeling, and the consumer-facing Gold layer.
  • Own dimensional design, covering grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across products.
  • Set the pod’s approach to where logic belongs, including what is handled in Silver, what belongs in Gold, and what is a consumer responsibility.
  • Keep models aligned to the questions they answer and push back on designs that will not hold over time.
  • Partner with ML engineering when Gold datasets act as training or feature sources so datasets are contracted, versioned, and reproducible.
  • Own ODCS data contracts as enforceable interfaces, including named owners and consumers, quality rules, freshness expectations, and an explicit breaking-change policy.
  • Assess compatibility for proposed contract changes and drive consumer notification when changes are genuinely breaking.
  • Represent pod contracts in cross-domain discussions where one pod’s Gold layer becomes another team’s dependency.
  • Set and uphold engineering standards across Python, SQL, dbt, testing, model structure, naming, and repository conventions, aligned with paved paths and Azure DevOps CI gates.
  • Lead code review and ensure quality rules are enforced through tests and asset checks, not just documentation.
  • Maintain observable pipeline health with instrumented freshness, volume, latency, and cost, plus alerting against contractual SLAs/SLOs.
  • Own operational readiness for pod pipelines, including failure diagnosis, triage, backfills, tuning, on-call coverage, escalation, root-cause analysis, and runbooks.
  • Keep delivery within regulated carrier control expectations, including change management via pull requests and pipeline gates, segregation of duties, least-privilege access, and CI/CD audit evidence.
  • Develop reusable frameworks, templates, and patterns to improve consistency and delivery speed.
  • Collaborate with the Product Owner and Scrum Master to refine and decompose work into estimable stories with testable acceptance criteria, including target layer and repository.
  • Hold Definition of Ready and Definition of Done expectations for merged, approved code, passing CI and coverage gates, and evidence that outcomes are real.
  • Identify unknowns that require spikes instead of estimates and elevate them during planning.
  • Grow engineers through design review, pairing, and code review, and reduce single points of knowledge.
  • Partner with the platform team to address capability gaps as demand signals instead of creating private workarounds.
  • Partner with the DataOps/MLOps Lead on shared standards for CI/CD, orchestration, and observability, and advocate for what the pod still needs.
  • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements.

Requirements

  • Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered.
  • 6+ years building and operating production data pipelines and consumer-facing data models from ingestion through published data products.
  • Strong hands-on Python and SQL, with credibility to make design calls and willingness to write and review production code.
  • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning.
  • Deep dimensional modeling experience, including grain, keys, slowly changing dimensions, facts and dimensions, and conformed dimensions.
  • Demonstrated technical leadership through setting standards, leading design, and elevating other engineers’ work.
  • Experience owning data other teams depend on, including breaking changes and production data incidents.
  • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaboration, code review, and CI/CD (Azure DevOps or comparable).
  • Experience defining observability and SLA/SLO expectations for shared data and managing incidents and communication when expectations are missed.
  • Ability to explain trade-offs clearly to engineers and business stakeholders, including saying no to designs that will not hold.

Technologies

  • Databricks, Azure, Spark-based lakehouse, Delta Lake, MERGE, Incremental processing
  • Python, SQL, dbt, Dagster, Databricks Workflows, Airflow
  • Git, CI/CD, Azure DevOps, Unity Catalog
  • ODCS, MLOps, MLflow, model registries, model serving, dlt (dltHub)
  • Great Expectations, Monte Carlo

Benefits

  • Comprehensive health, dental and vision insurance
  • Mental health benefits and an employee assistance program
  • Paid time away benefits (e.g., paid time off, paid parental leave, short-term disability, and a cultural observance day)
  • Contributions to healthcare accounts
  • Pension plan
  • 401(k) plan with Company matching
  • ProHealth Rewards to improve wellbeing while earning cash rewards

Preferred Qualifications

  • Databricks certification (Data Engineer Professional or equivalent demonstrated depth)
  • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage
  • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake
  • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments
  • Practical experience with data contracts, ODCS, or data-mesh style data product ownership
  • Experience with declarative Python ingestion such as dlt (dltHub) or comparable
  • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar
  • Familiarity with MLOps practice including MLflow, model registries, and model serving
  • Azure and Azure DevOps
  • Experience in financial services, insurance, or another regulated industry involving data access, lineage, and audit expectations
  • Experience introducing AI-assisted development into a team’s normal workflow in a disciplined way

Location: Remote (remote)

Compensation: USD 109,500 - 167,833 per year

If you require an accommodation to complete the application and recruitment process due to a disability, email [email protected].

Similar Jobs