EngineerJobs.io
← Back to all jobs

Job Description

Pyx Health Inc is hiring a Data Engineer to build and maintain Azure-based data infrastructure and dependable data pipelines. In this remote role, you will work with Databricks, Airflow, Python, and SQL to support healthcare data workflows and analytic use cases.

Key Responsibilities

  • Build and maintain data pipelines for ingesting, transforming, and cleaning healthcare data in the Azure cloud using Databricks, PySpark, and Delta Lake
  • Implement pipeline logic aligned to defined specifications, with architectural guidance from senior engineers
  • Create reusable testing frameworks to support pipeline reliability
  • Monitor pipelines for failures and performance issues, escalating complex problems appropriately
  • Develop and maintain Airflow DAGs with error handling and retry logic
  • Support pipeline deployment and configuration on Azure using Astronomer with the Astro CLI
  • Improve pipeline reliability and reduce manual intervention through ongoing enhancements
  • Implement data models in Delta Lake, including merge/upsert patterns for efficient storage and retrieval
  • Design and optimize Delta Lake tables for analytic workloads
  • Use Unity Catalog to support data governance, organization, and security
  • Apply Change Data Capture (CDC) patterns with direction from senior team members
  • Write and optimize T-SQL and Databricks SQL scripts, stored procedures, and ETL support
  • Work within the Azure ecosystem including ADLS Gen2, Key Vaults, Logic Apps, and Azure DevOps
  • Follow established security and scalability standards when building data infrastructure
  • Support infrastructure tasks as needed with guidance on architecture
  • Ensure scripts and datasets are documented to support enterprise data governance
  • Design, implement, and continuously improve automated data quality monitoring, reconciliation, and alerting for reliable downstream reporting
  • Troubleshoot and resolve pipeline failures, documenting root causes and resolutions
  • Collaborate with data scientists, analysts, and business stakeholders to clarify reporting requirements, address issues, and improve data assets
  • Participate in code reviews and incorporate feedback to improve code quality
  • Document pipelines, processes, and implementation decisions clearly and consistently
  • Stay current with data engineering technologies, with particular attention to the healthcare space

Required Qualifications

  • 2–4 years of experience as a Data Engineer or in a closely related data role
  • Hands-on experience with Azure cloud services (ADLS, Databricks, or similar)
  • Working knowledge of SQL for scripting and data modeling, including T-SQL, Spark, and Databricks SQL
  • Ability to contribute to technical projects with moderate oversight
  • Strong communication skills, including comfortable status updates and asking questions with cross-functional partners
  • Familiarity with CI/CD concepts, Git, and version control workflows
  • Solid problem-solving skills with a systematic approach to debugging
  • Familiarity with healthcare data standards and regulations (HIPAA, HL7, etc.) is a plus
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent work experience)

Technical Requirements (Must Haves)

  • Databricks SQL (Mid): Delta Lake fundamentals, basic merge/upsert patterns, familiarity with CDC concepts
  • Python (Mid): Pipeline logic, data transformation, SQL scripting; some PySpark/Spark DataFrame experience
  • Databricks Spark Notebooks (Mid): Notebook-based development, basic cluster usage, Delta table operations
  • Airflow Python Development (Mid): Write and maintain DAGs; understanding of error handling and retry patterns
  • Airflow Astro Configuration (Foundational): Familiarity with Astronomer or willingness to learn; basic Astro CLI usage
  • Azure Ecosystem (Foundational): Working knowledge of ADLS Gen2, Key Vaults, and Azure DevOps
  • T-SQL (Foundational–Mid): Basic stored procedures, SQL Server querying, and data manipulation

Nice to Have

  • Azure Data Factory (Foundational): Basic familiarity with pipeline authoring and triggers
  • Git + Azure DevOps CI/CD (Foundational): Experience with branching, pull requests, peer review processes, version control best practices, and deploying production data pipelines using CI/CD workflows

Technologies

  • Azure, ADLS Gen2, Azure DevOps, Key Vaults, Logic Apps
  • Databricks, Databricks SQL, Databricks Spark Notebooks, Delta Lake, Delta tables
  • Airflow, Astronomer, Astro CLI
  • Python, PySpark, Spark
  • SQL, T-SQL, SQL Server
  • Unity Catalog
  • ETL
  • Change Data Capture (CDC)
  • Git, CI/CD, Git + Azure DevOps CI/CD
  • Azure Data Factory

Role Details

  • Location: Remote
  • Minimum Experience: 2 years

Similar Jobs