Data Engineer
Apache Airflow
Application Security
Automation
Big Data
Bigdata
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Data Warehouse
Cloud Data Warehouse
Cloud Infrastructure
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Integration
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Data Warehousing
Database
Databases
Databricks
DevOps
DevSecOps
ETL
Facilities Management
Google Cloud Bigquery
Informatica
Information Technology (IT)
Infrastructure As Code
Programming Language
Programming Languages
Security
Security Automation
Snowflake
Software Security
Spark
SQL
Stream Processing
Streaming Data
Workflow Orchestration
Job Description
Function Health is hiring a Data Engineer to help build platform engineering for safe, fast changes across data, analytics, and ML.
Responsibilities
- Own the event pipeline that powers product analytics, experimentation, and feature gates, with schemas enforced at the source so invalid events never turn into invalid metrics.
- Build and maintain the Databricks lakehouse layers (Bronze, Silver, Gold), including automated schema evolution, contract tests, and backfills designed to be low-risk.
- Deliver freshness and volume monitoring driven from the data contract, rather than added later as separate checks.
- Enable end-to-end feature computation and serving plus training and evaluation pipelines.
- Implement the plumbing to move model outputs back into the product while maintaining the same standards for testing and observability.
- Advance the self-service story: templates, local dev and preview environments, policy-as-code for PHI, and ownership routing for alerts.
- Support progressive gates that allow exploratory model work to move quickly while member-facing changes receive more scrutiny.
Requirements
- Built internal platform or infrastructure that other engineers adopted; you’ve felt the difference between shipping a tool and getting it used.
- Operated production data or ML systems, including on-call responsibilities and fixing issues under pressure.
- Strong Python and SQL.
- Comfort working in a lakehouse environment; Databricks is used, and Snowflake or BigQuery knowledge transfers.
- Designed interfaces and schemas other teams depend on, and evolved them without breaking downstream consumers.
- Thought deeply about testing and CI for data or ML, including cases where correctness is statistical and failures may be silent.
- Experience: 1 to 4 years; emphasis is on what you’ve built, not just tenure.
Technologies
- Python, SQL, Databricks
- Snowflake, BigQuery
- dbt, DLT, Dagster, Airflow
- Kafka
- Spark Structured Streaming
- Terraform
Nice-to-Have Skills
- Declarative pipeline frameworks: dbt, DLT, Dagster, Airflow
- Streaming: Kafka, Spark Structured Streaming
- Data contracts, data diffing, or lineage tooling
- Terraform
- Feature stores
- MLOps and eval tooling
- Agentic coding workflows
- Healthcare experience
- PHI, including HIPAA experience
Core Values
- Ruthless Prioritization: move quickly to drive value, prioritize impact, and maintain standards of excellence.
- Member-First, Always: responsive delivery focused on peace of mind and outcomes.
- One Team, Moving Fast: aligned purpose, diverse perspectives, clear communication, and shared goals.
- Radical Ownership, Relentless Execution: urgency, precision, follow-through, pragmatism, and adopting new tech to improve outcomes.
- Mission Over Ego: alignment to mission, commit after disagreement, and operate with honesty and transparency.
- Sustained Integrity in Every Detail: clinical precision, accuracy, quality, and clarity.
Location: Remote
Minimum experience: 1 year