Lead Data Engineer
Big Data
Bigdata
Cloud Data Engineering
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Integration
Data Lake
Data Pipeline
Data Platform
Data Processing
Database
Databases
Databricks
Databricks Genie
Databricks Lakeflow
Databricks Pyspark
ETL
Informatica
Integration
Lakehouse
Spark
SQL
Job Description
Lead the engineering direction and delivery of a Databricks-based lakehouse, including streaming analytics capabilities.
Responsibilities
- Define the data architecture with team and stakeholders, setting technical direction for lakehouse, ingestion, and transformation patterns
- Identify, escalate, and remove technical blockers through hands-on problem solving, design decisions, and coordination across teams
- Design and build enterprise data models (conceptual, logical, physical) across domains, aligned to lakehouse and pipeline architecture
- Establish and enforce data modeling standards, best practices, and governance processes, including using Erwin
- Define data requirements and perform wrangling of large scale structured and unstructured data, validating outputs by running tools in the data environment
- Standardize and customize data analysis, develop mechanisms to ingest, analyze, validate, normalize, and clean data
- Create data policies and develop interfaces and retention models, including synthesizing or anonymizing data
- Implement statistical data quality procedures on new sources and support iterative analytics to enable data science and insights creation
- Develop and maintain data engineering best practices, contributing to insights across data analytics and visualization concepts and techniques
- Lead the data engineering team to build scalable data architecture and high-performance pipelines using state-of-the-art Big Data tooling
- Build pipelines to clean, transform, and aggregate data from disparate sources
- Use scripting languages and other tooling to connect systems together
- Apply data architecture knowledge to lead project teams from requirements to implementation
- Partner with data science and business intelligence teams on models and pipelines for research, reporting, and machine learning
- Acquire and process high-velocity real-time and batch data across different sizes and scales
Requirements
- 10+ years in data engineering with deep, hands-on Databricks experience (Spark/PySpark pipelines, Delta Lake merges and schema evolution, Delta Live Tables, Unity Catalog governance)
- Hands-on experience with Genie for self-service analytics and LakeFlow Designer for visual ETL orchestration
- 4+ years building reliable batch/streaming Spark workloads
- 3+ years with Kafka or Confluent for high-volume event processing
- 3+ years working with AWS analytics services (S3, IAM, Glue/Lambda/MSK)
- Proven ability to design conceptual, logical, and physical data models with strong governance and advanced SQL skills for business-ready datasets
- Track record leading technical design, mentoring engineers, driving architectural consensus, and unblocking teams on complex challenges
- Experience delivering ETL/ELT pipelines end-to-end in a lakehouse environment and working within Agile frameworks (Scrum/Kanban/SAFe) to manage iterative delivery and cross-team dependencies
Technologies
- Databricks, Spark, PySpark, Delta Lake, Delta Live Tables, Unity Catalog, Genie, LakeFlow Designer
- Kafka, Confluent
- AWS analytics services (S3, IAM, Glue/Lambda/MSK)
- SQL, Erwin, Python, Scala
- PySpark/Scala-Spark
- Snowflake, BigQuery, HBase, Cassandra
- Azure, Kinesis, TIBCO EMS, IBM MQ Series, MSK
- GIT, REST API, Web Services, NoSQL
- ETL/ELT, Agile frameworks (Scrum/Kanban/SAFe)
Work Conditions
- Location: Atlanta, GA, US
- Work model: Hybrid (two days in office; three days remote)
- Shift work: No
- On-call: Yes
- Weekend work: No