EngineerJobs.io
← Back to all jobs

Job Description

Lead the engineering direction and delivery of a Databricks-based lakehouse, including streaming analytics capabilities.

Responsibilities

  • Define the data architecture with team and stakeholders, setting technical direction for lakehouse, ingestion, and transformation patterns
  • Identify, escalate, and remove technical blockers through hands-on problem solving, design decisions, and coordination across teams
  • Design and build enterprise data models (conceptual, logical, physical) across domains, aligned to lakehouse and pipeline architecture
  • Establish and enforce data modeling standards, best practices, and governance processes, including using Erwin
  • Define data requirements and perform wrangling of large scale structured and unstructured data, validating outputs by running tools in the data environment
  • Standardize and customize data analysis, develop mechanisms to ingest, analyze, validate, normalize, and clean data
  • Create data policies and develop interfaces and retention models, including synthesizing or anonymizing data
  • Implement statistical data quality procedures on new sources and support iterative analytics to enable data science and insights creation
  • Develop and maintain data engineering best practices, contributing to insights across data analytics and visualization concepts and techniques
  • Lead the data engineering team to build scalable data architecture and high-performance pipelines using state-of-the-art Big Data tooling
  • Build pipelines to clean, transform, and aggregate data from disparate sources
  • Use scripting languages and other tooling to connect systems together
  • Apply data architecture knowledge to lead project teams from requirements to implementation
  • Partner with data science and business intelligence teams on models and pipelines for research, reporting, and machine learning
  • Acquire and process high-velocity real-time and batch data across different sizes and scales

Requirements

  • 10+ years in data engineering with deep, hands-on Databricks experience (Spark/PySpark pipelines, Delta Lake merges and schema evolution, Delta Live Tables, Unity Catalog governance)
  • Hands-on experience with Genie for self-service analytics and LakeFlow Designer for visual ETL orchestration
  • 4+ years building reliable batch/streaming Spark workloads
  • 3+ years with Kafka or Confluent for high-volume event processing
  • 3+ years working with AWS analytics services (S3, IAM, Glue/Lambda/MSK)
  • Proven ability to design conceptual, logical, and physical data models with strong governance and advanced SQL skills for business-ready datasets
  • Track record leading technical design, mentoring engineers, driving architectural consensus, and unblocking teams on complex challenges
  • Experience delivering ETL/ELT pipelines end-to-end in a lakehouse environment and working within Agile frameworks (Scrum/Kanban/SAFe) to manage iterative delivery and cross-team dependencies

Technologies

  • Databricks, Spark, PySpark, Delta Lake, Delta Live Tables, Unity Catalog, Genie, LakeFlow Designer
  • Kafka, Confluent
  • AWS analytics services (S3, IAM, Glue/Lambda/MSK)
  • SQL, Erwin, Python, Scala
  • PySpark/Scala-Spark
  • Snowflake, BigQuery, HBase, Cassandra
  • Azure, Kinesis, TIBCO EMS, IBM MQ Series, MSK
  • GIT, REST API, Web Services, NoSQL
  • ETL/ELT, Agile frameworks (Scrum/Kanban/SAFe)

Work Conditions

  • Location: Atlanta, GA, US
  • Work model: Hybrid (two days in office; three days remote)
  • Shift work: No
  • On-call: Yes
  • Weekend work: No

Similar Jobs