Senior Data Engineer
Job Description
At Deloitte, the Senior Data Engineer role sits at the intersection of data architecture and delivery for client engagements in Austin. This position blends hands-on pipeline development with architectural leadership, focusing on ETL/ELT design and optimization, PySpark and Spark SQL notebook work, and guiding data engineering efforts within the Project Delivery Model.
Location: Austin, TX (onsite)
Salary: USD 95,000 - 150,000 per year
Education: Bachelor's degree
Responsibilities
- Maintain regular communication with Engagement Managers (Directors), project teams, and stakeholders across functional and technical groups, escalating issues requiring higher-level input from engagement management.
- Design, build, and optimize end-to-end ETL/ELT pipelines using Azure Data Factory and Databricks.
- Author and tune PySpark and Spark SQL notebooks for large-scale data transformations.
- Architect data solutions spanning development, UAT, and production environments with Unity Catalog.
- Lead design discussions with client architects and other counterparts.
- Collaborate with multiple teams to establish data contracts and schema governance.
- Lead the design and optimization of high-volume data pipelines.
- Define and enforce data engineering standards including naming conventions, partitioning strategies, cluster configurations, and Spark tuning.
- Drive performance optimization through AQE tuning, liquid clustering, broadcast joins, and shuffle partition management.
- Design Databricks cluster policies, autoscaling configurations, and cost-optimization strategies.
- Conduct root-cause analysis on production incidents and implement durable fixes.
- Mentor junior and mid-level engineers via code reviews and pair programming.
- Evaluate new technologies and recommend adoption, including DABs, DLT, Auto Loader, Serverless Compute, and event hubs.
Requirements
- Proficiency in Python, PySpark, Spark SQL, and SQL Server.
- Experience with Azure services such as Azure Data Factory (ADF), ADLS Gen2, Key Vault, and Azure Monitor.
- Hands-on experience with Databricks including Delta Lake, Unity Catalog, and Workflows.
- Apache Airflow experience.
- Git and Azure DevOps for version control and CI/CD.
- Deep understanding of Spark internals, including DAG optimization, spill analysis, and data skew handling.
- Delta Lake advanced features such as time travel, deletion vectors, and predictive I/O.
- Unity Catalog governance covering row/column security, external locations, and system tables.
- Infrastructure as Code using Terraform and Azure ARM templates.
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or related IT discipline, or equivalent experience.
- Limited immigration sponsorship may be available.
- Ability to travel approximately 10% of the time, depending on client needs.
Technologies
- Python
- PySpark
- Spark SQL
- SQL Server
- Azure
- Azure Data Factory (ADF)
- ADLS Gen2
- Key Vault
- Azure Monitor
- Databricks
- Delta Lake
- Unity Catalog
- Workflows
- Apache Airflow
- Git
- Azure DevOps
- DABs
- DLT
- Auto Loader
- Serverless Compute
- Event Hubs
- Terraform
- Azure ARM templates