Senior Data Engineer
Job Description
Senior Data Engineer role with Deloitte in Boston, onsite, focused on delivering end-to-end data pipelines and solutions using Azure Data Factory, Databricks, and PySpark, along with client-facing collaboration and mentorship.
Responsibilities
- Maintain regular collaboration with Engagement Managers, project teams, and cross-functional stakeholders; escalate issues to engagement management when needed.
- Design, develop, and optimize ETL and ELT pipelines using Azure Data Factory and Databricks.
- Develop and tune PySpark and Spark SQL notebooks for large-scale data transformations.
- Architect end-to-end data solutions across development, UAT, and production environments with Unity Catalog governance.
- Lead design discussions with client architects and counterpart teams.
- Collaborate with multiple teams to establish data contracts and agreed schemas.
- Direct the design and optimization of high-volume data pipelines.
- Define and enforce data engineering standards, including naming conventions, partitioning strategies, cluster configurations, and Spark tuning.
- Drive performance enhancements through adaptive query execution tuning, clustering strategies, broadcast joins, and shuffle partition management.
- Design Databricks cluster policies, autoscaling configurations, and cost optimization strategies.
- Conduct root cause analysis on production incidents and implement durable fixes.
- Mentor junior and mid-level engineers via code reviews and pair programming.
- Evaluate new technologies and recommend adoption, including Delta Live Tables, Auto Loader, Serverless Compute, DABs, and Event Hubs.
Requirements
- Proficiency in Python, PySpark, Spark SQL, and SQL Server.
- Experience with Azure components such as Data Factory, Data Lake Storage Gen2, Key Vault, and Monitor.
- Experience with Databricks including Delta Lake, Unity Catalog, and Workflows.
- Familiarity with Apache Airflow for workflow orchestration.
- Git or Azure DevOps for version control and CI/CD.
- Deep knowledge of Spark internals, including DAG optimization, spill analysis, and skew handling.
- Delta Lake advanced features such as time travel, deletion vectors, and predictive I/O.
- Unity Catalog governance including row/column security, external locations, and system tables.
- Infrastructure as Code experience with Terraform and Azure ARM templates.
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related IT discipline, or equivalent experience.
- Limited immigration sponsorship may be available.
- Ability to travel approximately 10 percent, varying by client engagements.
Technologies
- Python
- PySpark
- Spark SQL
- SQL Server
- Azure: Data Factory, Data Lake Storage Gen2, Key Vault, Monitor
- Databricks: Delta Lake, Unity Catalog, Workflows
- Delta Live Tables (DLT)
- Auto Loader
- Serverless Compute
- DABs
- Event Hubs
- Apache Airflow
- Git
- Azure DevOps
- Terraform
- Azure ARM templates
Team
AI and Engineering leverages cutting-edge capabilities to build and operate integrated data, software, AI, network, and hybrid cloud solutions. The team focuses on transforming mission-critical operations and modernizing clients' technology and data platforms through tailored delivery models.
Additional Information
Accommodation information for applicants requiring assistance: https://www2.deloitte.com/us/en/pages/careers/articles/join-deloitte-assistance-for-disabled-applicants.html