Senior Data Engineer
Job Description
Senior Data Engineer at Deloitte in Arlington, VA onsite, focusing on designing and optimizing end-to-end data pipelines with Azure Databricks, ADF, and PySpark across development, UAT, and production environments.
Responsibilities
- Maintain ongoing collaboration with Engagement Managers, project teams, and stakeholders across functional and technical groups, escalating issues to engagement management when needed.
- Design, develop, and optimize ETL and ELT pipelines using Azure Data Factory and Databricks.
- Develop and tune PySpark and Spark SQL notebooks to perform large-scale data transformations.
- Architect comprehensive data solutions spanning development, UAT, and production environments with Unity Catalog.
- Lead design discussions with client architects and other counterparts to align on technical direction.
- Work with multiple teams to establish data contracts and agreed-upon schemas.
- Lead the design and optimization of high-volume data pipelines.
- Define and enforce data engineering standards, including naming conventions, partitioning strategies, cluster configurations, and Spark tuning.
- Drive performance improvements through AQE tuning, dynamic clustering considerations, broadcast joins, and shuffle partition management.
- Design Databricks cluster policies, autoscaling rules, and cost optimization strategies.
- Investigate and resolve production incidents via root-cause analysis and implement durable fixes.
- Mentor junior and mid-level engineers through code reviews and pair programming.
- Assess new technologies and recommend adoption, such as Delta Live Tables, Auto Loader, serverless compute, or event hubs.
Requirements
- Proficiency in Python, PySpark, Spark SQL, and SQL Server.
- Experience with Azure services including Data Factory, Data Lake Storage Gen2, Key Vault, and Monitor.
- Hands-on use of Databricks features such as Delta Lake, Unity Catalog, and Workflows.
- Familiarity with Apache Airflow for workflow orchestration.
- Version control and CI/CD using Git and Azure DevOps.
- Deep knowledge of Spark internals, including DAG optimization, spill analysis, and skew handling.
- Experience with advanced Delta Lake capabilities (time travel, deletion vectors, predictive I/O).
- Unity Catalog governance skills, including row and column security, external locations, and system tables.
- Infrastructure as Code experience with Terraform and Azure Resource Manager templates.
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related IT discipline, or equivalent experience.
- Limited immigration sponsorship may be available.
- Ability to travel approximately 10 percent, depending on client assignments.
Technologies
- Python, PySpark, Spark SQL
- SQL Server
- Azure Data Factory
- Azure Data Lake Storage Gen2
- Key Vault
- Azure Monitor
- Databricks (Delta Lake, Unity Catalog, Workflows)
- Delta Live Tables, Auto Loader, Serverless Compute
- Azure Event Hubs
- Git, Azure DevOps
- Terraform, Azure ARM templates
- Delta Lake advanced features
Team
The AI and Engineering group leverages advanced engineering practices to design, deploy, and operate integrated solutions across software, data, AI, networking, and hybrid cloud infrastructures. The team focuses on transforming mission-critical operations and enabling clients to modernize technology and data platforms while delivering client-centered outcomes.
Additional information
Accommodation for applicants with a need for assistance is available at the Deloitte accessibility page: https://www2.deloitte.com/us/en/pages/careers/articles/join-deloitte-assistance-for-disabled-applicants.html