EngineerJobs.io
← Back to all jobs

Job Description

Senior Data Engineer role with Deloitte in Boston, onsite, focused on delivering end-to-end data pipelines and solutions using Azure Data Factory, Databricks, and PySpark, along with client-facing collaboration and mentorship.

Responsibilities

  • Maintain regular collaboration with Engagement Managers, project teams, and cross-functional stakeholders; escalate issues to engagement management when needed.
  • Design, develop, and optimize ETL and ELT pipelines using Azure Data Factory and Databricks.
  • Develop and tune PySpark and Spark SQL notebooks for large-scale data transformations.
  • Architect end-to-end data solutions across development, UAT, and production environments with Unity Catalog governance.
  • Lead design discussions with client architects and counterpart teams.
  • Collaborate with multiple teams to establish data contracts and agreed schemas.
  • Direct the design and optimization of high-volume data pipelines.
  • Define and enforce data engineering standards, including naming conventions, partitioning strategies, cluster configurations, and Spark tuning.
  • Drive performance enhancements through adaptive query execution tuning, clustering strategies, broadcast joins, and shuffle partition management.
  • Design Databricks cluster policies, autoscaling configurations, and cost optimization strategies.
  • Conduct root cause analysis on production incidents and implement durable fixes.
  • Mentor junior and mid-level engineers via code reviews and pair programming.
  • Evaluate new technologies and recommend adoption, including Delta Live Tables, Auto Loader, Serverless Compute, DABs, and Event Hubs.

Requirements

  • Proficiency in Python, PySpark, Spark SQL, and SQL Server.
  • Experience with Azure components such as Data Factory, Data Lake Storage Gen2, Key Vault, and Monitor.
  • Experience with Databricks including Delta Lake, Unity Catalog, and Workflows.
  • Familiarity with Apache Airflow for workflow orchestration.
  • Git or Azure DevOps for version control and CI/CD.
  • Deep knowledge of Spark internals, including DAG optimization, spill analysis, and skew handling.
  • Delta Lake advanced features such as time travel, deletion vectors, and predictive I/O.
  • Unity Catalog governance including row/column security, external locations, and system tables.
  • Infrastructure as Code experience with Terraform and Azure ARM templates.
  • Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related IT discipline, or equivalent experience.
  • Limited immigration sponsorship may be available.
  • Ability to travel approximately 10 percent, varying by client engagements.

Technologies

  • Python
  • PySpark
  • Spark SQL
  • SQL Server
  • Azure: Data Factory, Data Lake Storage Gen2, Key Vault, Monitor
  • Databricks: Delta Lake, Unity Catalog, Workflows
  • Delta Live Tables (DLT)
  • Auto Loader
  • Serverless Compute
  • DABs
  • Event Hubs
  • Apache Airflow
  • Git
  • Azure DevOps
  • Terraform
  • Azure ARM templates

Team

AI and Engineering leverages cutting-edge capabilities to build and operate integrated data, software, AI, network, and hybrid cloud solutions. The team focuses on transforming mission-critical operations and modernizing clients' technology and data platforms through tailored delivery models.

Additional Information

Accommodation information for applicants requiring assistance: https://www2.deloitte.com/us/en/pages/careers/articles/join-deloitte-assistance-for-disabled-applicants.html

Similar Jobs