Senior Data Engineer - Databricks
Job Description
Senior Data Engineer to design and build enterprise-grade data platforms, services, and pipelines for clients, requiring strong communication, customer service, and problem-solving skills. The role is onsite in McLean, Virginia, with a salary range of USD 125,000 to 165,000 per year. A Master’s degree is required, and a minimum of 8 years of relevant experience is expected.
Overview
We are seeking a seasoned Senior Data Engineer to collaborate with our team and client stakeholders to develop robust data platforms, services, and pipelines. The ideal candidate combines technical excellence with strong communication and customer-service capabilities, along with a genuine passion for data and solving complex problems.
Responsibilities
- Lead and architect migration of data environments with emphasis on performance and reliability.
- Assess and interpret ETL jobs, workflows, BI tools, and reports to inform data platform decisions.
- Respond to technical inquiries related to customization, integration, enterprise architecture, and general feature functionality of data products.
- Design and implement database and data warehouse solutions on Databricks.
- Operate with core skills in Databricks, Python, and SQL.
- Support an Agile software development lifecycle.
- Contribute to the growth of the Data Exploitation Practice.
Requirements
- Ability to hold a position of public trust with the US government.
- 12+ years of experience with a Bachelor’s degree or 8+ years with a Master’s degree.
- 8-12 years of industry experience in data engineering and a demonstrated problem-solving mindset.
- 8-12 years of direct data engineering experience, including:
- Designing and building data pipelines and data products on modern lakehouse platforms (e.g., Databricks, Apache Spark), including Delta Lake, distributed processing, and performance optimization.
- Supporting or integrating GenAI/ML workflows (feature engineering, vector storage, RAW pipelines, or model lifecycle using MLFlow).
- Operating in enterprise or regulated cloud environments with platform constraints, security controls, and data governance frameworks.
- Building batch and streaming data pipelines using distributed processing frameworks (e.g., Apache Spark Structured Streaming) within a lakehouse architecture.
- Experience with search and retrieval patterns for analytics or AI use cases (e.g., indexing, semantic search, vector-based retrieval).
- Strong proficiency in SQL and Python for data engineering, including distributed data processing frameworks (e.g., Apache Spark).
- Advanced SQL knowledge and experience with relational databases, query authoring and optimization, and familiarity with diverse databases.
- Ability to lead technical discussions with both technical and non-technical stakeholders, translating business needs into actionable data solutions and guiding decisions.
- Proven ability to evaluate feasibility, identify risks, and make pragmatic recommendations in complex or ambiguous environments.
- Experience navigating cross-team environments and driving alignment through clear communication and structured problem-solving.
- Experience constructing complex queries to analyze results using databases or data processing development environments.
- Strong judgment in evaluating feasibility, balancing technical constraints, business needs, and long-term scalability.
- Experience architecting data systems that span transactional and data warehouses.
- Experience aggregating results and compiling information for reporting from multiple datasets.
- Experience working in an Agile environment.
- Experience supporting project teams of developers and data scientists who build web interfaces, dashboards, reports, and analytics or machine learning models.
Technologies
- Databricks
- Python
- SQL
- Apache Spark
- Delta Lake
- MLFlow
- Apache Spark Structured Streaming
Contributions
- Lead and design migration of data environments with a focus on performance and reliability.
- Assess and interpret ETL jobs, workflows, BI tools, and reports to guide data platform decisions.
- Address technical inquiries related to customization, integration, enterprise architecture, and data product functionality.
- Develop database and data warehouse solutions on Databricks using core skills in Databricks, Python, and SQL.
- Support an Agile software development lifecycle and contribute to the growth of the Data Exploitation Practice.
About Steampunk
Steampunk determines compensation based on geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this role is $125,000 to $165,000 per year. The presented estimate reflects a typical annual salary for this position, and annual salary is one element of Steampunk's total benefits package. Additional Steampunk benefits are available; more information can be provided upon request.
Identity Statement
As part of the application process, you are expected to participate on camera during interviews and assessments. Steampunk reserves the right to capture your image to verify identity and prevent fraud.