Databricks Data Engineer
Artificial Intelligence
Big Data
Bigdata
Cloud Data Engineering
Cloud Data Platform
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Governance
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Security
Data Warehouse
Database
Databases
Databricks
Databricks Lakehouse
Delta Lake
ETL
Informatica
Information Technology (IT)
Integration
Kafka
Programming
Programming Languages
Spark
Spark Structured Streaming
SQL
Stream Processing
Streaming Data
Job Description
This hands-on role supports the design, build, and operation of scalable data and analytics solutions on the Databricks Lakehouse platform for federal mission needs. The position focuses on batch and streaming pipelines, Delta Lake using medallion architecture, and governance and security through Unity Catalog.
Responsibilities
- Design, develop, and maintain scalable batch and streaming data pipelines using Databricks, Apache Spark, PySpark, and Spark SQL.
- Build and manage Delta Lake tables using the medallion (bronze/silver/gold) architecture to deliver reliable, analytics-ready data.
- Develop real-time and near-real-time ingestion solutions using Spark Structured Streaming and messaging platforms such as Kafka.
- Configure and manage Databricks clusters, jobs, and workflows in production environments.
- Implement data governance, access controls, and security best practices using Unity Catalog.
- Integrate data from multiple source systems and destinations, supporting ETL/ELT and pipeline orchestration.
- Optimize existing data workflows and Spark jobs for performance, reliability, and cost efficiency.
- Integrate Databricks development with CI/CD pipelines and enterprise SDLC tooling, including Git-based version control.
- Collaborate with data scientists and analysts to define data models and support machine learning and AI use cases, including model lifecycle management with MLflow.
- Support advanced analytics use cases such as anomaly detection, risk scoring, and fraud analytics.
- Monitor and troubleshoot data processing jobs, implementing data quality checks and observability to support high availability.
- Document data processes, frameworks, pipelines, and data mappings for technical and non-technical audiences.
- Work with scrum teams, product owners, and client stakeholders to deliver end-to-end data solutions.
- Stay current on Databricks platform capabilities and industry trends to recommend appropriate tools and technologies.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience).
- 4+ years of experience in data engineering, analytics engineering, or big data development.
- 2+ years of hands-on experience with the Databricks platform.
- Proficiency in Apache Spark, PySpark, and Spark SQL.
- Experience with Databricks clusters, jobs/workflows, Delta Lake, and Unity Catalog in production environments.
- Experience with medallion architecture and Spark Structured Streaming.
- Strong Python and SQL skills for data engineering and data analysis.
- Experience with ETL/ELT processes and data pipeline orchestration.
- Familiarity with cloud platforms such as AWS, Azure, or Google Cloud and their native data services.
- Experience integrating data solutions with CI/CD pipelines and Git-based version control workflows.
- Understanding of data governance, security, and access control best practices.
- Experience working in Agile development environments.
- Ability to obtain and maintain a Public Trust determination.
- Excellent analytical, problem-solving, and communication skills, including the ability to work with both technical and non-technical stakeholders.
Technologies
- Databricks, Apache Spark, PySpark, Spark SQL
- Delta Lake, Unity Catalog, medallion architecture
- Spark Structured Streaming, Kafka
- ETL/ELT, CI/CD, Git-based version control
- MLflow, Python, SQL
- AWS, Azure, Google Cloud
- ML/AI use cases
Preferred Qualifications
- Databricks certification (e.g., Databricks Certified Data Engineer Associate/Professional) or a cloud platform certification.
- Experience implementing ML or AI solutions in Databricks, including MLflow-based model lifecycle management.
- Knowledge of machine learning, AI, or Natural Language Processing (NLP) techniques, including text mining.
- Experience supporting fraud analytics, risk scoring, or anomaly detection.
- Experience with distributed data and streaming tools such as Kafka, Hadoop, Hive, or Amazon EMR.
- Experience with data quality frameworks and observability/monitoring tooling.
- Experience with NoSQL databases.
- Experience with visualization packages such as Plotly, Seaborn, or ggplot2.
- Experience supporting federal government or regulated-industry programs, especially the Department of Veterans Affairs.