Data Engineer - Managing Consultant
Job Description
Guidehouse is seeking a Data Engineer - Managing Consultant to design and deliver scalable, governed data pipelines and analytics platforms using Databricks and AWS. The role emphasizes security, observability, and enterprise data availability, and is based onsite in McLean, VA. A bachelor’s degree and extensive cloud data engineering experience are required.
Responsibilities
- Design, implement, and optimize scalable, production-grade data ingestion, transformation, and analytics pipelines with Databricks components such as Delta Lake, Delta Live Tables, and Auto Loader, alongside AWS services to provide trusted, timely data across enterprise use cases.
- Develop standardized pipelines supporting batch and near real-time processing, integrating legacy data sources with modern cloud-native services to improve data availability and close access gaps.
- Lead the full delivery lifecycle from intake and discovery through source profiling and technical design, ensuring requirements traceability and alignment with Architecture Review Board governance.
- Build and maintain governed data pipelines with embedded metadata, lineage, and data quality controls to meet technical, security, and documentation standards before production deployment.
- Create and operationalize data engineering frameworks that include observability, monitoring, alerting, and resilient error handling to sustain platform stability and target 99.9% availability.
- Collaborate with platform engineering and cloud operations teams to integrate pipelines with AWS services (S3, Glue, Kafka/Kinesis, APIs) for secure, scalable data movement and cross-platform interoperability.
- Enable governed analytics and self-service data access through curated datasets, semantic layers, and SQL warehouse integration to support enterprise reporting, dashboards, and advanced analytics.
- Apply security and compliance controls aligned with IRS cybersecurity policies, including RBAC/ABAC, data masking, encryption, and audit logging to protect sensitive data and maintain regulatory compliance.
Requirements
- Bachelor’s degree is required.
- Minimum eight (8) years of data engineering experience in cloud-based environments.
- Minimum four (4) years of hands-on Databricks experience designing scalable data pipelines.
- Advanced proficiency in Python, PySpark, and SQL for large-scale data processing within lakehouse architectures.
- Experience building and optimizing data pipelines using Delta Lake, medallion architecture, and modern ingestion frameworks for both batch and streaming data.
- Strong experience with AWS data platforms and services (S3, IAM, VPC, Glue, streaming frameworks) and integration with enterprise data ecosystems.
- Proven ability delivering production-ready data solutions with metadata management, lineage, data quality, and observability frameworks.
- Familiarity with DevSecOps and CI/CD practices, including automation, testing, and deployment in cloud data environments.
- Knowledge of data governance, security, and compliance requirements in regulated environments, including FISMA and FedRAMP High.
Technologies
- Databricks, Delta Lake, Delta Live Tables, Auto Loader
- AWS, S3, Glue, IAM, VPC
- Kafka, Kinesis, APIs, Python, PySpark, SQL
- Delta Sharing, Unity Catalog, Genie, Clean Rooms
- Informatica EDC/Axon, EventBridge, Terraform, GitHub Actions
- Databricks Asset Bundles, Photon, Z-ordering, liquid clustering
Benefits
- Medical, Rx, Dental & Vision Insurance
- Personal and Family Sick Time & Company Paid Holidays
- Discretionary variable incentive bonus eligibility
- Parental Leave and Adoption Assistance
- 401(k) Retirement Plan
- Basic and Supplemental Life Insurance
- Health Savings Account, Dental/Vision & Dependent Care Flexible Spending Accounts
- Short-Term & Long-Term Disability
- Student Loan PayDown
- Tuition Reimbursement, Personal Development & Certifications
- Employee Referral Program
- Corporate Sponsored Events & Community Outreach
- Emergency Back-Up Childcare Program
- Mobility Stipend
Travel
Up to 10% travel may be required.
Clearance
Ability to Obtain Public Trust.
What would be nice to have
- Experience supporting federal data platforms or large-scale enterprise data modernization in regulated environments such as IRS or Treasury.
- Hands-on experience with Databricks Unity Catalog, Delta Sharing, Genie, and Clean Rooms for governed data access and collaboration.
- Experience implementing streaming and near real-time data pipelines using Kafka, Kinesis, EventBridge, or similar technologies.
- Familiarity with Informatica EDC/Axon or enterprise metadata/catalog tooling.
- Proficiency in performance optimization techniques (Photon, Z-ordering, liquid clustering) for large-scale workloads.
- Exposure to CI/CD automation using Terraform, GitHub Actions, and Databricks Asset Bundles for infrastructure and pipeline deployment.
- Advanced cloud or Databricks certifications (AWS, Databricks) in good standing.