Data Engineer
Job Description
Based at Crumbl HQ in Provo, Utah, this Data Engineer role focuses on designing, building, and maintaining scalable data pipelines with dbt and Prefect to support data-driven decision making across the organization.
Responsibilities
- Design, implement, and maintain scalable data pipelines using ELT/ETL extraction methods to ensure reliability.
- Collaborate with data scientists, analysts, and other stakeholders to capture data requirements and uphold data quality.
- Create and maintain documentation such as data dictionaries, workflow diagrams, and data flow diagrams.
- Safeguard data integrity and security through appropriate controls and ongoing monitoring.
- Optimize pipelines for efficient processing and improved query performance.
- Implement and enforce data security policies, including access controls, encryption, and data masking.
- Design and implement data processing workflows with dbt and Prefect to support data science and machine learning initiatives.
- Develop and maintain data ingestion pipelines to bring external data into the organization’s data environment.
- Identify performance bottlenecks in pipelines and collaborate with infrastructure and operations teams to optimize performance.
- Test and validate data pipelines to verify correct operation and alignment with business requirements.
- Participate in code reviews and help establish engineering best practices.
- Stay informed about emerging data engineering and data science technologies and identify opportunities to adopt them internally.
Requirements
- Bachelor’s or Master’s degree in Data Science, Information Systems, or a related field
- 3+ years building and maintaining production data pipelines (degree in a related field or equivalent experience)
- Advanced SQL including window functions, CTEs, and performance tuning on large datasets
- Strong Python for data engineering with modular, testable pipeline code
- Hands-on dbt experience: models, tests, macros, and incremental materializations
- Production Snowflake experience: schema design, performance tuning, and warehouse/cost optimization
- AWS data services such as S3, Glue, and Lambda
- Data quality and observability with dbt plus Elementary
- Infrastructure as code with Terraform and version control with Git
- Dimensional data modeling (star/snowflake schemas, SCDs) and lakehouse concepts
- Strong problem-solving skills and clear communication with analysts, scientists, and stakeholders
Technologies
- dbt
- Prefect
- SQL
- Python
- Snowflake
- AWS S3
- AWS Glue
- AWS Lambda
- Terraform
- Git
- Elementary
Benefits
- Medical, dental, and vision benefits
- 15 days PTO per year
- 10 paid holidays
- Paid parental leave
- Personal phone bill reimbursement
- Gym reimbursement
- Corporate DoorDash DashPass membership
- Regular company and team activities
- 401k with competitive matching contribution plan
- Excellent opportunities for career growth
- Work in a hyper-growth company