Data Engineer
Azure Data Factory
Azure Data Platform
Azure Platform
Big Data
Bigdata
Data
Data Analysis
Data Architecture
Data Engineer
Data Factory
Data Factory Azure
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Database
Databases
ETL
Informatica
Information Technology (IT)
Integration
Microsoft Azure
Microsoft Fabric
Programming
Programming Language
Programming Languages
Pyspark
Rag Architectures
Spark
SQL
Job Description
Baker Group is seeking a Data Engineer to design, build, and maintain Microsoft Fabric data pipelines and infrastructure, enabling the organization’s data platform to function as a single source of truth. This role will own end-to-end ingestion, transformation, orchestration, governance, and preparation of datasets for reporting and AI/ML use cases.
Key Responsibilities
- Design, build, and maintain ETL/ELT pipelines to ingest data from enterprise systems into Microsoft Fabric.
- Architect and maintain the Fabric medallion Lakehouse structure, including bronze, silver, and gold layers, as the single source of truth for Baker Group.
- Implement best practices for the data infrastructure and environment, including Development/Test/Production environments and Git-based version control.
- Own pipeline orchestration, scheduling, and monitoring to support reliable, timely, and accurate data availability.
- Curate and maintain core datasets across employee, finance, project, service, and manufacturing domains.
- Establish and enforce data quality, validation, and reconciliation processes across all pipelines.
- Design and manage data models, schemas, and semantic layers that enable Data Analyst reporting and Data Scientist modeling.
- Define and maintain data ontologies and canonical business definitions to keep consistent meaning across systems and data consumers (for example, definitions of “project,” “employee,” or “cost code”).
- Prepare and structure data for AI and machine learning, including feature-ready datasets, retrieval-augmented generation (RAG) pipelines, and vector embedding storage.
- Manage Fabric capacity planning, workspace organization, and performance optimization.
- Implement data governance practices such as access controls, lineage tracking, and metadata management aligned with Baker Group data classification standards.
- Partner with business system owners (ERP, HRIS, MRP, etc.) to understand upstream structures and manage change impacts.
- Collaborate with Data Scientists to ensure pipeline outputs support analytical and machine learning use cases.
- Collaborate with Data Analysts to ensure data products support paginated reporting, dashboards, and self-service BI needs.
- Collaborate with Software Development and DevOps teams to ensure data products support application development needs.
- Coordinate with third-party consultants when needed to deliver data engineering projects and augment capacity for high-demand business needs.
- Develop and maintain documentation for pipelines, schemas, and integration logic.
- Troubleshoot and resolve pipeline failures, latency issues, and data quality incidents.
- Monitor and maintain data-specific infrastructure, including Fabric capacity, pipeline orchestration tools, and monitoring/alerting systems.
- Evaluate and recommend new data engineering tools, patterns, and best practices.
- Stay current with emerging trends in data engineering, cloud data platforms, and integration techniques.
Required Qualifications
- Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or another relevant quantitative field.
- Three to five years of experience in data engineering, ETL/ELT development, or a related field.
- Proficiency with SQL and database technologies for data extraction, transformation, and loading.
- Experience with Microsoft Fabric, Azure Data Factory, or similar cloud ETL/orchestration tools.
- Experience with medallion architecture and modern data warehousing patterns.
- Experience with a programming language such as Python, PySpark, or T-SQL for data transformation.
- Familiarity with data modeling techniques, including dimensional modeling and star schema.
- Understanding of data governance, data quality, and metadata management practices.
Preferred Qualifications
- Experience preparing data for AI/ML consumption (for example, vector embeddings and RAG architectures) is a plus.
- Business acumen and understanding of construction or related industries is a plus.
Technologies
- Microsoft Fabric
- ETL, ELT
- Azure Data Factory
- Git
- SQL
- Python, PySpark
- T-SQL
- Medallion architecture
- Dimensional modeling, star schema
- Retrieval-augmented generation (RAG)
- Vector embedding storage
Minimum Education and Experience Required
- Bachelor’s degree
- 3+ years of experience (data engineering, ETL/ELT development, or related work)
Certifications (Optional)
- No specific requirements; relevant certifications such as Microsoft Certified: Fabric Data Engineer Associate, Azure Data Engineer Associate, or similar cloud platform certifications are a plus.
Competencies and Work Style
- Strong analytical and troubleshooting skills to diagnose and resolve complex pipeline and data quality issues.
- Excellent time and project management skills to prioritize across multiple pipeline and infrastructure projects.
- Strong communication skills, including translating technical data structures for non-technical stakeholders.
- Team-oriented collaboration with Data Scientists, Data Analysts, and business system owners.
- Ability to focus on complex technical problems and work independently with minimal supervision.
- Meticulous attention to detail and a commitment to producing reliable, well-documented data infrastructure.
- Ability to work in a fast-paced environment and adapt to changing business priorities.
Location and Environment
Location: Ankeny, IA (onsite)
- Prolonged periods of sitting at a desk and working on a computer.
- Must be able to lift 10 pounds occasionally.
- May have occasional visits to a job site requiring periods of standing, walking, and/or climbing stairs.
Tools
Use a computer for 8 hours a day.