EngineerJobs.io
← Back to all jobs

Job Description

BioAgilytix is seeking a hands-on Data Engineer technical lead to design, build, and support its Enterprise Data Platform. This onsite role in Durham, NC focuses on delivering trusted, governed, production-ready data pipelines, dimensional and semantic models, and certified data products that support laboratory operations, analytics, sponsor reporting, regulatory compliance, and AI initiatives.

What you will help deliver

  • Scalable enterprise data platforms that enable high-quality, governed data products for analytics, scientific operations, sponsor reporting, regulatory compliance, and AI initiatives.
  • Production-grade ELT pipelines integrating data from laboratory information systems (LIMS), ERP, CRM, APIs, sponsor systems, cloud applications, and other enterprise sources.
  • Modular, reusable transformation frameworks using modern ELT practices, including automated testing, documentation, lineage, version control, and deployment automation.
  • Dimensional data models, semantic models, and governed datasets based on established enterprise business definitions.
  • Validated, traceable, and auditable pipelines supporting regulated laboratory operations and enterprise reporting with data integrity, reproducibility, lineage, and compliance with GxP, GLP, HIPAA, and 21 CFR Part 11, aligned with enterprise data governance standards.
  • Automated data validation and reconciliation, data quality controls, audit logging, monitoring, observability, and operational alerting to help ensure reliability.
  • Platform improvements across performance, scalability, security, governance, and operational efficiency.
  • Operational support including production monitoring, incident resolution, root cause analysis, and continuous reliability enhancements.
  • Sponsor-facing data delivery through harmonization, transformation, validation, lineage, and regulatory reporting processes.
  • Certified, governed, and AI-ready data products that support enterprise analytics, machine learning, semantic search, and generative AI initiatives.
  • Collaboration with Laboratory Operations, Quality teams, IT, and business stakeholders to deliver scalable, reusable, governed solutions.

Core requirements

  • Bachelor’s degree in computer science, Information Systems, Engineering, Mathematics, Data Science, or a related field (Master’s preferred).
  • 5+ years of experience in Data Engineering, Data Management, Software Engineering, Business Intelligence, or related technical disciplines, preferably within life sciences, biotechnology, pharmaceuticals, CROs, healthcare, or other regulated industries.
  • 3+ years hands-on experience designing, developing, and supporting enterprise-scale data engineering solutions in production environments.
  • 3+ years hands-on experience architecting, developing, and administering enterprise solutions using Snowflake, including performance optimization, security, governance, workload management, and operational support.
  • Strong hands-on experience with dbt Cloud or dbt Core for modular transformations, automated testing, documentation, lineage, and deployment.
  • Demonstrated expertise in enterprise dimensional data modeling (star schemas, conformed dimensions, slowly changing dimensions, snapshot fact tables, and analytical data warehouse design).
  • Experience designing semantic models, enterprise business vocabularies, ontology-driven data products, or knowledge graph concepts.
  • Strong proficiency in SQL and Python.
  • Experience with enterprise integration using ETL/ELT platforms such as Talend, Fivetran, or equivalent tools.
  • Experience integrating enterprise applications via REST APIs, GraphQL APIs, file-based interfaces, Change Data Capture (CDC), and event-driven messaging platforms.
  • Experience with AWS (S3, Lambda, ECS, Glue) and/or Azure.
  • Experience designing, validating, and maintaining certified enterprise data products with documented business definitions, transformation logic, lineage, ownership, and lifecycle management.
  • Experience implementing least-privilege security, RBAC, data masking, row-level security, encryption, secrets management, and secure data sharing.
  • Experience supporting enterprise production data platforms, including incident management, root cause analysis, operational monitoring, performance tuning, release management, and platform reliability engineering.
  • Experience working with Laboratory Information Management Systems, bioanalytical data, sponsor deliverables, and regulated laboratory environments is strongly preferred.
  • Expert proficiency in SQL and Python, including Snowflake architecture such as Snowpark, Dynamic Tables, Streams, Tasks, and data sharing/security/governance/workload management.
  • Strong DataOps and engineering practices including Git, GitHub Actions, CI/CD, Infrastructure-as-Code, automated testing, and observability.
  • Excellent communication, documentation, and collaboration skills, with the ability to work independently on complex assignments and escalate decisions when needed.
  • Ability to navigate a fast-paced, evolving data landscape, demonstrating resilience and flexibility.

Technologies you will work with

  • Snowflake, Snowpark, Dynamic Tables, Streams, Tasks, data sharing
  • dbt Cloud, dbt Core (models, snapshots, macros, tests, semantic models, documentation, lineage)
  • ELT, dimensional modeling, semantic modeling
  • Talend, Fivetran, REST APIs, GraphQL APIs, CDC, event-driven messaging
  • AWS (S3, Lambda, ECS, Glue) and/or Azure
  • RBAC, Power BI, Sigma, Tableau, semantic reporting platforms
  • Git, GitHub Actions, Infrastructure-as-Code, DataOps

Benefits

  • Medical Insurance (HDHP with HSA; PPO)
  • Dental Insurance
  • Vision Insurance
  • Flexible Spending Account (medical; dependent care)
  • Short Term Disability | Long Term Disability
  • Life Insurance
  • Paid Time Off (4 weeks per year)
  • Parental Leave
  • Paid Holidays (9 scheduled; 5 floating)
  • 401k with Employer Match
  • Employee Referral Program

Additional position details

  • Full-time role; some flexibility in hours with availability during core work hours per the BioAgilytix Employee Handbook.
  • Occasional weekend, holiday, and evening work needed.
  • No supervisory responsibilities.
  • Reports to the Data Management Lead; works independently on complex engineering initiatives with infrequent supervision and instructions.
  • Preferred credentials: Master’s degree.

Physical demands

  • Ability to work in an upright and/or stationary position for up to eight (8) hours per day.
  • Occasional mobility needed; occasional crouching and stooping with frequent bending and twisting of upper body and neck.
  • Light to moderate lifting and carrying up to 20 pounds, including luggage and a laptop computer.
  • Ability to use office equipment and computer software; ability to communicate information and ideas effectively.
  • Regular and consistent attendance; ability to perform under stress and multi-task.

Similar Jobs