Associate Director, Data Engineer: DSCS Digital Data Strategy
Senior
Analytics
Associate Director
Big Data
Cloud Operations
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Governance
Data Integration
Data Lake
Data Management
Data Platform
Data Processing
Data Quality
Data Strategy
Data Visualization
Data Warehouse
Database
Databricks
Metadata Management
Ontology Mapping
Reporting and Analytics
SQL
Job Description
This on-site opportunity in Boston centers on biologics data engineering within the DSCS Digital Data Strategy program. The role oversees the design, construction, and governance of biologics data pipelines to enable modeling, optimization, and decision support across Digital Insights. A Ph.D. and at least three years of relevant experience are required, with a salary range of USD 129,000 to 203,100 per year.
Responsibilities
- Act as the domain owner for biologics data engineering, maintaining visibility across all digital projects, data sources, systems, and data flows within the domain.
- Inform and influence the ongoing development of the DSCS Digital Data Strategy.
- Design and implement robust, scalable data pipelines that ingest experimental and process data from biologics source systems such as process historians, chromatography systems, electronic lab notebooks, and analytical instruments.
- Deliver analysis-ready datasets to support digital initiatives, including process characterization models, data lineage tracking, multivariate analytics, and cross-site manufacturing connectivity.
- Define and enforce data standards, metadata schemas, and ontology mappings to ensure biologics data is interoperable and readily consumable by modeling and optimization workflows.
- Align proactively with automation colleagues to anticipate new data streams from automated workflows that require pipeline development and ontology mapping.
- Own and govern system of record standards for biologics, ensuring consistent configuration and data entry practices across experiments, molecules, and sites.
- Catalog processes, analytical methods, instruments, and digital systems within the biologics domain to create a comprehensive data landscape map.
- Develop and maintain data visualizations, dashboards, and reports that enable scientists to explore process data across runs, molecules, scales, and manufacturing sites.
- Influence the biologics digital data strategy by identifying opportunities to improve data capture at the source and reduce friction between experimentation and modeling.
- Mentor and guide supporting data engineers across modalities, ensuring alignment with domain strategy and ontology governance.
- Maintain and version all pipeline code in GitHub, following team standards for code review, documentation, and deployment.
- Build strong partnerships with process development scientists, analytical scientists, and manufacturing teams to gather requirements and shape the domain’s digital data strategy.
- Demonstrate excellent interpersonal, communication, and collaboration skills.
- Embrace and model core values of diversity and inclusion, fostering a supportive culture where all can thrive.
- Collaborate effectively in a dynamic, integrated, multidisciplinary team environment.
Requirements
- Proficiency in Python and/or R, with comfort in development environments such as Jupyter, Posit/RStudio, or VS Code.
- Solid SQL skills with hands-on experience writing and optimizing queries against relational databases and data warehouses.
- Experience with ETL/ELT processes and building data pipelines in a scientific or pharmaceutical context.
- Familiarity with cloud platforms (AWS, Azure, or GCP) for data storage, processing, and integration.
- Working knowledge of how process models, multivariate analyses, and statistical tools depend on experimental data to anticipate modeler needs and deliver properly structured datasets.
- Experience defining or enforcing data standards, metadata schemas, or ontology mappings in a scientific or pharmaceutical context.
- Familiarity with version control systems (Git/GitHub) and collaborative software development practices.
- Proven ability to lead technical initiatives, mentor junior engineers, and influence data strategy across multiple stakeholders.
- Ability to deliver complex solutions under compressed timelines in a dynamic environment.
Technologies
- Python
- R
- Jupyter
- Posit/RStudio
- VS Code
- SQL
- AWS
- Azure
- GCP
- Databricks
- Delta Lake
- Git
- GitHub
- Streamlit
- Shiny
- PowerBI
- Spotfire
- Tableau
- Allotrope Simple Model
- ISA-88
- OPC-UA
Benefits
- Medical, dental, vision health coverage and other insurance benefits for employee and family
- Retirement benefits, including 401(k)
- Paid holidays
- Vacation
- Compassionate and sick days
- Annual bonus and long-term incentive, if applicable
Preferred Experience and Skills
- Hands-on experience in biologics process development such as chromatography, filtration, purification, or formulation, with a transition into data engineering, data science, or computational roles
- Experience with data pipeline and analytics platforms such as Databricks, including notebook development, workflow orchestration, and Delta Lake
- Experience with data visualization tools (Streamlit, Shiny, PowerBI, Spotfire, or Tableau) for scientist-facing dashboards
- Familiarity with ontology frameworks or standardized data models (e.g., Allotrope Simple Model, ISA-88, OPC-UA) and mapping instrument data to structured schemas
- Understanding of DoE methodologies and process characterization study designs for structuring data used in CPP and CQAs analysis
- Experience with data lineage concepts and building traceability across experimental systems, materials, and manufacturing steps
- Knowledge of regulatory expectations relevant to biologics process development and validation (ICH Q8-Q12, validation lifecycle, comparability studies)
- Experience partnering with lab automation teams to integrate data from automated workflows
- Experience with cross-site data integration and harmonizing data from multiple facilities
- History of cross-functional collaborations spanning laboratory, manufacturing, modeling, and digital teams