Senior Specialist: Data Engineer for Upstream Biologics
Job Description
Location: Boston, MA onsite with hybrid options. Merck offers a collaborative, inclusive culture and a comprehensive benefits package for a Senior Specialist, Data Engineer focused on Upstream Biologics. This role centers on designing and maintaining data pipelines to ingest Upstream Biologics Process Development data, enabling downstream analytics, process characterization, scale-up predictions, and multivariate analyses.
Benefits
- Medical, dental, and vision coverage for employees and their families
- Retirement benefits, including a 401(k) plan
- Paid holidays
- Vacation time
- Compassionate and sick days
Responsibilities
- Build and maintain robust, scalable data pipelines that ingest experimental and process data from upstream biologics source systems
- Deliver analysis-ready datasets to support upstream digital initiatives, including process characterization models, scale-up predictions, multivariate analytics, and high throughput process development workflows
- Map instrument outputs and experimental results ensuring ontology alignment and interoperability across upstream data sources
- Develop and maintain data visualizations, dashboards, and reports that enable upstream scientists to explore process data across runs, molecules, and scales
- Support system of record standards by ensuring consistent data entry practices
- Identify and flag data quality issues, gaps in metadata, and inconsistencies across source systems, contributing to continuous improvement of upstream data capture practices
- Collaborate with upstream process development scientists, analytical scientists, and engineers to understand evolving data needs and translate them into pipeline requirements
- Coordinate with adjacent domain engineers to ensure seamless data handoffs at domain boundaries
- Maintain and version all pipeline code in GitHub, following team standards for code review, documentation, and deployment
- Demonstrate excellent interpersonal, communication, and collaboration skills
- Embrace and model our core values of inclusion, including fostering a supportive culture where all can thrive
- Collaborate effectively in a dynamic, integrated, and multidisciplinary team environment
Requirements
- PhD required; minimum 2 years of experience
- Proficient in Python and/or R programming
- Comfortable in development environments such as Jupyter, Posit/RStudio, or VS Code
- Solid SQL skills with hands-on experience writing and optimizing queries against relational databases and data warehouses
- Experience with ETL/ELT processes and building data pipelines in a scientific or pharmaceutical context
- Familiarity with version control systems (Git/GitHub) and collaborative software development practices
- Ability to work in a team environment with cross-functional interactions
- Motivated to learn new skills, willing to take on new challenges, and scientifically curious
Technologies
- Python, R
- Jupyter, Posit/RStudio, VS Code
- SQL, Git, GitHub
- Databricks, Delta Lake
- Streamlit, Shiny
- PowerBI, Spotfire, Tableau
- Allotrope Simple Model, ISA-88, OPC-UA
Posting details
- Requisition ID: R404620
- End date: 07/9/2026
- Location: Boston, MA (onsite) with Hybrid arrangements
- Employee status: Regular
Inclusion, accommodations, and fair hiring
Merck supports inclusive hiring and will consider qualified applicants including those with arrest or conviction records in compliance with applicable laws. For accommodation during the application or hiring process, please request assistance. We observe fair chance hiring practices in accordance with local ordinances where applicable.
Additional notes
Posting is active through the stated end date above. This role emphasizes collaboration across upstream process development, analytical science, and engineering teams to meet evolving data needs.