Protocol Data Engineer
Job Description
A P Ventures LLC is supporting an FDA proof of concept for AI-assisted clinical trial protocol review. This remote role combines hands-on database engineering with structured protocol data modeling to help turn approved standards and technical models into scalable data structures, retrieval workflows, and traceable outputs within FDA-approved environments.
You will work alongside clinical protocol, data standards, architecture, interoperability, and AI/prompt engineering SMEs to operationalize structured protocol data models for both ICH M11-aligned and historical protocols, enabling reviewer-facing use supported by rigorous data quality practices.
Responsibilities
- Implement protocol data structures and physical schemas for extraction, normalization, storage, retrieval, comparison, and reviewer-facing use of ICH M11-aligned and historical protocols.
- Translate approved models including ICH M11, CDISC USDM, ontology, metadata, and terminology into practical database structures and data-engineering components for the proof of concept.
- Develop and support ETL/ELT pipelines and data transformations for protocol content and structured outputs, preserving source provenance, metadata, assumptions, validation status, and reviewer feedback.
- Collaborate with AI/prompt engineering and standards SMEs on RAG and semantic retrieval capabilities, including indexing, evidence caching, and protocol comparison components.
- Support integration with FDA-furnished Elsa/HALO capabilities and FDA-approved data services, including PostgreSQL and/or HALO/Databricks, for evidence caching, retrieval, analytics, and dashboard-ready outputs.
- Build advanced SQL and data-access logic for protocol extraction, historical retrieval and comparison, reviewer dashboards, and traceability to source evidence.
- Perform data-focused quality assurance, reconciliation, validation, and regression testing, including checks for completeness, mapping consistency, retrieval quality, unsupported values, ambiguity, and reproducibility.
- Document data models, interfaces, transformations, mappings, implementation decisions, technical limitations, and operational considerations to support FDA review, knowledge transfer, and future expansion.
- Participate in FDA technical working sessions, proof of concept demonstrations, validation activities, and iterative refinement based on reviewer feedback.
- Support implementation within FDA-provided platforms and approved data environments. Base effort is the proof of concept, without requiring a separate production platform.
Requirements
- Extensive experience in database engineering, data architecture, information modeling, and analytics in complex enterprise environments.
- Strong expertise in relational and dimensional data modeling, metadata management, ETL/ELT design, data warehousing, and data quality.
- Advanced SQL skills, including PostgreSQL/PLpgSQL and/or Oracle PL/SQL; SQL Server experience is beneficial.
- Experience with AWS data services and cloud database platforms such as Aurora PostgreSQL, Redshift, S3, Glue, Lambda, and DMS.
- Experience integrating JSON/XML data and RESTful APIs; working knowledge of Python and CI/CD practices is beneficial.
- Federal health, clinical research, or regulated-environment experience preferred; familiarity with FISMA/NIST controls is beneficial.
- Must be able to obtain a High-Risk Public Trust.
Technologies
- PostgreSQL, PLpgSQL, Oracle PL/SQL, SQL Server
- AWS, Aurora PostgreSQL, Redshift, S3, Glue, Lambda, DMS
- JSON, XML, RESTful APIs, Python, CI/CD
- Elsa/HALO, HALO/Databricks, ETL/ELT
- RAG, FISMA, NIST
- ICH M11, CDISC USDM
Education
BS degree in Computer Science, Mathematics, Data Science, or a relevant technical field.