Senior Data Engineer
Job Description
Xenon7 is partnering with a life sciences client in the Indianapolis, IN area to build scalable data platforms that connect scientific and clinical informatics with manufacturing process and OT data. This is a contract, full-time enterprise project engagement delivered in a hybrid setup with 3 days onsite per week, with an emphasis on hands-on engineering, architecture ownership, and governance in a highly monitored environment.
The Senior Data Engineer role is responsible for the architectural vision and day-to-day delivery of production-grade pipelines, structured datasets, and enterprise reporting foundations that support downstream machine learning enablement without requiring model training.
Responsibilities
- Design, build, and maintain production-grade data pipelines and data architecture for scientific, clinical trial, and research informatics across small and large molecules, genomics, proteomics, and LIMS.
- Structure complex multi-modal clinical and scientific datasets to support advanced analytics, enterprise reporting, and downstream machine learning models.
- Ingest, harmonize, and model operational technology (OT) and manufacturing process datasets, including API manufacturing pipelines, batch processing data, MES, SCADA, and OSIsoft PI systems.
- Unify laboratory and facility data pipelines into centralized, highly available enterprise data platforms.
- Build robust ETL/ELT pipelines using modern cloud platforms including Databricks, Snowflake, and AWS/Azure, with PySpark and SQL.
- Ensure pipeline integrations meet enterprise governance, data residency, and GxP regulatory standards in a heavily monitored environment.
- Collaborate directly with process engineers, chemical engineering leads, and research informatics directors to translate operational friction into clear technical specifications.
- Establish engineering best practices, data modeling standards, and pipeline monitoring frameworks across the enterprise data stack.
Requirements
- Senior-level proficiency with 10–20+ years across software development, data platform architecture, and complex ETL/ELT engineering.
- Proven ability to engineer pipelines across non-standard, specialized domains, including transitions between process/chemical engineering data and clinical/scientific research informatics.
- Unrestricted US Work Authorization with no sponsorship available, and ability to work 3 days onsite per week in the Indianapolis, IN area.
- Strong problem-solving, adaptability, and capability to articulate complex architecture to cross-functional engineering teams.
- Advanced Python, PySpark, and expert-level SQL.
- Hands-on expertise with Databricks, Snowflake, or AWS/Azure enterprise data ecosystems.
- Extensive experience with Airflow, dbt, Spark, and enterprise data orchestration engines.
- Experience building streaming and batch architectures via REST APIs, message brokers, and database integrations.
- Deep exposure to either scientific/clinical informatics (CDISC/SDTM, LIMS, clinical trials, multi-omics) or chemical/process engineering data (API manufacturing, batch data, SCADA, MES, OSIsoft PI).
Technology Stack
Python, PySpark, SQL, Databricks, Snowflake, AWS, Azure, Airflow, dbt, Spark, REST APIs, message brokers, CDISC/SDTM, LIMS, MES, SCADA, OSIsoft PI
Location & Contract Details
- Location: Indianapolis, IN Metro (Hybrid / 3-Day Onsite). Open to regional/EST candidates with onsite travel.
- Contract type: Contractor Full-Time / enterprise project engagement (outsourced via Xenon7).
Nice-to-Haves & Certifications
- Academic background in Chemical Engineering, Bio-process Engineering, Computer Science, or related STEM discipline.
- Direct experience working inside regulated GxP environments in Life Sciences or Specialty Chemicals.
- Certifications including Databricks Certified Data Engineer Senior/Professional, Snowflake SnowPro Core/Advanced, or AWS Data Engineer Associate/Professional.
Role Focus (What This Position Is Not)
- Not a pure Data Scientist or ML Researcher: the role designs and scales data architecture, pipelines, and platform infrastructure rather than building or training ML models.
- Not a non-coding architect: this is a 100% hands-on engineering lead position requiring direct pipeline construction and technical execution.
- Not fully remote: the role requires consistent hybrid commitment of 3 days onsite per week at the client site in Indianapolis.