Data Engineer with ETL/ELT pipelines and API integrations
Job Description
Dhanu Global Enterprises, Inc. is building a governed data and AI platform for a pharmaceutical environment where integration, traceability, and compliance matter. In this onsite role in Indianapolis, you will design and deliver production-ready data engineering solutions, spanning structured ETL/ELT and API-driven ingestion that supports regulatory reporting and AI-enabled scientific insights.
This position focuses on creating an Azure-based foundation that brings together device, laboratory, partner, and document data into a unified, governed digital thread. Delivery includes building pipelines end to end, implementing quality and lineage controls, and maintaining reliable operations in regulated GxP/FDA contexts.
Responsibilities
- Design, build, and maintain production ETL/ELT pipelines that integrate laboratory and operational systems, including Darwin, Teamcenter/PLM, LabVantage LIMS, Jama, TurboAC, and Qdocs/Veeva, into Microsoft Azure Fabric Lakehouse and PostgreSQL.
- Develop and optimize Bronze, Silver, and Gold medallion architecture, including schema mapping, data modeling, referential integrity, and performance optimization.
- Build scalable API integrations and cross-cloud data pipelines across Azure and AWS to support enterprise data integration.
- Implement automated data quality controls, controlled vocabulary normalization, schema validation, Q-gate/specification checks, data lineage, and audit trails to support GxP compliance.
- Monitor, troubleshoot, and optimize pipeline performance, reliability, error handling, and operational monitoring in production environments.
- Collaborate with business, engineering, and IT teams to integrate data sources and establish a governed, scalable digital thread supporting analytics, AI, and regulatory reporting.
Requirements
- 8+ years of hands-on Data Engineering experience building production ETL/ELT pipelines and API integrations.
- Expert-level Python and/or PySpark for data ingestion, transformation, orchestration, testing, and CI/CD.
- Experience with Microsoft Azure Fabric (Lakehouse, Data Factory, Fabric Pipelines, Delta Lake) delivering end-to-end production solutions.
- Hands-on AWS experience, including S3, Glue (or equivalent), and RDS/Aurora.
- Strong SQL and PostgreSQL experience, including normalized schema design, query optimization, indexing, and performance tuning.
- Experience with data modeling, medallion architecture, and Lakehouse design patterns.
- Experience implementing data lineage, quality controls, schema validation, error handling, and monitoring in regulated environments.
- Knowledge of GxP, GMP, GCP within pharmaceutical, biotechnology, or medical device environments.
- SCM, WMS, PLM, MM, QA, and/or any validated applications with FDA-regulated systems.
Location, Term, and Pay
- Location: Indianapolis, IN (onsite, 5 days per week)
- Duration: Through December 2026, with possible extension into 2027
- Compensation: USD 60 - 70 per monthly
- Minimum Experience: 8 years
Technology Focus
- ETL, ELT, Python, PySpark
- Microsoft Azure Fabric, Lakehouse, Data Factory, Fabric Pipelines, Delta Lake
- PostgreSQL, SQL
- AWS, S3, Glue, RDS, Aurora
- Bronze, Silver, Gold, Medallion architecture, Q-gate, Data lineage