Senior Bioinformatics Data Engineer
Job Description
Kaztronix seeks a Senior Bioinformatics Data Engineer for an onsite role in Wilmington, Delaware, to drive the modernization of a biomarker data lake on AWS. This position focuses on building and operating production pipelines across oncology clinical studies, partnering with the lead engineer across the full tech stack, and enabling downstream AI and visualization workflows with a strong emphasis on reproducibility and reliability.
Responsibilities
- Design and maintain ingestion pipelines orchestrated by Dagster for genomics vendors such as Caris, Predicine, Tempus, Olink, and CellCarta, incorporating IO managers, Iceberg writers, and row‑level accounting.
- Advance dbt Silver‑to‑Gold transformations with real data test coverage, store‑failures patterns, staging/intermediate/mart models, and macro consolidation.
- Build clinical data ingestion paths for SDTM and ADaM, implement reconciliation logic, and route subject dimensions as required.
- Deliver platform infrastructure including FastAPI endpoints, CI/CD pipelines, containerized deployments, observability instrumentation, and performance tuning for Redshift.
- Extract and reconcile transformation rules from legacy R and PySpark code against new platform implementations to maintain fidelity.
- Identify repetitive processes and convert them into automated workflows, guardrails, or reusable tooling to improve efficiency and consistency.
- Participate in adversarial design and code reviews, pinpoint edge cases, and push back on suboptimal patterns.
- Collaborate with the lead engineer on design decisions and maintain delivery velocity through paired working sessions and pull request reviews.
- Ensure reproducibility standards are met across the codebase with CI on every PR, automated tests, and avoidance of ad hoc notebook‑based production processes.
Requirements
- AI‑native engineering practice: proven experience building systems and workflows around AI coding agents such as Claude Code, Cursor, Codex, or equivalent. This goes beyond using prompts alone; you know when to automate, implement guardrails, and construct infrastructure that speeds future work.
- Education: Bachelor's or master's degree in computer science, data engineering, bioinformatics, or a related field.
- Experience: at least 5 years of professional data engineering experience delivering production pipelines on AWS (S3, ECS/Fargate, Redshift or equivalent MPP).
- Strong Python and SQL skills with working knowledge of modern data engineering libraries.
- Advanced proficiency with dbt and a workflow orchestration tool such as Dagster, Airflow, or Prefect.
- Data quality instinct: track record of catching silent failures, questioning data correctness, and identifying lossy joins or incomplete deliveries.
- Solid understanding of lakehouse architectures, ETL processes, and schema design for complex multi‑modal datasets.
- Ability to handle PHI‑adjacent clinical data under contractor policy, including background checks, compliance training, and VPN access.
- Willingness to work with legacy codebases (R, PySpark) to extract business rules and validate new implementations.
- Excellent communication skills and experience working in an embedded pair model with tight feedback loops.
Technologies
- Dagster
- dbt
- Apache Iceberg
- Amazon S3
- Redshift
- FastAPI
- Docker
- Amazon ECS
- AWS Fargate
- Python
- SQL
- R
- PySpark
- Airflow
- Prefect
- CloudFormation
- AWS Glue Catalog
- CI/CD
Preferred Qualifications
- Direct experience with Apache Iceberg, AWS Glue Catalog, or lakehouse table formats.
- Comfort reading genomic data such as VAF, HGVS nomenclature, VCFs, CNV/fusion semantics, or the ability to ramp quickly on unfamiliar scientific domains.
- Familiarity with SDTM, ADaM, and CDISC clinical data standards.
- Pharma, clinical research, or life sciences background.
- Experience with containerization (Docker/ECS) and infrastructure as code (CloudFormation).
- Proficiency in R for interoperability with bioinformatics teams.