IT Data Engineer β BB4352
Big Data
Bigdata
CI/CD
Cloud
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Architecture
Data Engineer
Data Governance
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Operations
Data Pipeline
Data Platform
Data Processing
Data Security
Data Warehouse
Database
Databases
Databricks
Databricks Workflows
Delta Lake
DevOps
DevSecOps
Engineering
ETL
Informatica
Information Technology (IT)
Infrastructure As Code
Kubernetes
Platform Engineering
Pyspark
Software Development
Spark
SQL
Job Description
IT Data Engineer will support Discovery R&D and data science efforts with Databricks-based, scalable data platforms and automated scientific pipelines.
Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies
- Build and optimize ETL/ELT for structured, semi-structured, and unstructured datasets
- Create data ingestion frameworks for research, laboratory, clinical, and external scientific data sources
- Implement data quality, validation, monitoring, and governance processes
- Support enterprise data lakehouse architecture and data platform modernization initiatives
- Develop and maintain Databricks notebooks, workflows, Delta Live Tables (DLT), and Databricks Jobs
- Create optimized Spark transformations and data processing solutions
- Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management
- Manage Delta Lake environments and optimize performance, scalability, and cost
- Integrate Databricks with cloud-native services and enterprise applications
- Deploy, configure, and support Seqera Platform (formerly Nextflow Tower)
- Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads
- Integrate Seqera workflows with AWS cloud infrastructure and compute environments
- Support containerized workflows using Docker and Kubernetes
- Enable reproducible, scalable, and compliant scientific data processing workflows
- Design and implement AWS-based cloud data solutions
- Manage cloud storage including S3 and associated data lifecycle policies
- Develop Infrastructure-as-Code using Terraform or CloudFormation
- Implement security controls and access management aligned with enterprise IT standards
- Collaborate with data scientists, researchers, bioinformaticians, and business stakeholders on data requirements
- Provide technical guidance on data engineering best practices and workflow automation
- Troubleshoot pipeline failures, performance issues, and workflow bottlenecks
- Contribute to platform roadmaps and continuous improvement initiatives
- Maintain technical documentation, SOPs, and knowledge articles
Requirements
- Bachelor's degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field
- 5+ years of experience in data engineering, cloud engineering, or analytics platform development
- 3+ years hands-on experience with Databricks and Apache Spark
- 2+ years experience with Seqera Platform (Nextflow Tower) and Nextflow workflows
- Experience supporting scientific research and life sciences environments is preferred (life sciences, pharmaceutical, biotech, healthcare)
- Databricks & Data Engineering, Databricks Lakehouse Platform
- Apache Spark (PySpark, Spark SQL)
- Delta Lake and Delta Live Tables (DLT)
- Databricks Workflows and Unity Catalog
- SQL and Python
- Seqera & Scientific Computing, including Seqera Platform / Nextflow Tower
- Nextflow pipeline development and bioinformatics workflow automation
- Docker and container technologies
- Kubernetes orchestration
- High-performance computing environments
- AWS required: S3, IAM, EC2, VPC, Lambda
- Terraform or CloudFormation
- Cloud monitoring and logging tools
- Data lake and lakehouse architectures, ETL/ELT frameworks, data modeling, and data cataloging/governance
- API integrations and data quality frameworks
- GitHub/GitLab
- CI/CD pipelines with Jenkins, GitHub Actions, or similar tools
- Infrastructure as Code and Agile/DevOps methodologies
Technologies
- Databricks
- Seqera Platform (Nextflow Tower)
- Nextflow
- Apache Spark, PySpark, Spark SQL
- Delta Lake, Delta Live Tables (DLT)
- Databricks Workflows
- Unity Catalog
- Medallion Architecture (Bronze, Silver, Gold)
- AWS, S3, IAM, EC2, VPC, Lambda
- Terraform, CloudFormation
- Docker, Kubernetes
- SQL, Python
- GitHub, GitLab
- CI/CD, Jenkins, GitHub Actions
- Infrastructure as Code
Benefits
- Dental insurance
- Health insurance
Location and Employment
- Location: Cambridge, MA (onsite)
- Job type: Contract
Compensation
- Pay: USD 70 - 95 per hour
- W2 rate question: What hourly rate are you seeking on a W2?
Application Questions
- Where are you currently located?
- What hourly rate are you seeking on a W2?