EngineerJobs.io
← Back to all jobs

Job Description

firstPROPhiladelphia offers a hybrid contract-to-hire Senior Data Engineer role in Philadelphia. You will focus on real-time streaming data pipelines for connected medical devices, leveraging Azure Databricks, Azure services, and ML/AI data workflows. The position is commissioned at USD 60 to 70 per hour, with a path toward permanent placement through a contract-to-hire arrangement and opportunities to collaborate across data, software, and analytics teams.

Responsibilities

  • Design, build, and operate batch and streaming data pipelines on Azure and Databricks, with a primary focus on real-time IoT and telemetry within a Medallion architecture.
  • Develop ETL and ELT workflows to ingest, transform, and validate large volumes of structured and unstructured data.
  • Build and maintain data services, APIs, and microservices for application, analytics, and ML/AI teams.
  • Implement real-time streaming solutions using Azure Event Hubs, Azure Stream Analytics, and related Azure integration patterns, prioritizing cost-effective throughput, thoughtful partitioning, and reliable downstream delivery to Databricks.
  • Optimize production Databricks pipelines with PySpark, Spark SQL, and Delta Lake, including performance tuning, reliability improvements, and cost awareness.
  • Troubleshoot and resolve complex pipeline issues across Databricks, Azure, and on premise systems, conducting root cause analysis and corrective actions.
  • Collaborate with data analysts, software engineers, ML engineers, and business stakeholders to translate requirements into technical designs and delivery priorities.
  • Apply data quality, validation, and privacy-first practices, delivering reliable pipelines through established software engineering standards, documentation, testing, and CI/CD.

Requirements

  • Bachelor's degree in Computer Science, Mathematics, Engineering, or a related technical field plus 6+ years of professional data engineering, analytics, or warehousing experience; or Master’s degree with 3+ years of relevant experience.
  • 5+ years designing, building, and operating big data and real-time streaming pipelines across cloud and on-prem environments.
  • 5+ years applying DevOps and CI/CD practices to data and analytics workloads.
  • Production experience building data services, APIs, or microservices for downstream data consumption.
  • Hands-on expertise designing and delivering production-grade data pipelines on Databricks.
  • Proven experience with Spark optimization and tuning in Databricks, including performance analysis, partitioning, caching, shuffle optimization, and cost-aware design.
  • Strong experience with Azure cloud services for data engineering and streaming workloads.
  • Experience ensuring data quality, validation, and privacy-aware handling for regulated or sensitive data.

Technologies

  • Databricks, Spark, PySpark, Spark SQL, Delta Lake
  • Azure, Azure Event Hubs, Azure Stream Analytics
  • Azure Databricks, Medallion architecture

Outcomes

  • Onboard to the Azure Databricks environment and contribute to troubleshooting, stabilization, and optimization of existing batch and streaming pipelines.
  • Stand up Azure streaming ingestion for telemetry data and deliver production-ready pipelines integrated with Databricks to support API, ML, and downstream analytics within the Medallion architecture.
  • Design and deliver data services or consumption patterns that enable business and ML teams to access near real-time telemetry data reliably, securely, and at scale.

Work hours and travel requirements

The role requires openness to assisting in off-hours production troubleshooting as needed. The IT team operates in a hybrid environment, with a minimum of two days per week in the downtown Philadelphia office.

Similar Jobs