EngineerJobs.io
← Back to all jobs

Job Description

INSPYR Solutions seeks a Lead Data Engineer to steer the design and optimization of distributed data pipelines within a high-volume, mission-critical setting. The role focuses on building and refining a Lakehouse architecture with Delta Lake, enabling robust data governance, reproducible ML workflows, and dependable production systems, all within a contract framework based in Doraville, GA with remote options. The position offers an hourly rate of USD 60 to 70.

Responsibilities

  • Design, develop, and tune distributed data pipelines with Apache Spark to support high-demand, mission-critical workloads.
  • Create and maintain an enterprise Lakehouse using Delta Lake, ensuring ACID compliance, lineage, auditability, and governance.
  • Build automated ingestion frameworks across batch, streaming, and event-driven modes spanning multiple cloud services and integration points.
  • Prepare feature-ready datasets and establish reproducible ML deployment patterns to enable machine learning workflows.
  • Lead data quality initiatives, access control, and cataloging across the platform.
  • Apply cost optimization, cluster tuning, and performance engineering to maximize efficiency.
  • Collaborate with Finance, BI, Operations, and ML teams to translate complex business needs into scalable data solutions.
  • Own production reliability, troubleshoot issues, and perform root cause analysis for data and ML pipelines.

Requirements

  • 7+ years of advanced data engineering experience with distributed compute technologies.
  • Expert Spark engineering skills, including performance tuning, cluster configuration, partition strategies, and handling large datasets.
  • Hands-on experience with Lakehouse architectures featuring ACID transactions, schema evolution, and governance frameworks.
  • Strong proficiency in Python and SQL for large-scale data transformations.
  • Experience supporting machine learning pipelines or model operationalization.
  • Proven ability to architect cloud-native data platforms on Azure, AWS, or GCP.
  • Ability to integrate diverse, complex data sources at enterprise scale.
  • Demonstrated track record of owning mission-critical production systems.
  • Experience with distributed streaming frameworks such as Kafka or Event Hubs.
  • Experience building or supporting ML platforms, feature stores, or experiment tracking systems.
  • Background in data security, compliance controls, or audit-ready governance.
  • Experience automating data operations with CI/CD and infrastructure as code.

Technologies

  • Databricks, Python, SQL, Apache Spark, Delta Lake, MLflow, Notebooks
  • Hugging Face Transformers, LangChain, LlamaIndex, LLMs including Anthropic Claude, Meta LLaMA, Google Gemini
  • Kafka, Event Hubs, ADLS, S3, GCS
  • Git, GitHub, GitLab, Azure Repos, Databricks Repos, GitHub Actions, Azure DevOps
  • Unity Catalog, RBAC, REST APIs

Core Tools

  • Databricks (Spark, Delta Lake, MLflow, Notebooks)
  • Python & SQL
  • Apache Spark via Databricks
  • Delta Lake for lakehouse architecture

Cloud Platforms

  • Azure, AWS, or GCP
  • Cloud Storage: ADLS, S3, GCS

Data Integration

  • Kafka or Event Hubs (streaming)
  • Auto Loader (Databricks file ingestion)
  • REST APIs, AI/ML workflows, MLflow
  • Hugging Face Transformers, LangChain / LlamaIndex for LLM integration
  • LLMs: Anthropic Claude, Meta LLaMA, Google Gemini

DevOps

  • Git ecosystems: GitHub, GitLab, Azure Repos
  • Databricks Repos
  • CI/CD: GitHub Actions, Azure DevOps

Security & Governance

  • Unity Catalog
  • RBAC

Benefits

  • Comprehensive medical benefits
  • Competitive pay
  • 401(k) retirement plan

Similar Jobs