EngineerJobs.io
← Back to all jobs

Job Description

UnitedHealth Group seeks a seasoned Principal AI / Machine Learning Data Engineer to shape how the organization leverages large-scale unstructured data for advanced analytics and Generative AI. This role sits at the intersection of data engineering and AI, delivering production-grade data platforms and scalable AI capabilities from ingestion to deployment. Based in Eden Prairie, MN with remote options, the position asks for hands-on leadership in building data-driven solutions that power enterprise analytics and AI initiatives.

Responsibilities

  • Design, develop, and maintain scalable data pipelines and data platforms supporting analytics, machine learning, and AI use cases
  • Build and optimize ingestion frameworks for large-scale structured and unstructured data, including streaming and event-driven sources
  • Partner with cross-functional stakeholders to understand evolving data and AI needs and define long-term technical solutions
  • Enable and support machine learning and AI workflows, including feature engineering, data preparation, and model deployment support
  • Drive strategic initiatives around Generative AI, data quality, observability, lineage, and governance
  • Develop and maintain frameworks that support rapid experimentation and deployment of AI/ML solutions
  • Introduce and evolve best practices in data modeling, orchestration, testing, and monitoring
  • Identify and champion opportunities for platform scalability, performance optimization, and cost efficiency
  • Collaborate with product, analytics, and infrastructure teams to deliver high-impact data and AI solutions
  • Build and maintain reusable parsing, enrichment, analytic, and service libraries to accelerate delivery across teams
  • Work under time-sensitive conditions while ensuring thoroughness
  • Maintain high ethical standards and confidentiality
  • Build and operate production data platforms and pipelines across batch and streaming workloads
  • Engage in hands-on engineering in Python and SQL; work with JVM languages (Java/Scala) within Spark ecosystems
  • Implement distributed processing and lakehouse/warehouse patterns (e.g., Spark/PySpark, Databricks, Snowflake)
  • Build pipelines for OCR, document parsing, and text extraction from image-based or scanned data sources
  • Enable Generative AI solutions in production (eg, retrieval-augmented generation), including retrieval patterns and evaluation/monitoring practices
  • Adopt knowledge-centric data approaches (metadata-driven systems, entity resolution, graph concepts) to improve discoverability
  • Emphasize data quality, observability, and monitoring through profiling, validation, alerting, and reliability improvements
  • Orchestrate, implement CI/CD, containerization, and infrastructure-as-code (Airflow, GitHub Actions, Docker, Terraform, Kubernetes)
  • Work in cloud environments (AWS, Azure, and/or GCP), securely handling sensitive data (PII/PHI) with compliance partners
  • Lead through influence, mentor engineers, and translate ambiguous problems into scalable technical roadmaps

Requirements

  • Bachelor’s degree or equivalent experience
  • 5+ years designing, building, and operating scalable data pipelines and platforms (batch + streaming)
  • 2+ years deploying Generative AI solutions to production (e.g., RAG, LLM-powered pipelines, semantic search)
  • Strong hands-on development in Python and SQL, with experience in Spark/PySpark and Databricks (or similar distributed platforms)
  • Experience building ingestion and processing frameworks for unstructured data (OCR, documents, images), including parsing and enrichment
  • Experience with cloud platforms (AWS/Azure/GCP), DevOps/CI/CD, and infrastructure-as-code, including secure handling of sensitive data (PII/PHI)
  • Proven ability to design scalable solutions, implement data quality/observability practices, and collaborate across stakeholders

Technologies

  • Python
  • SQL
  • Spark
  • PySpark
  • Java
  • Scala
  • Databricks
  • Snowflake
  • AWS
  • Azure
  • Google Cloud Platform (GCP)
  • Kafka
  • Kinesis
  • Event Hubs
  • Airflow
  • GitHub Actions
  • Docker
  • Terraform
  • Kubernetes
  • Delta Lake
  • Plotly
  • Seaborn
  • Chartjs
  • Great Expectations
  • Deequ

Benefits

  • comprehensive benefits package
  • incentive and recognition programs
  • equity stock purchase
  • 401k contribution

Preferred Qualifications

  • Experience with cloud platforms such as AWS, Azure, or Google Cloud, including managed data services
  • Experience with streaming and event-driven architectures (Kafka, Kinesis, Event Hubs)
  • Experience with data quality and validation frameworks (Great Expectations, Deequ) and/or data observability tooling
  • Experience enabling MLOps practices (feature stores, model registries, experiment tracking, deployment automation)
  • Experience with lakehouse architectures, Delta Lake, and advanced Spark optimization/performance tuning
  • Experience with data visualization tools and libraries such as Plotly, seaborn, and Chartjs
  • Experience with machine learning and predictive analytics
  • Familiarity with security and privacy concepts for data platforms (least privilege, PII/PHI handling) and working with compliance partners
  • Solid hands-on engineering in Python and SQL; familiarity with JVM languages (Java/Scala) in Spark ecosystems

Application Deadline

This posting will remain open for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.

Similar Jobs