EngineerJobs.io
← Back to all jobs

Job Description

Join CNA Insurance in Chicago as a Senior Data AI Engineer. This onsite role focuses on designing, building, and operating end-to-end AI and ML solutions that migrate CNA to a modern cloud data lakehouse. You will work across structured and unstructured data, unlocking analytical value through scalable pipelines, retrieval augmented generation architectures, vector databases, and knowledge graphs. As a senior individual contributor, you will influence data strategies and may guide others, contributing to AI readiness at scale.

Location: Chicago, IL (onsite). Salary range: USD 72,000 - 141,000 per year.

Responsibilities

  • Design and implement AI capabilities that accelerate legacy-to-cloud migrations, prioritizing scalability, reliability, and governance compliance.
  • Build scalable ingestion and transformation pipelines for both structured and unstructured data sources, applying OCR, NLP preprocessing, and document chunking optimized for large language model consumption.
  • Apply modern lakehouse patterns on Google Cloud Platform (GCP) with governance, cataloging, and lineage tracking to ensure data is discoverable, auditable, and AI-ready at scale.
  • Design and deploy vector databases, embedding pipelines, and knowledge graph structures that serve as the retrieval layer for RAG and AI applications.
  • Productionize AI solutions and advanced analytics within a DevOps/MLOps environment, including automated testing, monitoring, and rollback capabilities.
  • Foster innovation by proposing new ideas and selecting the right combination of tools and frameworks to convert business problems into analytics solutions.
  • Research and implement process improvements to address complex technology gaps, building a strong understanding of enabling technologies.

Requirements

  • Deep expertise building scalable ingestion and transformation pipelines across structured and unstructured data, with a proven track record migrating workloads to modern cloud platforms.
  • Experience parsing and normalizing diverse content types—PDFs, emails, images, call transcripts—using OCR and NLP preprocessing (tokenization, entity extraction, summarization) and document chunking optimized for LLMs.
  • Proven experience designing and implementing vector databases (Vertex AI Vector Search, Pinecone, pgvector), embedding pipelines, and knowledge graph structures for RAG and semantic search.
  • Strong SQL and data analytics skills; experience building data marts and feature datasets for data science and ML applications.
  • Strong Python coding skills; hands-on experience with BigQuery, Claude Code, RAG architectures, LLMs, ADK, and prompt engineering techniques.
  • Expertise in building ML platforms and data pipelines at scale; familiarity with ML algorithms, deep learning, NLP, information retrieval, and data mining techniques.
  • Experience with Google Cloud Platform services (Vertex AI, Dataflow, BigQuery, Cloud Run, Pub/Sub) and comfort with distributed computing frameworks (Apache Spark, Dataproc) for large-scale processing.
  • Solid experience managing diverse data sources with preprocessing, cleansing, and data integrity verification to meet ML requirements.
  • Demonstrated experience with machine learning, deep learning, NLP, information retrieval, or data mining—particularly applied to unstructured or semi-structured data.
  • Hands-on experience with vector databases, embedding models (text-embedding-gecko, OpenAI Ada, Cohere), and end-to-end RAG pipeline design.
  • Experience using Agile methods is preferred.
  • Strong communication and interpersonal skills; ability to collaborate effectively in a highly matrixed environment.
  • Preferred experience with the insurance industry, its products and services.
  • Experience implementing big data processing technologies; Apache Spark is preferred.

Technologies

  • Python
  • SQL
  • Java
  • BigQuery
  • Claude Code
  • Vertex AI
  • Vertex AI Vector Search
  • Pinecone
  • pgvector
  • OpenAI Ada
  • Cohere
  • text-embedding-gecko
  • Dataproc
  • Apache Spark
  • Dataflow
  • Cloud Run
  • Pub/Sub
  • Google Cloud Platform
  • ADK

Reporting relationship: Typically to a Director or above.

Similar Jobs