Senior Data AI Engineer
Senior
Artificial Intelligence
Big Data
Bigdata
Cloud
Cloud Native
Cloud Operations
Cloud Platforms
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Governance
Data Integration
Data Lake
Data Pipeline
Data Platform
Data Processing
Data Security
Data Warehouse
Database
Databases
ETL
Google Cloud
Google Cloud Platform
Machine Learning
Machine Learning Engineer
Programming Language
Programming Languages
Spark
SQL
Vertex Ai
Job Description
Join CNA Insurance in Chicago as a Senior Data AI Engineer. This onsite role focuses on designing, building, and operating end-to-end AI and ML solutions that migrate CNA to a modern cloud data lakehouse. You will work across structured and unstructured data, unlocking analytical value through scalable pipelines, retrieval augmented generation architectures, vector databases, and knowledge graphs. As a senior individual contributor, you will influence data strategies and may guide others, contributing to AI readiness at scale.
Location: Chicago, IL (onsite). Salary range: USD 72,000 - 141,000 per year.
Responsibilities
- Design and implement AI capabilities that accelerate legacy-to-cloud migrations, prioritizing scalability, reliability, and governance compliance.
- Build scalable ingestion and transformation pipelines for both structured and unstructured data sources, applying OCR, NLP preprocessing, and document chunking optimized for large language model consumption.
- Apply modern lakehouse patterns on Google Cloud Platform (GCP) with governance, cataloging, and lineage tracking to ensure data is discoverable, auditable, and AI-ready at scale.
- Design and deploy vector databases, embedding pipelines, and knowledge graph structures that serve as the retrieval layer for RAG and AI applications.
- Productionize AI solutions and advanced analytics within a DevOps/MLOps environment, including automated testing, monitoring, and rollback capabilities.
- Foster innovation by proposing new ideas and selecting the right combination of tools and frameworks to convert business problems into analytics solutions.
- Research and implement process improvements to address complex technology gaps, building a strong understanding of enabling technologies.
Requirements
- Deep expertise building scalable ingestion and transformation pipelines across structured and unstructured data, with a proven track record migrating workloads to modern cloud platforms.
- Experience parsing and normalizing diverse content types—PDFs, emails, images, call transcripts—using OCR and NLP preprocessing (tokenization, entity extraction, summarization) and document chunking optimized for LLMs.
- Proven experience designing and implementing vector databases (Vertex AI Vector Search, Pinecone, pgvector), embedding pipelines, and knowledge graph structures for RAG and semantic search.
- Strong SQL and data analytics skills; experience building data marts and feature datasets for data science and ML applications.
- Strong Python coding skills; hands-on experience with BigQuery, Claude Code, RAG architectures, LLMs, ADK, and prompt engineering techniques.
- Expertise in building ML platforms and data pipelines at scale; familiarity with ML algorithms, deep learning, NLP, information retrieval, and data mining techniques.
- Experience with Google Cloud Platform services (Vertex AI, Dataflow, BigQuery, Cloud Run, Pub/Sub) and comfort with distributed computing frameworks (Apache Spark, Dataproc) for large-scale processing.
- Solid experience managing diverse data sources with preprocessing, cleansing, and data integrity verification to meet ML requirements.
- Demonstrated experience with machine learning, deep learning, NLP, information retrieval, or data mining—particularly applied to unstructured or semi-structured data.
- Hands-on experience with vector databases, embedding models (text-embedding-gecko, OpenAI Ada, Cohere), and end-to-end RAG pipeline design.
- Experience using Agile methods is preferred.
- Strong communication and interpersonal skills; ability to collaborate effectively in a highly matrixed environment.
- Preferred experience with the insurance industry, its products and services.
- Experience implementing big data processing technologies; Apache Spark is preferred.
Technologies
- Python
- SQL
- Java
- BigQuery
- Claude Code
- Vertex AI
- Vertex AI Vector Search
- Pinecone
- pgvector
- OpenAI Ada
- Cohere
- text-embedding-gecko
- Dataproc
- Apache Spark
- Dataflow
- Cloud Run
- Pub/Sub
- Google Cloud Platform
- ADK
Reporting relationship: Typically to a Director or above.