Lead Data Engineer – AI/Machine Learning
Manager
Apache Airflow
Artificial Intelligence
Big Data
Bigdata
Cloud Data Engineering
Cloud Data Platform
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Science
Data Warehouse
Database
Databases
Databricks
DevOps
ETL
Informatica
Llm Operations
Machine Learning
Ml Ops
Programming Language
Programming Languages
Spark
Streaming Data
Vector Databases
Job Description
Core Specialty is hiring a Lead Data Engineer – AI/Machine Learning for a hybrid role based in Cincinnati, OH. Reporting to the VP, Head of Data, this position helps shape and drive organizational AI and ML enablement, moving from hands-on work on data platforms and pipelines to leading how AI/ML capabilities are built, deployed, and governed.
The role starts with direct technical ownership, then progressively expands into defining frameworks, advising on emerging AI/ML practices, and partnering across engineering and governance to close AI/ML readiness gaps.
What you’ll do
- Design, build, and optimize data pipelines, ingestion frameworks, and data platform components that support analytics, reporting, and AI/ML use cases.
- Own complex engineering initiatives independently, from technical design through implementation and rollout, with minimal oversight.
- Identify and resolve performance, scalability, and reliability issues within the existing data platform.
- Propose innovative, well-reasoned improvements to data engineering challenges, including proactively identifying gaps.
- Write clean, well-tested, well-documented code and infrastructure-as-code while maintaining strong engineering hygiene.
- Contribute to defining the organization’s AI/ML frameworks, including evaluating and recommending tools, platforms, and standards for AI/ML delivery.
- Build working prototypes that deliver value to engineering teams.
- Help shape and implement an MLOps strategy, covering model deployment, monitoring, versioning, and lifecycle management approaches.
- Collaborate with Data Governance to ensure AI/ML frameworks align with data governance, security, and compliance standards.
- Design and advocate for scalable data infrastructure patterns for AI/ML, such as feature stores, curated/governed datasets, and streaming access for training and inference.
- Partner with Data Science, Data Engineering, and business stakeholders to assess AI/ML readiness gaps and build a roadmap to address them.
- Serve as a subject-matter expert to advise the VP, Head of Data on emerging AI/ML technologies, practices, and industry trends.
- Document AI/ML standards, frameworks, and decisions to support consistent adoption as practices mature.
- Act as a senior technical resource by guiding architecture, design patterns, and best practices for AI readiness and ML Ops frameworks.
- Work closely with Enterprise Architecture to establish architectural blueprints for AI readiness.
- Other duties as assigned.
Key requirements
- 7+ years of experience in data engineering, including work on large-scale, mature data platforms.
- 3+ years developing ML or AI deliverables, including deployment to production.
- Strong data engineering fundamentals, including expertise in data pipeline design, optimization, and distributed data processing (for example: Spark, dbt, Airflow, Kafka, or equivalent).
- Hands-on experience with Snowflake, Databricks, and/or Azure Synapse Analytics, with ability to architect and optimize workloads on one or more of these platforms.
- Strong cloud knowledge across AWS, Azure, or GCP, plus modern data warehouse and lakehouse architecture experience.
- Strong Python skills and solid software engineering practices (testing, version control, code review) for production system delivery.
- API design and integration experience, including orchestration, tool-calling, and retrieval-oriented systems around models.
- Practical experience with LLM APIs (for example, OpenAI) and open-weight models.
- Prompt engineering and prompt evaluation as a disciplined practice.
- Knowledge of context windows, tokenization, embeddings, and model limitations such as hallucination, latency, and cost tradeoffs.
- Experience with vector databases (for example: Pinecone, Weaviate, pgvector) and embedding models.
- Experience with chunking strategies, hybrid search, and reranking.
- Experience with frameworks such as LangChain, LangGraph, LlamaIndex, or custom orchestration.
- Design experience for tool use and function-calling, multi-step reasoning chains, and agent memory or state management.
- Understanding of when to fine-tune versus prompt versus RAG.
- Familiarity with parameter-efficient methods such as LoRA, as well as MLOps/LLMOps.
- Experience with model evaluation frameworks, including A/B testing for model outputs and observability (tracing, logging model calls).
- Deployment pattern knowledge including latency/cost optimization, caching, streaming responses, and fallback handling.
- Experience with versioning prompts and models, not only code.
- Awareness of safety, evaluation, and governance, including bias/safety evaluation and appropriate handling of PII.
- Agentic workflows for engineering and architecture: working knowledge required.
- Demonstrated ability to independently own complex technical projects from design to delivery with minimal oversight.
- Experience shaping AI/ML enablement (framework definition, evaluating MLOps tooling, or building infrastructure for model training and deployment).
- Experience partnering with Data Governance, Data Science, or Compliance to align technical practices with governance and regulatory needs.
- Experience proposing and driving technical solutions rather than only executing predefined plans.
Preferred experience
- Experience designing or implementing agentic workflows for data engineering.
- Experience with Property & Casualty insurance carriers.
- Experience with Data Vault 2.0 or ensemble data modeling techniques.
Education
Bachelor’s degree in a related field or demonstrated equivalent experience in a related field is required.
Technologies
- Spark, dbt, Airflow, Kafka
- Snowflake, Databricks, Azure Synapse Analytics
- AWS, Azure, GCP
- Python, OpenAI
- Pinecone, Weaviate, pgvector
- LangChain, LangGraph, LlamaIndex
- LoRA, MLOps, LLMOps
- Data Vault 2.0, Ensemble data modeling techniques
Benefits
- Medical, dental, vision, and life insurance
- Short- and long-term disability
- 401(k) plan with 100% company match of a 6% contribution
- Employee Assistance Plan
- Health Savings Account
- Flexible Spending Account
- Health Reimbursement Account
- Wellness program
- Opportunities for professional development and advancement