Data & AI Engineer
Job Description
The Carlyle Group is looking for a hands-on Data & AI Engineer to translate data and AI architecture into production-ready systems. This role focuses on building and operating AI-ready data pipelines, retrieval and semantic components, and data products that support analytics and generative AI applications across the firm.
Working as an individual contributor from New York, NY (onsite), you will help enable reliable retrieval for LLMs and agents, productionize priority RAG and agent-grounding use cases, and establish reusable engineering patterns that can be adopted across the federated data platform.
Responsibilities
- Build and operate AI-ready data pipelines including embedding generation, chunking, indexing, and refresh workflows so enterprise data is reliably retrievable by LLMs, agents, and generative AI applications.
- Implement retrieval-augmented generation (RAG) components such as vector store integrations, hybrid search, re-ranking, and grounding logic aligned to architectural patterns set by the Senior AI & Data Architect.
- Develop and maintain tool and function interfaces for agents and copilots to query and act on enterprise data with guardrails, logging, and evaluation hooks.
- Partner with Data Science and AI Engineering teams to operationalize feature stores, evaluation datasets, and reusable AI data products.
- Contribute to semantic and context engineering for natural-language analytics, conversational reporting, and AI-driven insights.
- Design, build, and maintain production-grade ELT, streaming, and transformation pipelines using tools including dbt, Fivetran, and Snowflake.
- Implement ingestion, modeling, and consumption patterns that meet enterprise standards for scalability, performance, security, resiliency, and cost efficiency.
- Write clean, well-tested Python and SQL, applying software engineering best practices such as version control, code review, CI/CD, modular design, and automated testing.
- Productionize new sources and domains under the federated operating model, partnering with domain data engineers to apply shared platform capabilities consistently.
- Implement semantic models, data contracts, and analytical/dimensional models that support trusted self-service analytics and reliable AI grounding.
- Build and maintain reusable data products with clear ownership, documented contracts, and contextual metadata for human and AI consumers.
- Collaborate with the Senior AI & Data Architect to refine and extend enterprise semantic standards based on production outcomes.
- Support discovery and consumption tooling so analysts, applications, and agents can find and use data products with minimal friction.
- Implement data quality checks, lineage capture, and pipeline observability across both data and AI workloads.
- Build logging, evaluation, and monitoring components for AI systems, including prompt and response capture, retrieval metrics, and model performance signals aligned with governance standards.
- Partner with Data Governance to operationalize metadata, stewardship, and access controls so AI systems consume enterprise data with the same rigor as human users.
- Surface issues early, propose remediations, and feed lessons learned back into architectural patterns.
- Participate in architectural design reviews and contribute a hands-on engineering perspective to evolving patterns and standards.
- Mentor junior data engineers and analysts on modern data and AI engineering practices.
- Document patterns, write runbooks, and share knowledge across the federated organization to accelerate adoption of reusable platform capabilities.
Requirements
- Bachelor’s degree (required).
- 6+ years of overall relevant technical experience (required).
- Experience in data engineering, analytics engineering, or platform engineering, with 1-2 years of hands-on experience building generative AI or AI/ML systems in production.
- Proven experience implementing retrieval, grounding, and semantic components for LLM- or agent-based applications, including RAG pipelines, vector stores, embedding workflows, and structured tool use.
- Hands-on experience with one or more modern AI platforms and tooling categories, such as AWS Bedrock, Databricks ML, Snowflake Cortex, OpenAI/Anthropic APIs, LangChain/LlamaIndex (or equivalents), MLflow, and vector databases including Databricks Vector Search, pgvector, or Pinecone.
- Strong expertise in Python and SQL, with working knowledge of distributed processing frameworks such as Spark.
- Deep hands-on experience with a modern data stack including dbt, Fivetran, Snowflake in AWS-based environments.
- Experience building data pipelines and products used by AI systems, not only BI tools and human analysts.
- Experience operating within federated data operating models and complex, regulated enterprise environments; financial services experience preferred.
Technologies
- Python, SQL, dbt, Fivetran, Snowflake, ELT, streaming
- Snowflake Cortex, AWS Bedrock, Databricks ML
- OpenAI/Anthropic APIs, LangChain/LlamaIndex, MLflow
- RAG, vector stores, hybrid search, re-ranking, embedding workflows, structured tool use
- Feature stores, vector databases: Databricks Vector Search, pgvector, Pinecone
- Spark, prompt and response capture, lineage capture, CI/CD
What Success Looks Like
- Within the first 12 months: deliver foundational AI-ready data pipelines and retrieval components defined in the target-state architecture.
- Productionize one or more priority RAG or agent-grounding use cases.
- Establish reusable engineering patterns for other domain teams to adopt across the federated data platform.
In-Office Requirement
- 4 days per week.
Location
- Washington, D.C. or New York, NY.
Benefits / Compensation
- Anticipated base salary range: $160,000 to $180,000.
- Discretionary incentive program eligibility based on individual and organizational performance.
- Comprehensive benefits package including retirement benefits, health insurance, life and disability insurance, paid time off, paid holidays, family planning benefits, and wellness programs.
- Compensation range varies by applicable office location and reflects required and preferred skill sets, prior experience and training, and licenses and/or certifications.