Data Engineer Lead
Apache Airflow
Application Security
Automation
Big Data
Bigdata
CI/CD
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Infrastructure
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
Databricks
Datadog
Dbt
Delta Lake
DevOps
Devops Tools
DevSecOps
ETL
Informatica
Information Technology (IT)
Infrastructure As Code
Project Management
Pyspark
Security Automation
Software Development
Spark
SQL
Workflow Orchestration
Job Description
Capital Group Companies is hiring a senior, hands-on Data Engineer Lead to guide the CSGT data platform and AI-first data products.
Responsibilities
- Own the data engineering strategy and roadmap for CSGT, including Lakehouse architecture on Databricks and AWS, and make decisions on scalability, security, reliability, and cost
- Set and influence standards and practices across adjacent teams and the wider Capital Group technology organization, contributing to centers of excellence and firm-wide engineering practices
- Design and build ingestion, transformation, and serving pipelines using Databricks, PySpark, Delta Lake, dbt, and Airflow, creating reusable patterns and frameworks for consistent, maintainable data products
- Evaluate new structured and unstructured datasets at the business-capability level and map them into platform data domains and subject areas
- Lead complex cross-team data initiatives end-to-end, covering requirements through production support, including estimates, work breakdown, sequencing, dependencies, and cost; surface risks early while balancing durable platform needs with business urgency
- Drive continued evolution of AI-first engineering by translating business outcomes into specifications and directing agents to plan, build, test, and document changes in small, reviewable increments
- Build reusable agent workflows, skills, and tool integrations for profiling, source-to-target mapping, pipeline and test generation, schema-change analysis, and incident investigation
- Define agent autonomy and human approval boundaries, plus how agent actions are reviewed and traced
- Make governed data understandable and reliable for AI using Unity Catalog metadata, lineage, business definitions, semantic models, and access controls; connect datasets to Databricks Genie and other AI applications used by investment professionals
- Define evaluation datasets and acceptance criteria for agents and AI-generated SQL and code
- Set and validate the testing strategy across platform layers, including performance, stability, and availability; review and approve quality metrics before release
- Embed data quality, reconciliation, freshness, observability, and recovery controls into automated testing and CI/CD, including security and policy checks from the start
- Partner with investment professionals and product managers on shared product vision and ownership of business outcomes, demonstrating how data and AI support research and portfolio construction at scale
- Raise the engineering bar through design and code reviews, day-to-day leadership for engineers on initiatives, and guidance through complex data and performance issues
- Develop engineers by teaching how to inspect and challenge AI-generated work, share reusable patterns and context through forums, and help managers identify strengths and development needs
Requirements
- 10+ years of experience in data or software engineering, including technical leadership of complex production data platforms delivered across multiple teams
- Strong hands-on Python and SQL, sound software design judgment, and deep understanding of distributed data processing, query performance, and automated testing
- Production experience with Databricks on AWS, including PySpark, Delta Lake, Unity Catalog, Databricks Jobs, Databricks SQL, and Databricks Asset Bundles; understand security, access, and cost implications of designs
- Experience orchestrating production pipelines with Apache Airflow (including Astronomer) and building tested transformations with dbt, including reliable retries, backfills, and dependency management
- Strong data modeling and governance experience, including dimensional and time-series models, slowly changing dimensions, bi-temporal history, data contracts, lineage, and semantic metadata
- Implemented data quality, observability, and CI/CD for data platforms using tools such as Deequ, dbt tests, Lakehouse Monitoring, Datadog, Terraform, and Harness
- Experience preparing governed data for AI through natural-language-to-SQL tools such as Databricks Genie, semantic metadata, or other governed data-access patterns
- Use AI coding agents beyond code completion by writing specifications, supplying context, running tests, and reviewing generated changes via source control
- Evaluate AI-generated output with representative test cases, regression tests, execution traces, and human review; distinguish plausible answers from verified results
- Understand prompt injection, sensitive data handling, and least-privilege access; design approval boundaries and audit trails for agents working against enterprise systems
- Lead architecture discussions, influence without formal authority, develop other engineers, and explain technical choices and trade-offs clearly to investment professionals and technology leaders
- Act as an agent of change with urgency, questioning how work is done and automating or removing what does not add value while respecting existing context
- Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience)
Technologies
- Databricks, AWS
- Python, SQL, PySpark, Delta Lake
- Unity Catalog, Databricks Jobs, Databricks SQL, Databricks Asset Bundles
- Apache Airflow, Astronomer, dbt, Airflow
- Deequ, Lakehouse Monitoring, Datadog
- Terraform, Harness
- Databricks Genie, CI/CD
Preferred Qualifications
- Experience with investment management data such as portfolios, positions, returns, exposures, benchmarks, and attribution, or multi-asset portfolio construction
- Experience building LLM applications or agent workflows that call tools and APIs, including context management, retrieval, state, and error handling
- Familiarity with Model Context Protocol (MCP)
- Experience with PostgreSQL, SQL Server, or Lakebase, or modernizing legacy data platforms onto a Lakehouse
Benefits
- Generous time-away and health benefits from day one, with the opportunity for flexible work options
- 2-for-1 matching gifts for charitable contributions
- Opportunity to secure annual grants for the organizations you love
- Access on-demand professional development resources
- Competitive salary, bonuses and benefits
- Company-funded retirement contribution
- Individual annual performance bonus
- Capital’s annual profitability bonus
- Retirement plan where Capital contributes 15% of eligible earnings
Location
- Los Angeles, CA (onsite)
Salary
- USD 201,683 - 342,072 per yearly
- Southern California base salary range: $201,683-$322,693
- New York base salary range: $213,795-$342,072
Minimum Experience: 10 years
Education: Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience)