Lead Data Engineer
Job Description
Navanta, LLC builds governed data foundations that support public Call Report data and secure bank-core ingestion as client environments are stood up. As the Lead Data Engineer, you’ll own the lakehouse backbone, working closely with AI/ML, Security, and platform teams to enable reliable, reconcilable data products for Navanta AI platforms.
What you’ll work on
- Design and operate the lakehouse using Apache Iceberg (or similar) on object storage, including a catalog for table management and per-bank isolation, plus dbt models and a query engine.
- Build secure, least-privilege ingestion from bank systems, using log-based CDC where permitted, with query-based and batch/SFTP fallbacks, and an in-bank collector pattern.
- Own data modeling for the semantic and metric layer, covering deposits, concentration, uninsured exposure, asset quality, and peer groups.
- Ensure correctness and resilience by handling schema drift, data quality, and reconciliation, while making ingestion observable and recoverable.
- Partner across teams on the structured-query path with AI/ML and on PII classification at landing with Security, aligned with regulatory data-handling requirements.
- Document for audit readiness by maintaining data lineage, transformation logic, and access controls to support exam and audit processes.
- Define and enforce data contracts, including quality thresholds and alerting for pipeline failures.
Core expectations
This is a full-time role that combines ongoing pipeline operations with initiative-based lakehouse build-out and new bank onboarding. You will collaborate closely with AI/ML, platform engineering, and Security, including participation in an on-call rotation for data pipeline reliability.
What you bring
- 8–12+ years in data engineering with end-to-end ownership from ingestion through serving, including 2+ years in a lead or senior role.
- Strong Python and expert SQL, with rigorous data modeling for analytics.
- Hands-on lakehouse experience (Iceberg/Delta/Hudi or equivalent) and modern transformation tooling.
- Experience building reliable pipelines from messy operational and transactional source systems.
- Comfort with CDC mechanics and extracting data from databases you do not control.
- Bachelor’s degree in computer science, mathematics, information systems, or a related field, or equivalent hands-on experience.
- Financial services or regulated data environment experience strongly preferred.
Key KPIs
- Data freshness and pipeline reliability: SLAs met for data ingestion and bank-core feeds.
- Data quality score across key metrics versus source reconciliation.
- Time to onboard a new bank’s data environment from kickoff to a queryable lakehouse.
- PII classification coverage at landing and zero unauthorized data-access incidents.
- Semantic layer adoption: percentage of assistant queries resolved via governed metrics versus ad hoc SQL.
Technologies you’ll use
Python, SQL, Apache Iceberg (plus Polaris / Nessie / Lakekeeper), dbt, Trino / Presto / DuckDB, Debezium, Kafka / Redpanda, Dagster (or Airflow), and storage such as S3 / MinIO. Source and modeling include SQL Server and PostgreSQL, plus pgvector (or equivalent).
Nice to have
- Experience with financial or core-banking data, including FFIEC / Call Report data specifically.
- Strong SQL Server familiarity.
- Experience with data contracts, lineage, and governance practices.
Location and work environment
Alpharetta, GA (onsite). Typical office environment with up to 20% travel time may be required.