EngineerJobs.io
← Back to all jobs

Job Description

Navanta, LLC builds governed data foundations that support public Call Report data and secure bank-core ingestion as client environments are stood up. As the Lead Data Engineer, you’ll own the lakehouse backbone, working closely with AI/ML, Security, and platform teams to enable reliable, reconcilable data products for Navanta AI platforms.

What you’ll work on

  • Design and operate the lakehouse using Apache Iceberg (or similar) on object storage, including a catalog for table management and per-bank isolation, plus dbt models and a query engine.
  • Build secure, least-privilege ingestion from bank systems, using log-based CDC where permitted, with query-based and batch/SFTP fallbacks, and an in-bank collector pattern.
  • Own data modeling for the semantic and metric layer, covering deposits, concentration, uninsured exposure, asset quality, and peer groups.
  • Ensure correctness and resilience by handling schema drift, data quality, and reconciliation, while making ingestion observable and recoverable.
  • Partner across teams on the structured-query path with AI/ML and on PII classification at landing with Security, aligned with regulatory data-handling requirements.
  • Document for audit readiness by maintaining data lineage, transformation logic, and access controls to support exam and audit processes.
  • Define and enforce data contracts, including quality thresholds and alerting for pipeline failures.

Core expectations

This is a full-time role that combines ongoing pipeline operations with initiative-based lakehouse build-out and new bank onboarding. You will collaborate closely with AI/ML, platform engineering, and Security, including participation in an on-call rotation for data pipeline reliability.

What you bring

  • 8–12+ years in data engineering with end-to-end ownership from ingestion through serving, including 2+ years in a lead or senior role.
  • Strong Python and expert SQL, with rigorous data modeling for analytics.
  • Hands-on lakehouse experience (Iceberg/Delta/Hudi or equivalent) and modern transformation tooling.
  • Experience building reliable pipelines from messy operational and transactional source systems.
  • Comfort with CDC mechanics and extracting data from databases you do not control.
  • Bachelor’s degree in computer science, mathematics, information systems, or a related field, or equivalent hands-on experience.
  • Financial services or regulated data environment experience strongly preferred.

Key KPIs

  • Data freshness and pipeline reliability: SLAs met for data ingestion and bank-core feeds.
  • Data quality score across key metrics versus source reconciliation.
  • Time to onboard a new bank’s data environment from kickoff to a queryable lakehouse.
  • PII classification coverage at landing and zero unauthorized data-access incidents.
  • Semantic layer adoption: percentage of assistant queries resolved via governed metrics versus ad hoc SQL.

Technologies you’ll use

Python, SQL, Apache Iceberg (plus Polaris / Nessie / Lakekeeper), dbt, Trino / Presto / DuckDB, Debezium, Kafka / Redpanda, Dagster (or Airflow), and storage such as S3 / MinIO. Source and modeling include SQL Server and PostgreSQL, plus pgvector (or equivalent).

Nice to have

  • Experience with financial or core-banking data, including FFIEC / Call Report data specifically.
  • Strong SQL Server familiarity.
  • Experience with data contracts, lineage, and governance practices.

Location and work environment

Alpharetta, GA (onsite). Typical office environment with up to 20% travel time may be required.

Similar Jobs