Data Engineer 5
Apache Airflow
Artificial Intelligence
Azure
Big Data
Bigdata
Cassandra
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Data Warehouse
Cloud Data Warehouse
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Science
Data Warehouse
Data Warehousing
Database
Databases
Databricks
ETL
Google Cloud
Informatica
Information Technology (IT)
MongoDB
Programming Language
Programming Languages
Snowflake
SQL
Job Description
Capital One is hiring a Data Engineer 5 in McLean, VA (onsite) to design and deliver cloud-first data solutions and lead end-to-end large-scale initiatives.
Responsibilities
- Collaborate with and across Agile teams to design, develop, test, implement, and support technical solutions
- Guide a team of developers, data analysts, and data scientists with experience across machine learning, distributed microservices, lakehouse architecture, and full-stack systems
- Build with Python and Spark using open-source relational and NoSQL databases, plus cloud data warehousing platforms including Databricks and Snowflake
- Stay current with data trends; experiment with and learn new technologies; participate in internal and external technology communities; mentor the data engineering community
- Partner with product managers and software engineers to deliver robust cloud-first data solutions used by millions of Americans
- Independently design, build, and deliver cloud data solutions and applications with little or no support from supervisors or managers
- Architect and apply consistent data engineering design patterns to improve code quality, maintainability, and reusability across platforms and pipelines
- Design and build data pipelines and platforms for scalability, resilience, and operational efficiency under growing data volume and business demands
- Act as a force multiplier by balancing hands-on technical work, innovation, and mentoring to raise peer and junior engineer capabilities
- Lead end-to-end, large-scale transformative data initiatives by driving key architectural decisions and evaluating platform options (for example, Snowflake vs Databricks) against technical and business requirements
- Serve as an ambassador for data engineering by communicating technical concepts and data outcomes clearly to internal and external stakeholders
Requirements
- Bachelor’s Degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- 6+ years of experience in application development (internship experience does not apply)
- 4+ years of experience in distributed data
- 4+ years of experience with SQL
- 4+ years programming with at least one of: Python, Java, Scala
- 4+ years designing and developing data pipelines
- 2+ years in data modeling and designing end-to-end data solutions using both relational and non-relational databases
Technologies
- Python, Spark
- Databricks, Snowflake
- SQL, NoSQL, open-source relational databases
- Distributed microservices, lakehouse architecture
- Machine learning
- EMR, Glue, Airflow, Dagster
- Monte Carlo, Splunk
- AWS, Microsoft Azure, Google Cloud
- MongoDB, Cassandra, DynamoDB, Redshift
- Scala, Java
Preferred Qualifications
- Master’s Degree in Computer Science or related field
- 8+ years of experience in data engineering
- 4+ years of data modeling experience
- 9+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java
- 5+ years hands-on experience designing, deploying, and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud)
- 5+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks
- 5+ years of experience designing, implementing, and operating real-time or streaming data pipelines
- 3+ years of experience in data observability (e.g., Monte Carlo, Splunk) or data orchestration (e.g., Airflow, Dagster)
- 5+ years of experience with unstructured or semistructured data using NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB)
- 5+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift)
- 3+ years of experience working in an Agile development environment
- 3+ years of experience developing user-centric reusable data products
Compensation
- USD 229,900 - 262,400 per year
Benefits
- Eligible for performance-based incentive compensation, potentially including cash bonuses and/or long-term incentives (LTI)
- Comprehensive, competitive, and inclusive health, financial, and other benefits supporting total well-being