Data Engineer 4 (Risk Tech)
Job Description
Capital One is building proprietary risk management solutions powered by state-of-the-art AI, and this role will help bring those cloud-first data capabilities to life. You will work in an environment that emphasizes scalable engineering patterns, strong security and compliance practices, and clear communication across teams delivering experiences that support financial empowerment for millions of Americans.
Location: McLean, VA (onsite). Salary: USD 197,300 - 225,100 per year.
Responsibilities
- Partner with Agile teams to design, develop, test, implement, and support technical solutions across full-stack development tools and technologies.
- Drive technical influence within a team of developers, data analysts, and data scientists with deep experience in machine learning, distributed microservices, lakehouse architecture, and full stack systems.
- Use Python and Spark, along with open-source relational and NoSQL databases and cloud data warehousing platforms such as Databricks and Snowflake.
- Stay current with data and AI trends by experimenting with and learning new technologies, participating in internal and external technology communities, and mentoring members of the data community.
- Collaborate with product managers and software engineering to deliver robust cloud-first data solutions.
- Independently design, build, and deliver world-class cloud data solutions and applications with little or no support from supervisors or managers.
- Architect and enforce common data engineering design patterns to improve code quality, maintainability, and reusability across platforms and pipelines.
- Design and build data pipelines and platforms focused on scalability, resilience, and operational efficiency to meet increasing data volume and business demands.
- Implement data security standards, including encryption at rest and in transit and fine-grained access control, to ensure compliance with data privacy regulations.
- Serve as an ambassador for the data engineering team by communicating technical concepts and data outcomes clearly to internal and external stakeholders.
Requirements
- Bachelor’s degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 4 years of experience in application development (internship experience does not apply).
- At least 2 years of experience in distributed data.
- At least 2 years of experience with SQL.
- At least 2 years of experience with one programming language: Python, Java, or Scala.
- At least 2 years of experience in data pipeline design and development.
- At least 1 year of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems.
Technologies
Python, Spark, Databricks, Snowflake, SQL, NoSQL, relational databases, distributed microservices, lakehouse architecture, encryption at rest/transit, fine-grained access control, EMR, Glue, Airflow, Dagster, Monte Carlo, Splunk, Mongo, Cassandra, DynamoDB, Redshift, AWS, Microsoft Azure, Google Cloud, Pinecone, Milvus, Qdrant, pgvector, Databricks Vector Search, LangChain, LlamaIndex
Benefits
- A comprehensive, competitive, and inclusive set of health, financial, and other benefits supporting your total well-being.
- Performance-based incentive compensation, which may include cash bonuses and/or long-term incentives (LTI).
Preferred Qualifications
- 7+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java.
- 4+ years of hands-on experience designing, deploying, and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud).
- 4+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks.
- 4+ years of experience designing, implementing, and operating real-time or streaming data pipelines.
- 2+ years of experience working on data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster).
- 4+ years of experience working with unstructured or semistructured data using NoSQL databases (e.g., Mongo, Cassandra, DynamoDB).
- 4+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift).
- 2+ years of experience working in an Agile development environment.
- 2+ years of experience developing user-centric reusable data products.
- Experience building and maintaining pipelines for vector embeddings or managing vector search solutions (e.g., Pinecone, Milvus, Qdrant, pgvector, Databricks Vector Search).
- Experience designing data pipelines and context retrieval frameworks to support large language model (LLM) integrations using tools such as LangChain or LlamaIndex.
- Experience preparing unstructured text, logs, or document datasets for Generative AI workflows, including text chunking and automated feature extraction.
Additional information
- This role is expected to accept applications for a minimum of 5 business days.
- No agencies please.
- Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws.
- Capital One promotes a drug-free workplace.
- Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with applicable laws.
- For accommodation requests during the application process, contact Capital One Recruiting at 1-800-304-9102 or [email protected].
- For technical support or questions about the recruiting process, email [email protected].