Data Engineer
Apache Kafka
Apache Kafka Connect
API
Big Data
Cloud Data Engineering
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Integration
Data Lake
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
ETL
Informatica
Kafka
Kafka Schema Registry
Pandera
Protocol Buffers
SQL
Stream Processing
Job Description
This Data Engineer role at Infinitive Inc (onsite in McLean, VA) focuses on designing, building, and scaling next-generation event-driven data platforms. The position combines Apache Kafka streaming with Temporal durable workflow orchestration, along with production-grade schema controls and batch and streaming ETL/ELT using Python and Apache Spark.
Role Responsibilities
- Architect, deploy, and maintain high-volume distributed data streams using Apache Kafka, including producers, consumers, Kafka Connect, and Schema Registry.
- Define and enforce schema design standards, versioning strategies, and automated schema validation across systems. Use frameworks such as Avro, Protocol Buffers, and JSON Schema to maintain data contracts for microservices, streaming consumers, and lakehouse storage.
- Implement durable execution workflows with Temporal to coordinate long-running distributed pipelines, apply transaction compensation using the Saga pattern, and manage cross-system ETL tasks.
- Build end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark (including PySpark).
- Design and optimize analytical data models, including dimensional/star schema patterns, in cloud data warehouses and lakehouses such as Snowflake, BigQuery, Databricks, and Redshift.
- Create and maintain automated testing, continuous schema validation, data drift detection, and observability for both streaming and batch workflows.
- Collaborate with software engineers, machine learning engineers, and analysts to establish standard schema definitions, data contracts, and production-ready CI/CD release patterns.
Requirements
- 4+ years of professional experience in data engineering, backend distributed systems, or software engineering.
- Hands-on experience with Temporal (or Cadence) demonstrating knowledge of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration.
- Deep expertise with Apache Kafka, including message partitioning, consumer groups, offset management, and topic design.
- Practical experience with schema definition frameworks such as Apache Avro, Protocol Buffers (including gRPC), or JSON Schema.
- Experience managing schema evolution and compatibility modes (backward, forward, full), including working with schema registries such as Confluent Schema Registry and AWS Glue Schema Registry.
- Experience enforcing data validation rules, contract testing, and data quality checks using tools such as Great Expectations, Pandera, Pydantic, and dbt tests.
- Strong programming skills in Python (Go or Java is a plus), with emphasis on clean code, design patterns, and unit/integration testing practices.
- Hands-on development with Apache Spark using PySpark and Spark SQL for large-scale dataset processing.
- Strong experience with relational databases, dimensional modeling, and query performance tuning.
Technology Stack
- Apache Kafka, Kafka Connect, Schema Registry
- Temporal, Cadence
- Apache Avro, Protocol Buffers, gRPC, JSON Schema
- Confluent Schema Registry, AWS Glue Schema Registry
- Great Expectations, Pandera, Pydantic, dbt
- Python, SQL
- Apache Spark, PySpark, Spark SQL
- Snowflake, BigQuery, Databricks, Redshift
Compensation
A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00 per year.
Similar Jobs
S
S