Senior Data Engineer - Technology
Job Description
Neptune Technology Group Inc. is seeking a Senior Data Engineer to own and modernize our data pipeline architecture. The role centers on migrating ETL workloads from Redshift stored procedures and legacy SSIS into scalable, maintainable pipelines built with AWS Glue and S3, while guiding the move toward near real-time analytics using stream processing technologies such as Apache Flink and ClickHouse.
This on-site position is based in Duluth, GA.
Responsibilities
- Design and build modern ETL/ELT pipelines using AWS Glue, S3, and related services to replace legacy stored procedures and SSIS jobs.
- Architect data transformation workflows that are testable, version-controlled, and observable.
- Optimize and maintain the Redshift data warehouse, including materialized views, query performance, and cost efficiency.
- Drive the evolution from batch ETL to near real-time stream processing, evaluating and implementing technologies such as Apache Flink, ClickHouse, Kafka, Kinesis, or equivalent platforms.
- Design pipelines that support both near real-time and batch workloads as the platform transitions.
- Partner with product and analytics teams to ensure data models support reporting, AI/ML, and customer-facing features.
- Establish patterns and best practices for pipeline development that the broader team can adopt.
- Participate in production support and incident response for data infrastructure.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field. Minimum of 5+ years in data engineering roles.
- Deep experience with AWS Glue (PySpark/Python), S3, and Redshift.
- Proven track record migrating ETL workloads from legacy tools (SSIS, stored procedures, or similar) to modern cloud-native pipelines.
- Strong SQL skills with best practices and SQL linting, particularly in Redshift or other columnar/MPP databases.
- Experience with or strong interest in stream processing frameworks (Flink, Spark Streaming, Kafka Streams, or similar).
- Familiarity with data pipeline orchestration, monitoring, and error handling patterns.
- Experience with infrastructure-as-code and CI/CD for data pipelines.
Technologies
- Python
- PySpark
- SQL
- AWS Glue
- S3
- Redshift
- SSIS
- Apache Flink
- ClickHouse
- Kafka
- Kinesis
- Apache Spark Streaming
- Kafka Streams
- Aurora MySQL
- DynamoDB
- Apache Druid
- dbt
- Airflow
- Step Functions
- MQTT
Location
Locations: Tallassee, Alabama or Duluth, Georgia. This is an onsite role in Duluth, GA, with travel to manufacturing or customer locations up to 20% as needed.