Data Engineer
Job Description
The Data Engineer role at Amazon.com Services LLC in Arlington, VA, onsite, focuses on building and maintaining data pipelines and analytics that reveal factors driving changes in employee sentiment, actions, and business outcomes. This position sits within Amazon’s PXT Central Science team and emphasizes scalable data infrastructure, model productionization, and cross-team collaboration.
Responsibilities
- Data Pipeline Development: Architect and sustain scalable data pipelines using AWS native services (Glue, EMR, Lambda); implement robust monitoring and error handling for data workflows; optimize for performance, reliability, and cost efficiency.
- Model Productionization & API Development: Build APIs and data serving layers to productionize science models for downstream consumption; design batch and real-time inference pipelines.
- Data Integration & Quality: Create scalable feature extraction and processing frameworks for diverse data types; implement strong data quality checks; design flexible schemas to support evolving requirements.
- Cross-team Collaboration: Partner with economics, data science, and software engineering teams to translate analytical requirements into production-ready solutions; contribute to technical design reviews and architecture discussions.
- Analytics & Infrastructure: Maintain layered data systems used by economists and scientists; develop automated reporting solutions; operate across multiple interconnected AWS accounts with security best practices.
Requirements
- Knowledge of professional software engineering practices across the full software development life cycle, including coding standards, architectures, code reviews, source control, continuous deployment, testing, and operational excellence.
- 3+ years of data engineering experience.
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS.
- Experience with data modeling, warehousing, and building ETL pipelines.
- Experience with AWS technologies such as Redshift, S3, AWS Glue, EMR, Kinesis, Firehose, Lambda, and IAM roles and permissions.
- Experience with non-relational databases or data stores (object storage, document or key-value stores, graph databases, column-family databases).
- Bachelor's degree or foreign equivalent in computer science, engineering, mathematics or an equivalent field.
Technologies
- Python
- Java
- Scala
- NodeJS
- AWS Glue
- EMR
- Lambda
- Redshift
- S3
- Kinesis
- Firehose
- IAM
- Hadoop
- Hive
- Spark
- Generative AI
Benefits
- Health insurance (medical, dental, vision, prescription)
- 401(k) matching
- Paid time off
- Parental leave
- Sign-on payments
- Restricted stock units (RSUs)
- Adoption and Surrogacy Reimbursement coverage
- Employee Assistance Program (EAP) and Mental Health Support
- Medical Advice Line
- Flexible Spending Accounts (FSA)
- Basic Life & AD&D insurance
- Optional supplemental life plans
About the Team
The Central Science Team within Amazon’s People Experience and Technology organization (PXTCS) combines economics, behavioral science, statistics, machine learning, and Generative AI to proactively identify mechanisms and process improvements that enhance Amazon and the lives, well-being, and value of work for Amazonians. This interdisciplinary group blends science, engineering, and UX to deliver solutions with measurable impact.