Senior Data Engineer
Senior
3d Design Tools
Amazon Athena
Amazon Web Services
AWS
Aws Glue
Big Data
Bigdata
Cloud
Cloud Computing
Cloud Data Warehouse
Cloud Operations
Cloud Platforms
Data
Data Analysis
Data Architecture
Data Engineer
Data Integration
Data Lake
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
EMR
ETL
Informatica
Information Technology (IT)
IT Services
Kafka
Rendering Engines
Spark
SQL
Stream Processing
Job Description
The Senior Data Engineer role at LTM focuses on designing, building, and optimizing scalable data solutions on AWS. The position emphasizes PySpark-based big data processing, data warehousing concepts with Apache Hive, and modern data lake table formats including Apache Iceberg.
Responsibilities
- Design, develop, and deploy scalable, cost-effective data solutions on AWS using services including S3 for data lakes, EC2, EMR, Glue, Athena, Lambda, Redshift, and Kinesis
- Build and maintain robust ETL/ELT pipelines using PySpark for data ingestion, transformation, and loading into data stores, including environments that use open table formats such as Iceberg
- Develop and optimize big data processing jobs using PySpark on AWS EMR or AWS Glue, handling large datasets and integrating with Iceberg table formats
- Implement and manage data warehousing solutions, including schema design, data modeling, and query optimization, with an approach that supports historical and analytical workloads using Hive and Iceberg
- Provide secure and robust cloud infrastructure components, including VPCs, subnets, routing, and security groups, to support proper connectivity and isolation for data solutions
- Design, deploy, and manage containerized data processing applications on Amazon EKS
- Optimize performance and efficiency by tuning AWS resources and big data applications across Spark, Hive, and Iceberg
- Apply data governance, security, and compliance best practices in AWS, including IAM policies, S3 bucket policies, and encryption
- Set up monitoring, logging, and troubleshooting for data pipelines and AWS infrastructure to resolve issues promptly
- Develop and maintain automation scripts using Python and shell scripting for infrastructure provisioning, deployment, and operational tasks
- Collaborate with data scientists, analysts, and other engineering teams to understand data requirements and deliver reliable data solutions
Required Qualifications
- At least one AWS certification, for example: AWS Certified Solutions Architect – Associate, AWS Certified Data Analytics – Specialty, or AWS Certified Developer – Associate
- Hands-on experience with key AWS services for data processing and storage, including S3, EC2, EMR, Glue, Athena, and Lambda
- Experience with VPC, subnets, routing, and security groups
- Experience with EKS
- Strong proficiency in PySpark for complex data transformations and analytics
- Practical experience with Apache Iceberg for managing and querying data lakes
- In-depth knowledge and practical experience with Apache Hive for data storage, querying, and schema management
- Expert-level Python proficiency for scripting data manipulation and AWS automation using Boto3
- Proficiency in shell scripting for automation and operational tasks
- Strong SQL skills for data querying and manipulation
- Solid understanding of ETL/ELT, data modeling, distributed computing, and data governance
Technologies
- AWS, PySpark, Amazon S3, EC2, EMR, AWS Glue, AWS Athena, AWS Lambda, Amazon Redshift, Amazon Kinesis
- Apache Iceberg, Apache Hive, VPC, subnets, routing, security groups
- Amazon Elastic Kubernetes Service (EKS), Kubernetes
- IAM policies, encryption, Python, shell scripting, SQL, Boto3
- Apache Spark, SparkSQL, data lakes, ETL/ELT, Iceberg, Spark Hive Iceberg
- Apache Airflow, Git, Apache Kafka, Flink, Presto, AWS CodePipeline, GitHub Actions, GitLab CI
Benefits
- Comprehensive Medical Plan covering medical, dental, and vision
- Short Term and Long-Term Disability coverage
- 401(k) plan with company match
- Life insurance
- Vacation time, sick leave, and paid holidays
- Paid paternity and maternity leave
Good to Have Skills
- Orchestration experience with workflow orchestration tools such as Apache Airflow
- CICD experience with tools and practices including AWS CodePipeline, GitHub Actions, and GitLab CI
- Version control proficiency using Git
- Exposure to other big data technologies including Apache Kafka, Flink, or Presto
- Containerization and orchestration experience with Kubernetes for deploying and managing containerized applications
Certifications
- AWS Certified Solutions Architect – Associate
- AWS Certified Data Analytics – Specialty
- AWS Certified Developer – Associate
Role Details
- Work location: Tampa, FL (onsite)
- Salary: USD 83,912 - 113,900 per year
- Mandatory: Karat interview