EngineerJobs.io
← Back to all jobs

Job Description

The Senior Data Engineer role at LTM focuses on designing, building, and optimizing scalable data solutions on AWS. The position emphasizes PySpark-based big data processing, data warehousing concepts with Apache Hive, and modern data lake table formats including Apache Iceberg.

Responsibilities

  • Design, develop, and deploy scalable, cost-effective data solutions on AWS using services including S3 for data lakes, EC2, EMR, Glue, Athena, Lambda, Redshift, and Kinesis
  • Build and maintain robust ETL/ELT pipelines using PySpark for data ingestion, transformation, and loading into data stores, including environments that use open table formats such as Iceberg
  • Develop and optimize big data processing jobs using PySpark on AWS EMR or AWS Glue, handling large datasets and integrating with Iceberg table formats
  • Implement and manage data warehousing solutions, including schema design, data modeling, and query optimization, with an approach that supports historical and analytical workloads using Hive and Iceberg
  • Provide secure and robust cloud infrastructure components, including VPCs, subnets, routing, and security groups, to support proper connectivity and isolation for data solutions
  • Design, deploy, and manage containerized data processing applications on Amazon EKS
  • Optimize performance and efficiency by tuning AWS resources and big data applications across Spark, Hive, and Iceberg
  • Apply data governance, security, and compliance best practices in AWS, including IAM policies, S3 bucket policies, and encryption
  • Set up monitoring, logging, and troubleshooting for data pipelines and AWS infrastructure to resolve issues promptly
  • Develop and maintain automation scripts using Python and shell scripting for infrastructure provisioning, deployment, and operational tasks
  • Collaborate with data scientists, analysts, and other engineering teams to understand data requirements and deliver reliable data solutions

Required Qualifications

  • At least one AWS certification, for example: AWS Certified Solutions Architect – Associate, AWS Certified Data Analytics – Specialty, or AWS Certified Developer – Associate
  • Hands-on experience with key AWS services for data processing and storage, including S3, EC2, EMR, Glue, Athena, and Lambda
  • Experience with VPC, subnets, routing, and security groups
  • Experience with EKS
  • Strong proficiency in PySpark for complex data transformations and analytics
  • Practical experience with Apache Iceberg for managing and querying data lakes
  • In-depth knowledge and practical experience with Apache Hive for data storage, querying, and schema management
  • Expert-level Python proficiency for scripting data manipulation and AWS automation using Boto3
  • Proficiency in shell scripting for automation and operational tasks
  • Strong SQL skills for data querying and manipulation
  • Solid understanding of ETL/ELT, data modeling, distributed computing, and data governance

Technologies

  • AWS, PySpark, Amazon S3, EC2, EMR, AWS Glue, AWS Athena, AWS Lambda, Amazon Redshift, Amazon Kinesis
  • Apache Iceberg, Apache Hive, VPC, subnets, routing, security groups
  • Amazon Elastic Kubernetes Service (EKS), Kubernetes
  • IAM policies, encryption, Python, shell scripting, SQL, Boto3
  • Apache Spark, SparkSQL, data lakes, ETL/ELT, Iceberg, Spark Hive Iceberg
  • Apache Airflow, Git, Apache Kafka, Flink, Presto, AWS CodePipeline, GitHub Actions, GitLab CI

Benefits

  • Comprehensive Medical Plan covering medical, dental, and vision
  • Short Term and Long-Term Disability coverage
  • 401(k) plan with company match
  • Life insurance
  • Vacation time, sick leave, and paid holidays
  • Paid paternity and maternity leave

Good to Have Skills

  • Orchestration experience with workflow orchestration tools such as Apache Airflow
  • CICD experience with tools and practices including AWS CodePipeline, GitHub Actions, and GitLab CI
  • Version control proficiency using Git
  • Exposure to other big data technologies including Apache Kafka, Flink, or Presto
  • Containerization and orchestration experience with Kubernetes for deploying and managing containerized applications

Certifications

  • AWS Certified Solutions Architect – Associate
  • AWS Certified Data Analytics – Specialty
  • AWS Certified Developer – Associate

Role Details

  • Work location: Tampa, FL (onsite)
  • Salary: USD 83,912 - 113,900 per year
  • Mandatory: Karat interview

Similar Jobs