EngineerJobs.io
← Back to all jobs

Job Description

Caterpillar is seeking a Lead Data Engineer to help shape a physical AI platform by designing and delivering scalable data pipelines, microservices, and cloud data capabilities. This onsite role in Chicago brings together data architecture and production engineering to support reliable, high-quality data for business and engineering teams.

In this position, you will lead solution design and operational excellence across enterprise platforms, partnering with engineering and architecture leaders to translate complex needs into dependable workflows, integrations, and event-driven systems using AWS and Python.

Responsibilities

  • Collaborate with Principal Software Engineers and Data Architects to define solution architecture
  • Lead the design and optimization of scalable data pipelines and microservices in Python for both real-time and batch processing across enterprise platforms
  • Drive cloud-native data ingestion and streaming solutions using AWS services including Kinesis, S3, DynamoDB, EventBridge, and related technologies
  • Own data integration frameworks and source data pipelines that support CI Autonomy initiatives, including operational excellence
  • Partner with business, product, and engineering stakeholders to translate requirements into scalable data architectures, workflows, mappings, and system designs
  • Establish and enforce automated testing, data quality controls, and validation frameworks to support integrity, reliability, and compliance
  • Lead production monitoring, performance tuning, and root-cause analysis using observability tools such as CloudWatch to maintain high availability and service reliability

Requirements

  • Decision Making and Critical Thinking: lead analysis and resolution of complex issues across distributed data platforms, designing scalable and resilient solutions
  • Effective Communications: communicate across teams through constructive feedback, active listening, and documentation that makes data systems and processes easier to understand and support
  • Software Development: experience leading backend systems and data pipeline design and development using Python, Java, and modern frameworks, including technical direction to deliver reliable and scalable solutions
  • Software Development Life Cycle: guide delivery of data engineering solutions in an Agile environment, translating requirements into technical plans and ensuring quality, reliability, and business value
  • Software Integration Engineering: ability to lead design and integration of APIs, data pipelines, streaming platforms, and databases for reliable data exchange across enterprise systems
  • Software Product Design/Architecture: expertise in scalable, event-driven data systems and architectures, ensuring solutions are reliable, maintainable, and aligned to business needs
  • Software Product Technical Knowledge: strong knowledge of AWS services and data engineering tools to define requirements, support testing and deployment, troubleshoot issues, and operate data solutions across environments
  • Software Product Testing: define and implement testing strategies including functional, performance, and data quality testing across the development lifecycle

Technologies

  • Python, Java
  • AWS: Kinesis, S3, DynamoDB, EventBridge, CloudWatch
  • CI/CD, Azure DevOps, Jira, Jenkins
  • SQL, relational databases, NoSQL databases
  • Microservices, APIs
  • CI Autonomy
  • Helios Data Platform

Benefits

  • Medical, dental, and vision benefits
  • Paid time off plan (Vacation, Holidays, Volunteer, etc.)
  • 401(k) savings plans
  • Health Savings Account (HSA)
  • Flexible Spending Accounts (FSAs)
  • Health Lifestyle Programs
  • Employee Assistance Program
  • Voluntary Benefits and Employee Discounts
  • Career Development
  • Incentive bonus
  • Disability benefits
  • Life Insurance
  • Parental leave
  • Adoption benefits
  • Tuition Reimbursement
  • These benefits also apply to part-time employees

Top Candidates Will Have

  • Bachelor’s degree in Computer Science, Computer Engineering, or related field
  • 8+ years of experience in data engineering or related disciplines with increasing responsibility
  • Extensive experience on modern, large scale, complex Caterpillar data platforms such as Helios Data Platform
  • Strong foundation developing and deploying Python solutions to a production environment
  • Experience leading teams to build high-throughput, scalable data pipelines
  • Strong hands-on experience with AWS data services (Kinesis, S3, DynamoDB, EventBridge, etc.) at scale
  • Strong SQL skills, including data quality and validation practices
  • Experience deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc.
  • Experience developing microservices that support real-time data ingestion
  • Experience developing software applications using relational and NoSQL databases
  • Ability to ensure data integrity across distributed and streaming systems
  • Experience with monitoring, testing, and automation in large-scale data environments

Additional Details

  • Full-time position based in the Chicago, IL office
  • Onsite work is required five days a week
  • Domestic relocation assistance is available
  • Visa sponsorship is available for eligible applicants

Salary range: USD 128,470 - 208,770 per year

Similar Jobs