EngineerJobs.io
← Back to all jobs

Job Description

GE Aerospace’s Commercial Engine Services BI team is seeking a Staff Data Engineer to help turn raw operational data into reliable, analytics-ready datasets used for real-time and batch reporting, as well as AI/ML applications. In this role, you will design and run production data pipelines with a strong focus on scalability, data quality, monitoring, and close collaboration with BI and software and AI engineering partners.

Based in Evendale, OH, this position is eligible for fully remote arrangements across the United States, with an in-person requirement for New Hire Orientation on Day 1. The base pay range is $112,000 to $150,000 per year, with potential eligibility for an annual discretionary bonus.

Role focus

  • Design, build, and maintain production-grade data pipelines that transform raw operational inputs into datasets for applications, reports, and AI/ML models.
  • Implement transformation logic using a multi-layer medallion architecture, including data cleaning, enrichment, aggregation, and business-rule implementation.
  • Create incremental loading patterns, support schema evolution, and use data versioning strategies to maintain reliability and backward compatibility.
  • Schedule and orchestrate automated refresh workflows to support real-time reporting and recurring batch processing.
  • Optimize performance for large datasets using partitioning, caching, indexing, and aggregation approaches that meet dashboard performance needs.

Data quality, monitoring, and reliability

  • Troubleshoot pipeline failures, data quality problems, and performance bottlenecks, then implement fixes and preventive measures.
  • Build automated data quality checks for null validation, range checks, referential integrity, business rule enforcement, and schema drift detection.
  • Implement data validation frameworks to surface issues early before they impact downstream dashboards or models.
  • Monitor data quality metrics and alerts, investigate anomalies, communicate with stakeholders, and coordinate remediation with source system owners.
  • Create monitoring systems and reporting that provide visibility into pipeline health, freshness, record counts, and quality trends over time.
  • Implement monitoring and alerting for pipeline execution details including failures, freshness, quality issues, compute costs, and execution times.
  • Perform root cause analysis for data incidents, document findings, and implement preventive actions.
  • Document known data quality issues, workarounds, and resolution plans, and maintain a knowledge base for the BI team.

Collaboration and engineering practices

  • Partner with BI analysts to understand dashboard and report requirements, translating business logic into transformation code.
  • Collaborate with software engineers to prepare training datasets, build feature pipelines, and help ensure data quality for forecasting and machine learning models.
  • Work with the Data Platform Architect to follow architectural patterns, coding standards, and adopt platform capabilities such as data cataloging, monitoring frameworks, and CI/CD pipelines.
  • Support the BI team with data questions, query optimization, and troubleshooting, including guidance on efficient dataset querying.
  • Coordinate with the CDAIO team on source system integrations, data contracts, and ingestion layer requirements.
  • Maintain comprehensive documentation covering business logic, transformation steps, data lineage, dependencies, refresh schedules, and SLAs.
  • Create data dictionaries for datasets, including column definitions, data types, expected values, refresh frequency, and usage examples.
  • Follow software engineering best practices including Git-based version control, code review, automated testing, and CI/CD integration.
  • Contribute reusable SQL and Python utilities, templates, and patterns that accelerate pipeline development across the team.

Minimum qualifications

  • Bachelor’s degree in Computer Science, Information Systems, or a related field (or a high school diploma/GED with 4 years of relevant data engineering experience).
  • Minimum 5 years of hands-on experience building data pipelines and ETL/ELT processes in production environments.
  • Expert-level SQL skills, including complex joins, window functions, CTEs, aggregations, and query optimization for large datasets.
  • Strong Python skills and familiarity with the PySpark DataFrame API, including transformations, actions, and optimization techniques.
  • Proven experience building ETL/ELT pipelines on cloud data platforms such as Databricks, Snowflake, AWS Glue, or similar.
  • Understanding of dimensional modeling, slowly-changing dimensions, aggregate tables, and analytics-optimized data structures.
  • Experience implementing automated data validation, schema checks, and data quality frameworks.
  • Familiarity with cloud data services, including compute optimization and cost management.
  • Experience with Git workflows, code review practices, and automated testing for data pipelines.

Technologies

  • SQL, Python, PySpark, Databricks, Snowflake, AWS Glue, Git, CI/CD
  • Medallion architecture, ETL, ELT

Benefits

  • Healthcare benefits including medical, dental, vision, and prescription drug coverage.
  • Access to a Health Coach from GE Aerospace.
  • Employee Assistance Program with 24/7 confidential assessment, counseling, and referral services.
  • GE Aerospace Retirement Savings Plan (401(k)) with company matching contributions and company retirement contributions.
  • Access to Fidelity resources and planning consultants.
  • Tuition assistance.
  • Adoption assistance.
  • Paid parental leave.
  • Disability insurance and life insurance.
  • Paid time-off for vacation or illness.

Additional notes

  • Remote eligibility: fully remote arrangements across the United States; in-person attendance required for New Hire Orientation on Day 1.
  • Relocation assistance: not provided.
  • Posting close date: expected to close on Friday October 2, 2026.

Similar Jobs