J
Software Engineer III - Senior Databricks/Spark/AWS Data Engineer
Artificial Intelligence
AWS
Aws Cloudwatch
Big Data
Bigdata
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Operations
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Lake
Data Lakehouse
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
Databricks
Databricks Pyspark
Databricks Workflows
Delta Live Tables
ETL
Informatica
Information Technology (IT)
Programming
Programming Languages
S3
Spark
SQL
Job Description
Lead secure, scalable data engineering efforts on the Databricks-on-AWS lakehouse, delivering production pipelines that support workforce analytics and BI reporting.
Responsibilities
- Design, build, and maintain new Databricks data pipelines using PySpark, focusing on secure, high-quality production code
- Review and debug data engineering processes implemented by others to improve reliability and delivery
- Optimize and tune PySpark jobs and Databricks clusters for performance, scalability, and cost efficiency (including partitioning, caching, and resource management)
- Build scalable data frameworks for end-to-end Databricks pipelines using medallion lakehouse patterns (bronze/silver/gold) for workforce data analytics
- Implement data quality checks and validation using Delta Lake and Delta Live Tables expectations to ensure accuracy and reliability
- Set up robust monitoring and alerting to detect and address data ingestion issues early, using Databricks and AWS CloudWatch to improve performance and throughput
- Identify recurring pipeline issues and opportunities to eliminate or automate remediation using Databricks Workflows and AWS-native automation
- Use AI and agentic AI solutions to accelerate pipeline development and apply AI-assisted engineering tools (e.g., Claude, GitHub Copilot) to improve productivity and code quality
- Provision and deliver curated, reliable datasets to BI partners using Sigma, Tableau, and Alteryx
- Collaborate with business stakeholders to understand requirements, then create architecture and design artifacts for complex applications
- Participate in software engineering communities of practice exploring new and emerging technologies, supporting a culture of diversity, opportunity, inclusion, and respect
Requirements
- 3+ years of applied experience in data engineering with formal training or certification in software engineering concepts, including design, application development, testing, and operational stability
- Advanced hands-on expertise in Apache Spark (PySpark) for large-scale distributed processing
- Strong proficiency building and operating production pipelines on Databricks using Delta Lake and lakehouse patterns
- Strong expertise across AWS data ecosystem including S3, EMR, Glue, Lambda, and Athena, plus AWS storage and compute services
- Experience with data formats including Parquet and Iceberg
- Strong Python skills for data processing and application development (with Java or Scala as a plus)
- Experience with automation and continuous delivery using CI/CD pipelines and tooling such as Git/Bitbucket, Jenkins, or Spinnaker
- Hands-on experience across system design, application development, testing, and operational stability, with advanced understanding of agile methodologies, application resiliency, and security
- In-depth knowledge of the financial services industry and its IT systems
- SQL and data modeling skills for efficient data management and retrieval (experience with Oracle is a plus)
- Practical experience scheduling and automating job execution using Airflow and Autosys
Preferred Qualifications / Capabilities & Skills
- Databricks certifications such as Databricks Certified Data Engineer Associate/Professional
- Familiarity with generative AI and agentic AI frameworks, including AI coding assistants like Claude and GitHub Copilot
- Deeper expertise in the AWS cloud platform and its broader service catalog
Technologies
- Databricks, Apache Spark, PySpark, AWS
- Delta Lake, Delta Live Tables, Delta Live Tables expectations
- Databricks Workflows, AWS CloudWatch
- AI, Agentic AI, Claude, GitHub Copilot
- Sigma, Tableau, Alteryx
- Python, Java, Scala, Git, Bitbucket, Jenkins, Spinnaker
- S3, EMR, Glue, Lambda, Athena
- Parquet, Iceberg, SQL, Oracle
- Airflow, Autosys
Benefits
- Competitive total rewards package including base salary determined based on role, experience, skill set, and location
- Eligible roles may receive commission-based pay and/or discretionary incentive compensation in cash and/or forfeitable equity
- Comprehensive health care coverage
- On-site health and wellness centers
- A retirement savings plan
- Backup childcare
- Tuition reimbursement
- Mental health support
- Financial coaching
- Additional details about total compensation and benefits provided during the hiring process
- Equal opportunity employer; values diversity and inclusion
Location
- Columbus, OH (onsite)
About the Team
- Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing
- Corporate teams help set businesses, clients, customers, and employees up for success
Similar Jobs
J
J
J
J