Systems Software Engineer
Job Description
Systems Software Engineer (MTS3) role focused on upgrade automation and guided self-service workflows for a next-generation hardware upgrade framework.
Responsibilities
- Architect, implement, and maintain high-reliability systems software components for controller-upgrade automation and the next-generation hardware upgrade framework.
- Convert complex hardware replacement and upgrade tasks into repeatable, self-service workflows to reduce support and field dependencies.
- Design robust state machines, health checks, workflow orchestration, and pre/post-upgrade validations.
- Implement failure-handling mechanisms with automated recovery paths for upgrade execution.
- Collaborate with internal services, Pure1/Skyline cloud teams, and global CPBU engineering units to deliver an end-to-end guided experience.
- Develop comprehensive test strategies, regression coverage, and CI/CD automation (for example, Jenkins) across varied hardware and software combinations.
- Define actionable operational metrics for upgrade execution time, manual intervention, error modes, and telemetry.
Requirements
- Hands-on development and coding experience in C++ and Python.
- Strong foundation in distributed systems, high availability (HA), storage architectures, hardware lifecycle, or embedded/platform software.
- Proven experience designing state machines, workflow orchestration, health checks, error handling, and operator guidance for critical systems.
- Exceptional debugging capability across multi-layer boundaries spanning software, hardware, and distributed services in reliability-sensitive environments.
- Experience building test strategies and automated validation across hardware/software combinations.
- Working knowledge of CI/CD tooling such as Jenkins or GitLab.
- Strong written and verbal communication skills, with a record of driving cross-functional alignment across technical partners.
- Ability to work on-site in Santa Clara, CA to engage directly with hardware and lab environments.
- Plus: familiarity with storage protocols/architecture (NVMe, iSCSI, PCIe, NUMA), HA failover design, array controller replacement dynamics, or server/CPU hardware architecture.
Key Technologies
- C++, Python
- Jenkins, GitLab
- Pure1, Skyline
- NVMe, iSCSI, PCIe, NUMA
Compensation
- USD 149,000 - 224,000 per year
Benefits
- Flexible time off
- Wellness resources
- Company-sponsored team events
Accommodation & Accessibility
- Candidates with disabilities may request accommodations for all aspects of the hiring process.
- For more information, contact: [email protected]
Inclusive Team Commitment
- Committed to fostering growth and development of every person through community-building and Employee Resource Groups.
- Advocates for inclusive leadership.
- Equal opportunity employer; does not discriminate based on legally protected characteristics.
What You’ll Do (Summary)
- Design & Build Upgrade Infrastructure: architect, implement, and maintain high-reliability systems software components for controller-upgrade automation and the hardware upgrade framework.
- Automate Upgrade Workflows: transition hardware replacement and upgrades into repeatable self-service workflows to reduce support and field dependencies.
- Engineered Safety & Resilience: build state machines, health checks, orchestration, validations, failure handling, and automated recovery paths.
- Cross-Functional Collaboration: work with internal services, Pure1/Skyline cloud teams, and CPBU engineering units for an end-to-end guided experience.
- Testing & Pipeline Validation: develop test strategies, regression coverage, and CI/CD automation (for example, Jenkins) across diverse hardware/software combinations.
- Observability & Telemetry: establish operational metrics for execution time, manual intervention, error modes, and telemetry.
Similar Jobs
U