Software Engineer III, Diags Infrastructure
Job Description
Build production and manufacturing diagnostic software that improves hardware uptime by reducing Mean Time To Repair (MTTR) and increasing useful hardware coverage.
Responsibilities
- Design, build, and scale next-generation diagnostic infrastructure for production and manufacturing environments
- Contribute to lowering MTTR and maximizing useful hardware coverage across global fleets
- Develop foundational software primitives and API capabilities to model diagnostic capabilities and validate test constraints
- Engineer diagnostic infrastructure, qualification frameworks, and dynamic distribution systems as part of the Diags Infrastructure team
- Create systems that describe, qualify, and distribute diagnostics across both test and production environments
- Drive automation for the repair ecosystem by delivering coverage-based recommendations to reduce operational intervention
- Support the execution of diagnostics by developing and providing critical infrastructure
Requirements
- Bachelor’s degree or equivalent practical experience
- 2 years of experience building and developing large-scale infrastructure and distributed systems
- 2 years of experience in programming
- 2 years of experience testing and launching software products
- 1 year of experience with distributed computing and large-scale data processing
- Master’s degree or PhD in Computer Science or related technical field
- 2 years of experience with data center architecture
- 2 years of experience with tools development
- 2 years of experience with monitoring systems
- 2 years of experience coding in Python and C++
Technologies
- Python
- C++
Location and Compensation
- Location: Sunnyvale, CA (onsite)
- Salary: USD 147,000 - 210,000 per year
- Compensation details: US: $147000 - $210000 (USD) + 15% bonus target + equity + benefits
About the team: Google’s Diags Infrastructure mission is to engineer diagnostic infrastructure, qualification frameworks, and dynamic distribution systems to minimize MTTR and maximize useful hardware coverage. The team uses coverage-based recommendations to automate parts of the repair ecosystem across both test and production environments.