EngineerJobs.io
← Back to all jobs

Job Description

Join Okta’s Resilience Team within the Workforce Identity Cloud (WIC) and help build a cloud-native Identity as a Service platform designed for reliability at scale. This Staff role is focused on leading technical initiatives, owning dependable backend systems, and partnering across teams to address complex reliability challenges in a high-availability, multi-tenant SaaS environment. The position is hybrid in San Francisco, CA, with an annual salary range of USD 194,000 - 243,000.

What you’ll work on

  • Partner with engineering teams to design, develop, and deliver cloud-based infrastructure projects that resolve complex reliability issues and enable Okta’s services to scale.
  • Design and implement backend frameworks, tools, microservices, and library components that the broader engineering organization can use at scale.
  • Lead and advise on fault-tolerant system design, championing resilient, scalable solutions across engineering.
  • Explore and integrate AI-driven developer tooling, automated diagnostics, and predictive monitoring to streamline performance testing, accelerate debugging, and improve service reliability.
  • Perform rigorous design and code reviews, raising the engineering bar with high-quality unit and functional tests to support robust programming standards.
  • Participate in and lead incident Root Cause Analysis (RCA) processes to identify systemic weaknesses and build long-term mitigation strategies.
  • Collaborate with Architects, QA, Product Owners, Engineering Services, and Tech Ops to deliver secure, dependable software.
  • Support team growth by mentoring new engineering hires and interns.

Required experience and skills

  • 8+ years of experience as a software developer with deep, expert-level knowledge of Java (Golang experience is a strong plus).
  • Demonstrated ability to architect, implement, tune, and debug global cloud software.
  • Deep understanding of microservices, cloud infrastructure ( AWS preferred, Azure, or GCP), and container ecosystems (Docker, Kubernetes).
  • Proven track record addressing scaling, redundancy, and performance bottlenecks in high-availability, multi-tenant SaaS environments.
  • Practical experience or strong interest in leveraging generative AI utilities and developer assistance tools (examples listed: AI coding assistants, automated code-generation, intelligent log analytics, or predictive profiling tools ) to boost development productivity and optimize systems.
  • Outstanding communication and leadership skills, including coordinating high-impact technical projects, managing stakeholder alignment, and influencing technical direction.
  • Previous experience in a dedicated platform reliability, resilience, or core infrastructure role is highly valued.
  • Eligibility requirement: the employee must be either a U.S. citizen (42 U.S. Code § 9102); a lawful permanent resident (8 U.S.C. 1101(a)(20)) or a protected individual (8 U.S.C. 1324b(a)(3)); or a U.S. national (including citizens and non-citizens born in outlying possessions such as American Samoa and Swains Island). Working on a U.S. visa does not qualify as a U.S. person.

Education

B.S. or M.S. in Computer Science, a related field, or equivalent industry experience.

Technologies

  • Java, Golang
  • AWS, Azure, GCP
  • Docker, Kubernetes
  • Microservices
  • AI coding assistants, automated code-generation, intelligent log analytics, predictive profiling tools

Similar Jobs