EngineerJobs.io
← Back to all jobs

Job Description

Palo Alto Networks is seeking a Principal Machine Learning Engineer in Santa Clara, CA to provide technical leadership for designing and delivering robust, next-generation cloud security solutions. The role owns the machine learning lifecycle from development and training through production deployment and real-time inference, with an emphasis on scalable architecture, CI/CD, monitoring, and practical MLOps evaluation.

Responsibilities

  • Provide technical leadership for end-to-end solution delivery by partnering with cross-functional teams across Product, SRE, QA, and Support to align engineering execution with business objectives.
  • Lead the development of scalable cloud security architecture through a blend of strategic planning and hands-on coding.
  • Set and promote best practices for model versioning, reproducibility, auditing, and compliance to support code quality and data privacy across the organization.
  • Architect and drive the complete ML lifecycle, from initial development and training to production deployment and real-time inference.
  • Build and maintain automated, resilient systems for continuous integration and continuous delivery (CI/CD) as well as monitoring across backend and machine learning components.
  • Continuously assess and incorporate emerging MLOps tools and frameworks to improve scalability, reliability, and operational efficiency.
  • Design and implement next-generation cloud security solutions that address complex backend infrastructure and machine learning model challenges.
  • Strategically manage and optimize ML infrastructure and pipelines to improve performance, support production integration, and reduce operational costs.

Requirements

  • Strong background in machine learning and ML frameworks, including TensorFlow and PyTorch.
  • Experience with Infrastructure-as-Code (IaC) tools such as Terraform or CloudFormation.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 10+ years of software development experience with a focus on cloud-native and SaaS applications.
  • Proven experience designing and building large-scale distributed systems on public cloud platforms including AWS, GCP, or Azure.
  • Strong proficiency in at least one modern programming language such as Python, Go, or Java.
  • Demonstrated experience across the full machine learning lifecycle, including model deployment and MLOps.

Required Technologies

  • TensorFlow, PyTorch
  • Terraform, CloudFormation
  • Python, Go, Java
  • AWS, GCP, Azure
  • CI/CD, MLOps
  • Docker, Kubernetes
  • Kafka, Flink

Team

The engineering team is central to the organization’s products and connected to the mission of preventing cyberattacks. The team focuses on ongoing innovation in cybersecurity, building products to solve problems that may not have been addressed elsewhere, and defining the industry through proactive engineering decisions. Collaboration is conducted in person, with most teams working from the office full time and flexibility as needed.

Preferred Qualifications

  • Master’s or PhD in Computer Science or a related technical field.
  • Experience in the cybersecurity domain or with network security products.
  • Expertise with containerization and orchestration, particularly Docker and Kubernetes.
  • Experience with real-time data processing and streaming technologies such as Kafka and Flink.
  • Contributions to open-source projects in cloud-native or MLOps.

Location and Compensation

  • Location: Santa Clara, CA, United States (onsite)
  • Salary: USD 163,200 - 264,000 per year

Similar Jobs