EngineerJobs.io
← Back to all jobs

Job Description

NVIDIA is expanding its Kubernetes-native topology platform to gather, normalize, and expose topology data for provisioning systems and schedulers across multiple cloud providers. This Senior Software Engineer role focuses on interfacing with NVIDIA hardware to optimize GPU-to-GPU scheduling and contributes to the open-source Topograph project. The position is based in Santa Clara, CA, with remote work options, and offers an annual salary range of USD 184,000 - 287,500.

Responsibilities

  • Build a system that collects topology information from multiple sources
  • Aggregate and normalize data so it is available to provisioning systems and workload schedulers
  • Direct contributor in the critical open-source project Topograph
  • Collaborate with hardware teams to ensure new product launches have efficient scheduling capabilities

Requirements

  • Eight or more years of relevant experience
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical field, or equivalent experience
  • Strong production engineering experience in Go or another systems language
  • Experience with distributed systems, Kubernetes, Slurm/Slinky, Linux, containers, APIs, and CI
  • Ability to design clean interfaces between discovery logic, data models, and scheduler output
  • Familiarity with networking, cluster topology, cloud infrastructure, or large-scale compute systems
  • Excellent testing, debugging, documentation, and code review habits

Technologies

  • Go
  • Kubernetes
  • Slurm
  • Slinky
  • Kueue
  • DRA
  • Topograph

Benefits

  • Equity
  • Benefits

Ways to Stand Out From the Crowd

  • Experience with GPU clusters, NVLink, InfiniBand, Ethernet fabrics, or HPC
  • Hands-on work with Kubernetes scheduling, Slurm/Slinky topology, DRA, Kueue, or device plugins
  • Experience integrating with cloud provider topology APIs or cluster metadata systems

Similar Jobs