EngineerJobs.io
← Back to all jobs

Job Description

Hewlett Packard Enterprise is building HPE AI Essentials, an inference platform designed to run large language models on customer-owned infrastructure, including air-gapped and sovereign environments. In this hybrid role in Spring, TX, you will help lead the model runtime and the Kubernetes orchestration approach that supports efficient, low-latency, high-utilization inference.

What you will do

  • Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
  • Collaborate with inference performance engineering teams, with responsibility for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
  • Set the distributed inferencing strategy, including disaggregated prefill and decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Assess emerging runtimes and methods, including quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether to adopt, develop in-house, or decline each approach.
  • Define the orchestration layer for the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
  • Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences.

Required qualifications

  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including the ability to modify engine internals.
  • Deep understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Knowledge of tensor and pipeline parallelism, NCCL collective operations, and how GPU memory hierarchy and interconnect characteristics affect performance.
  • Expert proficiency with Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling.
  • Strong programming ability in Go and Python, plus the ability to read, debug, and profile C++/CUDA using tools such as Nsight.
  • Experience debugging and profiling multi-tier application workloads such as RAG and Agents.
  • Excellent analytical, debugging, and problem-solving skills.

Technologies involved

  • vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM
  • Go, Python, C++, CUDA, Nsight
  • Kubernetes, NCCL, RDMA, GPUDirect Storage, InfiniBand, RoCE
  • RAG, Agents

Education and experience

  • Minimum experience: 12 years
  • Education: Degree in Computer Science or a related field

Benefits

  • A comprehensive benefits suite intended to support physical, financial, and emotional wellbeing.
  • Programs designed to support career goals, whether you want to deepen expertise or apply your skills elsewhere within HPE.
  • Unconditional inclusion, with flexibility to manage work and personal needs.

Preferred background

  • Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe.
  • Experience with disaggregated prefill/decode serving, or KV cache offload and reuse at scale.
  • Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
  • Experience delivering on-premises, air-gapped, or regulated enterprise software.

Location and work model

  • Location: Spring, TX (hybrid)
  • The role is designed as hybrid, with an expectation to work on average 2 days per week from an HPE office.
  • Primary work location is Spring, TX, with the possibility of other HPE sites in the US; remote work options may be considered.

Compensation

  • Annual base salary range: USD 160,000 - 303,000 (listed)
  • Annual salary range also varies by location: USD 160,000 - 303,000 in Colorado; USD 152,000 - 349,000 in North Carolina & Texas
  • Variable incentives may also be offered.

Similar Jobs