Principal Software Engineer, Inference
Job Description
Hewlett Packard Enterprise is building HPE AI Essentials, an inference platform designed to run large language models on customer-owned infrastructure, including air-gapped and sovereign environments. In this hybrid role in Spring, TX, you will help lead the model runtime and the Kubernetes orchestration approach that supports efficient, low-latency, high-utilization inference.
What you will do
- Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
- Collaborate with inference performance engineering teams, with responsibility for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
- Set the distributed inferencing strategy, including disaggregated prefill and decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.
- Assess emerging runtimes and methods, including quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether to adopt, develop in-house, or decline each approach.
- Define the orchestration layer for the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
- Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences.
Required qualifications
- Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including the ability to modify engine internals.
- Deep understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
- Knowledge of tensor and pipeline parallelism, NCCL collective operations, and how GPU memory hierarchy and interconnect characteristics affect performance.
- Expert proficiency with Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling.
- Strong programming ability in Go and Python, plus the ability to read, debug, and profile C++/CUDA using tools such as Nsight.
- Experience debugging and profiling multi-tier application workloads such as RAG and Agents.
- Excellent analytical, debugging, and problem-solving skills.
Technologies involved
- vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM
- Go, Python, C++, CUDA, Nsight
- Kubernetes, NCCL, RDMA, GPUDirect Storage, InfiniBand, RoCE
- RAG, Agents
Education and experience
- Minimum experience: 12 years
- Education: Degree in Computer Science or a related field
Benefits
- A comprehensive benefits suite intended to support physical, financial, and emotional wellbeing.
- Programs designed to support career goals, whether you want to deepen expertise or apply your skills elsewhere within HPE.
- Unconditional inclusion, with flexibility to manage work and personal needs.
Preferred background
- Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe.
- Experience with disaggregated prefill/decode serving, or KV cache offload and reuse at scale.
- Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
- Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
- Experience delivering on-premises, air-gapped, or regulated enterprise software.
Location and work model
- Location: Spring, TX (hybrid)
- The role is designed as hybrid, with an expectation to work on average 2 days per week from an HPE office.
- Primary work location is Spring, TX, with the possibility of other HPE sites in the US; remote work options may be considered.
Compensation
- Annual base salary range: USD 160,000 - 303,000 (listed)
- Annual salary range also varies by location: USD 160,000 - 303,000 in Colorado; USD 152,000 - 349,000 in North Carolina & Texas
- Variable incentives may also be offered.