Build the private cloud AI foundation behind HPE AI Essentials. In the Private Cloud AI organization, this hybrid Senior Software Engineer role focuses on evolving the model runtime for inference with goals like low tail latency and high GPU utilization, including support for distributed execution in enterprise LLM serving scenarios such as air-gapped and sovereign environments.
What you’ll work on
You will design, implement, and take ownership of key runtime and serving components used to deploy large language model workloads. The role spans engine integration and performance-critical inference optimizations, then extends into distributed execution and Kubernetes orchestration.
- Design, implement, and own major components of LLM serving deployments, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
- Partner with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
- Build and operate distributed execution, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
- Evaluate emerging runtimes and serving approaches such as quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption based on evidence
- Contribute to orchestration capabilities for the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
- Triage and resolve customer issues end-to-end by identifying root causes and improving systems and processes to prevent recurrence
- Deliver impactful code and design reviews, mentor team members, and lead by example on engineering practices
Core requirements
- Experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including ability to modify engine internals
- Strong understanding of inference internals: continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
- Knowledge of tensor and pipeline parallelism, NCCL collectives, and how GPU memory hierarchy and interconnect characteristics affect performance
- Advanced proficiency with Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
- Programming proficiency in Go and Python, plus ability to read, debug, and profile C++/CUDA using Nsight
- Familiarity with debugging and profiling multi-tier workloads such as RAG and Agents
- Excellent analytical, debugging, and problem-solving skills
Technologies you’ll use
vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM, Go, Python, C++, CUDA, Nsight, Kubernetes, NCCL
Preferred background
- Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe
- Experience with disaggregated prefill/decode or KV cache offload and reuse at scale
- Exposure to RDMA, GPUDirect Storage, InfiniBand, or RoCE
- Knowledge of MIG, fractional GPU allocation, and multi-tenant GPU isolation
- Experience delivering on-premises, air-gapped, or regulated enterprise software
Salary, location, and work model
- Location: Spring, TX (Hybrid)
- Hybrid expectation: work on average 2 days per week from an HPE office
- Remote: remote work options may be considered
- Salary (base): USD 144,000 - 273,000 per year
- Salary range by state: USD 144,000 - 273,000 in Colorado; 137,000 - 315,000 in North Carolina & Texas
- Incentives: variable incentives may also be offered
Health, inclusion, and growth
HPE offers a comprehensive benefits suite designed to support your physical, financial, and emotional wellbeing. The company also invests in professional development with programs aligned to your career goals. HPE is unconditionally inclusive, celebrates individual uniqueness, provides flexibility to manage work and personal needs, and emphasizes doing bold moves together.
Additional details
- Minimum experience: 8 years
- Education: Degree in Computer Science or related field
- Job application period closure (estimated): December 30, 2027
- Recruitment fraud alert: HPE and authorized agencies/vendors will never charge registration, hiring, or other fees, and will never request sensitive personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications. Verify any third party claiming to represent the company through official channels.
Stay connected: Follow @HPECareers on Instagram to see the latest on people, culture, and tech at HPE.