This position is no longer accepting applications
Closed on September 5, 2026.
This role is filled — get an email when new Engineering roles open on EngineerJobs.io:
AI Engineer
Artificial Intelligence
Deep Learning
Engineering
Generative AI
Large Language Models
Machine Learning
Machine Learning Engineer
Ml Ops
PyTorch
Rag Architectures
View similar jobs
Get alerted when similar jobs are posted — set up a New Engineering jobs on EngineerJobs.io alert.
See other roles at Entertainment Partners.
Job Description
Entertainment Partners is building production-grade AI and agentic systems that power an intelligent product suite, with a focus on applied ML engineering, LLM product development, and production system design. In this hybrid role based in Tempe, AZ, you will help design, train, evaluate, and deploy AI/ML capabilities that integrate with EP’s enterprise data and product APIs.
What You’ll Do
- Design, develop, train, fine-tune, and evaluate machine learning models using PyTorch and related libraries such as torchvision, torchaudio, torch.nn, and torch.optim.
- Build and maintain ML training pipelines, experiment tracking workflows, and model evaluation frameworks.
- Implement transformer-based models and production LLM integrations for use cases including NLP, information extraction, classification, and generation.
- Use parameter-efficient fine-tuning approaches such as LoRA, QLoRA, and PEFT to adapt foundation models to EP domains including payroll, residuals, and production management.
- Design and implement RAG architectures using vector databases and semantic search pipelines with tools including pgvector, Pinecone, and Weaviate.
- Optimize inference for latency and throughput using quantization, batching, and caching strategies for production serving.
- Develop and maintain AI evaluation frameworks, including automated evals as unit tests, to support reliable, safe, production-grade model behavior.
- Build LLM-powered agentic workflows using LangChain and LangGraph, along with EP’s internal MCP (Model Context Protocol) server architecture.
- Create multi-step reasoning pipelines, tool-calling agents, and autonomous task execution systems that integrate with EP enterprise data and product APIs.
- Apply prompt engineering strategies such as few-shot templates, chain-of-thought scaffolding, and structured output validation.
- Apply and maintain EP AI quality engineering standards including failure taxonomy, runtime guardrails, and evidence-driven release gates.
- Contribute to EP’s Enterprise Context Engine, described as a governed, zero-data-retention AI context layer exposed via MCP to Tabnine Agent and Claude Code.
- Build and maintain MLOps infrastructure for training, experiment tracking with MLflow and Weights & Biases, model versioning, and deployment.
- Containerize and deploy ML services using Docker and Kubernetes, integrating with CI/CD pipelines such as GitHub Actions and Azure DevOps.
- Monitor model performance in production, including drift detection, feedback loops, and automated retraining triggers.
- Support security, privacy, and compliance requirements, including data minimization and access control for sensitive payroll data.
- Collaborate with data engineering on feature stores, data pipelines, and training data infrastructure.
- Work with the Chief Architect AI & Data and CAIO to define AI architecture patterns and best practices for the EP engineering organization.
- Partner with product managers, UX designers, and full stack engineers to translate AI capabilities into product features.
- Conduct code reviews for AI/ML code, emphasizing reproducibility, correctness, and production readiness.
- Mentor engineers on AI engineering fundamentals, LLM integration patterns, and responsible AI practices.
- Stay current with the AI/ML landscape and evaluate new models, frameworks, and techniques for potential application at EP.
- Contribute to EP’s PE AI Maturity Scorecard (S1–S3) by advancing organizational AI capability maturity.
- Represent EP AI engineering practices during Architecture Review Board discussions.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Statistics, Mathematics, or a related quantitative field.
- 6–10+ years of professional software engineering experience, including 3+ years focused on ML/AI engineering in production environments.
- Expert-level proficiency in Python and deep familiarity with the Python ML/AI ecosystem.
- Hands-on production experience with PyTorch, including model definition (nn.Module), custom training loops, autograd, GPU acceleration (CUDA), and model serialization (TorchScript, ONNX).
- Experience with Hugging Face Transformers, Datasets, and PEFT, including fine-tuning and adapting foundation models.
- Demonstrated experience building RAG pipelines, including chunking strategies, embedding models, vector store selection, and retrieval evaluation.
- Production experience integrating LLM APIs (OpenAI, Anthropic, and open-source via vLLM/Ollama) and building reliable prompt engineering systems.
- Experience with LangChain or LangGraph for multi-step agents and tool-calling workflows.
- Strong grounding in ML fundamentals, including learning paradigms, loss functions, regularization, evaluation metrics, and statistical validation.
- Experience with experiment tracking tools such as MLflow, Weights & Biases, or Comet, with a focus on reproducible ML workflows.
- Working knowledge of containerization (Docker) and cloud ML services including AWS SageMaker, Azure ML, or OCI Data Science.
- Experience with SQL and NoSQL databases, including designing data pipelines for ML training and inference.
Technologies
- Python, PyTorch, torchvision, torchaudio, torch.nn, torch.optim
- LoRA, QLoRA, PEFT, RAG, pgvector, Pinecone, Weaviate
- Quantization, MLflow, Weights & Biases, Docker, Kubernetes
- GitHub Actions, Azure DevOps, LangChain, LangGraph
- MCP (Model Context Protocol), Tabnine Agent, Claude Code
- TorchScript, ONNX, Hugging Face Transformers, Hugging Face Datasets, CUDA
- OpenAI, Anthropic, vLLM, Ollama, Comet
- AWS SageMaker, Azure ML, OCI Data Science, SQL, NoSQL, vector databases
- Triton Inference Server, TorchServe, Ray Serve, TensorFlow, JAX, OpenCV, Kubeflow, KFServing
- Enterprise Context Engine, MLOps, CI/CD
Compensation and Benefits
- Salary range: USD 140,000 - 180,000 per year
- Health, Dental, and Vision options
- 401(k) retirement savings plan with company match
- Paid holidays, vacation time, and sick time
- Participation in company equity plans
- Employee Assistance Program, mental health and wellness programs
- Training and development
- Annual bonus and merit reviews
Preferred Qualifications
- Experience with additional deep learning frameworks (TensorFlow, JAX) or framework interoperability (ONNX).
- Familiarity with computer vision (torchvision, OpenCV) or speech/audio processing (torchaudio).
- Experience with model compression techniques: quantization (INT8, FP16, BF16), pruning, distillation.
- Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar.
- Contributions to open-source ML projects or published research (papers, patents, or technical blog posts).
- Experience with responsible AI frameworks, bias evaluation, and AI governance practices.
- Familiarity with MCP server development for exposing tools to AI agents.
- Prior domain experience in payroll, fintech, media, or enterprise SaaS environments.
- Experience with Kubernetes-based ML workload orchestration (Kubeflow, KFServing, or similar).
- Hybrid work environment familiarity: Burbank, CA headquarters with flexible remote schedule.
- On-call availability as needed for production AI system incidents and model deployment events.
- Access to GPU-accelerated compute environments (cloud-based) for training workloads.
- Ability to sit for extended periods at a computer workstation, and to use a keyboard and mouse.
- Occasional participation in early-morning or evening sessions to coordinate with distributed teams or international partners.
Similar Jobs
U