Staff AI Engineer - Global Infrastructure
Backend Developer
Manager
Agentic Ai
Ai Agent
Ai Agent Platform
Application Security
Artificial Intelligence
Artificial Intelligence Engineer
Automation
Cloud
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Cloud Technology
Data Architecture
Data Platform
DevOps
DevSecOps
Engineering
Engineering Software
Facilities Management
Gen Ai Platform
Generative AI
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning
Machine Learning Engineering
Platform Engineering
Programming
Security Automation
Staff Ai Engineer
Job Description
American Express seeks a Staff AI Engineer to lead enterprise-scale AI architecture and implementation across infrastructure and agentic AI systems. This hybrid role in New York focuses on scalable, reliable, secure, and responsible AI platforms while guiding technical direction, platform risk, and continuous improvements in AI engineering practices.
Role Responsibilities
- Lead enterprise-scale AI architecture and implementation spanning infrastructure and agentic AI systems.
- Set technical direction for core enterprise AI platform initiatives.
- Ensure AI platforms are scalable, reliable, secure, and responsible.
- Mentor senior technical talent and support engineering excellence across teams.
- Drive continuous improvement in AI engineering practices and platform capabilities.
- Provide 24x7 support to help maintain an uninterrupted, high-quality customer and colleague experience.
- Lead technology risk and information security, enterprise data governance and platforms, digital product and design, and enterprise AI platforms on behalf of the company.
- Provide product management for core enterprise platforms.
Required Qualifications
- Deep knowledge of machine learning and deep learning systems, including model architectures, training, evaluation, and optimization.
- Advanced understanding of Generative AI and LLM ecosystems, including embeddings, fine-tuning, prompt design, retrieval-augmented generation, and inference at scale.
- Strong understanding of AI infrastructure, including GPU platforms, accelerated compute, workload scheduling, capacity optimization, model serving, and inference performance.
- Advanced understanding of agentic AI design, including planning, reasoning, tool use, memory, multi-agent coordination, and autonomy controls.
- Strong foundation in distributed systems and cloud-native architecture, including Kubernetes, APIs, microservices, event-driven design, observability, and platform reliability.
- Knowledge of DevOps and CI/CD platforms such as Harness, including automated deployment, environment promotion, governance, and operational controls.
- Knowledge of enterprise AI governance, including model risk management, explainability, bias detection, safety, and regulatory compliance.
- MS/PhD in Artificial Intelligence, Machine Learning, Computer Science, or related discipline (preferred).
- 12+ years of experience in AI/ML engineering, AI infrastructure, platform engineering, data engineering, or related fields, with a track record of delivering complex production systems at scale.
- Proven experience architecting end-to-end AI/ML platforms across data pipelines, training, deployment, serving, monitoring, and inference optimization.
- Experience designing and operating GPU-based AI infrastructure, including accelerated compute platforms, workload scheduling, utilization optimization, capacity management, and performance tuning.
- Experience with cloud-native and Kubernetes-based platforms supporting training, batch workloads, inference, orchestration, observability, reliability, and cost efficiency.
- Hands-on experience with Python and modern AI/ML frameworks including PyTorch, TensorFlow, Scikit-learn, Hugging Face, and related tooling.
- Deep experience with agentic AI systems, including planning, tool use, memory, evaluation, retrieval, embeddings, vector databases, and agent frameworks.
- Experience leading complex cross-functional technical initiatives and influencing architecture and engineering direction without direct authority.
- Demonstrated ability to mentor and develop engineers at all levels, including Staff-level engineers.
- Experience working in regulated environments such as financial services, including AI governance, risk management, and compliance considerations.
Technologies
- Python; PyTorch; TensorFlow; Scikit-learn; Hugging Face
- Kubernetes; Harness; CI/CD
- LLM ecosystems; embeddings; fine-tuning; prompt design; retrieval-augmented generation
- GPU platforms; accelerated compute; model serving
- Vector databases; agent frameworks
- Distributed systems; cloud-native architecture; APIs; microservices; event-driven design; observability
Compensation and Location
Location: New York, NY (hybrid). Salary: USD 144,250 - 256,250 per year.
Benefits
- Competitive base salaries
- Bonus incentives
- 6% Company Match on retirement savings plan
- Free financial coaching and financial well-being support
- Comprehensive medical, dental, vision, life insurance, and disability benefits
- Flexible working model with hybrid, onsite or virtual arrangements depending on role and business need
- 20+ weeks paid parental leave for all parents, regardless of gender, for pregnancy, adoption or surrogacy
- Free access to global on-site wellness centers staffed with nurses and doctors (depending on location)
- Free and confidential counseling support through Healthy Minds program
- Career development and training opportunities