EngineerJobs.io
← Back to all jobs

Job Description

American Express is hiring a Senior AI Engineer (onsite) in Sunrise, FL to design and build foundational Agentic AI Platform capabilities that support safe, reliable agent operation across the enterprise.

Responsibilities

  • Contribute to the architecture and implementation of capabilities across the Agentic AI Platform, including Agent Runtime & Execution
  • Design and build scalable agent runtime infrastructure using Kubernetes and cloud-native technologies
  • Develop runtime abstractions enabling agents to execute across internally managed sandboxed environments and third-party AI platforms
  • Build orchestration capabilities for tool execution, state and context management, memory, asynchronous workloads, event-driven execution, and multi-agent workflows
  • Create a unified control plane to manage agents across heterogeneous execution environments and AI providers
  • Develop APIs and services for agent registration, configuration, deployment, versioning, lifecycle management, policy enforcement, and runtime management
  • Build abstractions that reduce provider-specific complexity while preserving access to differentiated capabilities across AI ecosystems
  • Engineer platform-level controls for enterprise AI governance, including:
    • Agent identity, authentication and authorization
    • Tool permissions
    • Policy enforcement
    • Data boundaries
    • Auditability
    • Lifecycle controls
  • Build mechanisms for governing models, prompts, tools, MCP servers, knowledge sources, agent-to-agent interactions, and external integrations
  • Partner with security, risk, privacy, and governance teams to translate enterprise requirements into scalable technical controls
  • Build registries and catalogs so developers and AI systems can discover reusable agents, tools, skills, prompts, models, knowledge sources, and other platform capabilities
  • Develop metadata systems for ownership, versioning, dependencies, certification, and discovery to support a healthy enterprise agent ecosystem
  • Build end-to-end telemetry for agent execution, including traces, events, model interactions, tool calls, latency, token consumption, cost, failures, policy decisions, and quality signals
  • Develop capabilities to make complex and multi-agent workflows explainable and debuggable
  • Enable platform and application teams to understand agent behavior across multiple models, runtimes, tools, and external systems
  • Develop evaluation frameworks to measure quality, reliability, safety, and task performance
  • Build automated evaluation pipelines including offline evaluations, production signals, human feedback, regression testing, and experimentation
  • Create continuous learning loops that convert production telemetry and feedback into improvements to agents, prompts, tools, models, and platform capabilities
  • Build SDKs, APIs, CLIs, templates, local development environments, and self-service workflows for productive agent development
  • Create opinionated paved roads that incorporate enterprise security, governance, observability, and operational standards automatically
  • Build CI/CD capabilities for AI agents, including automated evaluation, policy validation, security checks, artifact/version management, deployment, promotion, rollback, and release controls
  • Design GitOps and infrastructure-as-code patterns for deploying and managing agent workloads
  • Help establish engineering standards for taking agents from experimentation to production safely and repeatedly

Requirements

  • Strong software engineering experience building production systems using Python, Java, Go, or TypeScript
  • Strong understanding of Generative AI, LLMs, agent architectures, tool/function calling, retrieval, context management, and agent orchestration
  • Experience building production applications or platforms using major model providers or AI platforms
  • Experience with Kubernetes, containers, microservices, distributed systems, and cloud-native architecture
  • Experience designing production APIs and event-driven or asynchronous systems
  • Strong understanding of modern cloud infrastructure and infrastructure-as-code practices
  • Experience with CI/CD, automated testing, production observability, and software delivery practices
  • Strong understanding of security fundamentals: identity, authentication, authorization, secrets, and least-privilege access
  • Ability to work through ambiguous technical problems and turn emerging technologies into reliable production systems
  • Strong communication skills and collaboration across engineering, architecture, product, security, and governance organizations

Technologies

  • Python, Java, Go, TypeScript
  • Generative AI, LLMs
  • Agent architectures
  • Tool/function calling, retrieval, context management, agent orchestration
  • Kubernetes, cloud-native technologies, microservices, distributed systems
  • Event-driven systems, asynchronous systems
  • CI/CD, infrastructure-as-code, GitOps
  • OpenTelemetry
  • LangGraph, Semantic Kernel
  • Google ADK, OpenAI Agents SDK
  • Model Context Protocol (MCP), A2A
  • Vector databases, embeddings
  • LLM-as-judge techniques
  • Policy-as-code

Preferred Qualifications

  • Experience with agent frameworks and orchestration technologies such as LangGraph, Semantic Kernel, Google ADK, OpenAI Agents SDK, or similar frameworks
  • Model Context Protocol (MCP) and emerging agent interoperability protocols like A2A
  • Multi-agent systems and agent-to-agent communication patterns
  • Kubernetes operators, controllers, service meshes, or sophisticated Kubernetes platform engineering
  • Agent/model gateways and intelligent model routing
  • Vector databases, retrieval systems, embeddings, and enterprise knowledge architectures
  • LLM and agent evaluation frameworks, LLM-as-judge techniques, human feedback systems, experimentation, and quality measurement
  • AI observability and distributed tracing, including OpenTelemetry, Langfuse, and production monitoring
  • Experience with Evals
  • Policy engines and policy-as-code
  • AI security: prompt injection defenses, tool security, data-loss prevention, and AI-specific threat modeling
  • Platform engineering, internal developer platforms, developer portals, and enterprise service catalogs
  • Large-scale distributed systems operating in highly regulated environments

What Success Looks Like

  • Enable developers across American Express to move from an agent idea to a governed production deployment through a consistent, self-service platform
  • Make the secure and reliable path the easiest path, so teams can build once and operate agents across multiple AI ecosystems with governance, identity, observability, evaluation, deployment, and operational capabilities by default
  • Establish engineering foundations for an enterprise agent ecosystem where agents can be built, discovered, trusted, governed, observed, evaluated, deployed, and continuously improved at scale
  • Depending on business unit requirements, role details, cost, and applicable laws, American Express may provide visa sponsorship for certain positions

Similar Jobs