EngineerJobs.io
← Back to all jobs

Job Description

This on-site role in Nashville, TN, is for a Senior Core Infrastructure Software Engineer at Oracle to design, build, and operate AI systems on Oracle Cloud Infrastructure (OCI) for large-scale environments.

Responsibilities

  • Contribute to the design and development of distributed systems that scale horizontally and vertically, leveraging distributed state management tools.
  • Design, implement, and deliver scalable agentic AI systems capable of reasoning, planning, tool use, workflow execution, multi-step orchestration, and safe human-in-the-loop escalation.
  • Ensure scalability, performance, availability, and durability for owned services and components.
  • Define, implement, and verify functional requirements and tests for features within an existing system.
  • Optimize services for high throughput, low latency, and large-scale cloud workloads.
  • Build systems that tolerate network unreliability, service disruptions, and partition scenarios while meeting SLOs.
  • Establish telemetry, KPIs, dashboards, and alerts to monitor health, performance, reliability, and customer impact.
  • Lead performance testing, load testing, fault injection, brownout testing, and other validation strategies to ensure resiliency.
  • Implement secure infrastructure controls for multi-tenant cloud environments, including access controls, encryption, and remediation of security gaps.
  • Develop automation, infrastructure as code, and deployment tooling to support safe patching, updates, rollbacks, and operational recovery.
  • Design automation scripts and tooling to troubleshoot operational issues.
  • Own production operations, including troubleshooting, incident response, root cause analysis, and ongoing service improvements.
  • Adhere to change management plans for patching, updating, and rolling back applications.
  • Strengthen operational readiness by improving runbooks, monitoring, change management, deployment safety, and recovery processes.
  • Apply advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
  • Collaborate to ensure cloud infrastructure complies with relevant industry standards and that documentation is up to date.

Requirements

  • PhD or Master’s degree, or Bachelor’s degree in Computer Science, AI/ML, Engineering, or a related field, or equivalent practical experience.
  • 3-7+ years of professional software engineering experience.
  • Experience designing and developing high-scale distributed systems, cloud services, infrastructure platforms, or AI/ML platform services.
  • Strong programming skills in Java, Go, or Python with the ability to contribute production-grade code, reviews, tests, and debugging in complex distributed environments.
  • Expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
  • Proven use of Agile methodologies to drive continuous improvement and product delivery.
  • Excellent written and verbal communication skills.

Technologies

  • Java
  • Golang
  • Python
  • Kubernetes
  • Docker
  • LangGraph
  • LangChain
  • CrewAI
  • AutoGen
  • LlamaIndex
  • AWS
  • Azure
  • Google Cloud
  • Oracle Cloud
  • Oracle Cloud Infrastructure (OCI)
  • Codex
  • Claude Code
  • Cursor
  • Copilot

Preferred Qualifications

  • Experience with large scale cloud platforms such as AWS, Azure, Google Cloud, or Oracle Cloud.
  • Practical experience with orchestration frameworks like LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
  • Understanding of LLM application patterns, including prompt design, structured outputs, function/tool calling, context management, RAG, memory, tool safety, and evaluation.
  • Experience using AI assisted software development tools such as Codex, Claude Code, Cursor, Copilot, or similar systems in large-scale engineering environments.

Similar Jobs