EngineerJobs.io
← Back to all jobs

Job Description

Oracle's OCI AI Infrastructure team seeks a Principal Software Engineer to design, develop, and optimize cloud software for AI/ML workloads and GPU delivery.

Location: Austin, TX (onsite)

Responsibilities

  • Operate autonomously in ambiguous contexts while upholding established standards and practices.
  • Design, build, troubleshoot, and debug software for diverse cloud infrastructure components, including databases, applications, tools, and networks.
  • Contribute to defining and evolving standard software engineering practices with a focus on AI driven development.
  • Create software for tasks involved in developing, designing, and debugging applications or operating systems, leveraging AI and ML techniques.
  • Lead the development of strategic initiatives, including:
    • Implement spike detection for provisioning failures using ML to minimize operational disruptions.
    • Expand Kafka integrations to enable near real-time actions supporting 1-Day SLO hardware repairs, using event driven architecture and stream processing.
    • Develop an automated ticket routing framework to streamline workflows with NLP and ML.
    • Drive initiatives through collaboration with cross-functional teams and customers, applying AI driven insights and recommendations.
    • Leverage AI and ML to build tools that automate testing, simulate complex environments, and reproduce incidents, freeing engineers for higher-value work and better customer outcomes.
    • Collaborate and lead technical discussions across teams to ensure seamless integrations and effective problem solving.
    • Provide guidance and mentorship to junior engineers to foster growth and development.

Requirements

  • Python
  • Java
  • TypeScript
  • Agile Principles
  • Data modeling
  • Data warehousing
  • Data governance
  • OCI
  • AWS
  • Azure
  • Google Cloud Platform (GCP)
  • Linux
  • MacOS
  • Bash
  • Perl
  • Ruby
  • Docker
  • RESTful APIs
  • API gateways
  • API security
  • Swagger/OpenAPI
  • Chatbots
  • Virtual assistants
  • Predictive analytics
  • Kafka
  • RoCE
  • Infiniband
  • NLP
  • ML

Technologies

  • Oracle Cloud Infrastructure (OCI)
  • AWS
  • Azure
  • Google Cloud Platform (GCP)
  • Linux
  • MacOS
  • Docker
  • Python
  • Java
  • TypeScript
  • Bash
  • Perl
  • Ruby
  • RESTful APIs
  • API gateways
  • API security
  • Swagger/OpenAPI
  • Kafka
  • RoCE
  • Infiniband
  • NLP
  • ML
  • Chatbots
  • Virtual assistants
  • Predictive analytics
  • Data modeling
  • Data warehousing
  • Data governance

Description

OCI AI Infrastructure is building a cutting edge, ultra high performance GPU platform to support AI, ML, and HPC workloads at scale. The GPU Availability and Monitoring team focuses on architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services across thousands of GPUs, leveraging RoCE and InfiniBand.

This role offers the opportunity to shape the future of cloud infrastructure and automation, collaborating across teams and with customers to continuously advance our technology stack and drive meaningful impact in the AI infrastructure space.

Similar Jobs