EngineerJobs.io
← Back to all jobs

Job Description

Oracle is seeking an onsite Principal Software Engineer to lead architecture and development of scalable, elastic, fault-tolerant distributed systems.

Responsibilities

  • Lead development and initiate architecture of scalable, elastic distributed system components.
  • Define and enforce scalability requirements for owned components; optimize code and data paths for high-throughput, hyper-scale workloads.
  • Leverage data plane platform components for large-scale retrieval, storage, and processing.
  • Design fault-tolerant, in-service-upgradable systems using redundancy, replication, and failover.
  • Apply partition policies and operational techniques such as load-shedding, throttling, and rate-limiting to handle network unreliability while meeting SLOs.
  • Establish KPIs and telemetry; build proactive dashboards and alerts.
  • Design complex validation approaches including fault injection and brownouts to ensure correctness and durability.
  • Proactively diagnose and resolve production issues, mentor peers, and support operational readiness.
  • Implement robust security controls, execute remediation, and maintain compliance documentation.
  • Develop IaC and automation to support safe patching, updates, and rollbacks aligned to change-management plans.
  • Build distributed-system components enabling horizontal and vertical scaling using distributed-state-management tools.
  • Implement performance and load testing; design, develop, test, deploy, and maintain production software and cloud services.
  • Build fault-tolerant components that withstand in-service updates through redundancy, replication, and automatic failover.
  • Apply recovery-oriented computing patterns including retries, circuit breakers, and timeouts for network unreliability.
  • Implement tests, alarm configurations, fault-injection, and brown-out scenarios to detect failures and validate correctness.
  • Use standard data-replication and synchronization techniques to preserve integrity and availability.
  • Draft and execute runbooks and operational procedures; build dashboards, telemetry, and alerting for component health.
  • Diagnose, debug, and resolve component issues; develop strategies to prevent interruptions and avoid customer maintenance windows.
  • Design, implement, and maintain automation scripts, tooling, and IaC for troubleshooting and cloud-infrastructure management.
  • Participate in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
  • Apply advanced multi-tenant security measures, including encryption and access controls; implement remediation plans.
  • Ensure compliance with relevant industry standards and regulations; keep documentation current and follow change-management plans for patching, updates, and rollbacks.
  • Partner with product managers, architects, and engineering teams to translate requirements into technical solutions.
  • Create technical and operational documentation and mentor junior team members.

Requirements

  • 8 years of software-development experience, or a Bachelor’s degree in specified fields (Computer Science, Computer Engineering, Software Engineering, Electrical/Electronics Engineering, Computer/Information Systems, Information Technology, Telecommunications, Mathematics, Physics, or related) plus 4 years software-development experience, or a Master’s degree plus 2 years software-development experience, or a Doctorate in those fields.
  • Demonstrated ability or knowledge in distributed systems; prototyping; computer-science programming; software engineering; web development; innovation; cross-functional collaboration; information vulnerabilities; operating systems; API development and integration; applied algorithm engineering; and source control.
  • Proficiency in Java, Python, Go, C++, C#, or a similar language.
  • Strong skills in data structures, algorithms, object-oriented design, software engineering, problem-solving, communication, and collaboration.
  • Experience building, testing, and debugging production-quality software; familiarity with REST APIs, databases, cloud applications, and agile methodologies.
  • Cloud-platform experience (AWS, Azure, Google Cloud, or Oracle Cloud); system-level testing and automation; delivering and operating large-scale distributed systems.

Technologies

  • Java, Python, Go, C++, C#
  • REST APIs
  • AWS, Azure, Google Cloud, Oracle Cloud
  • Encryption, access controls
  • Data plane platforms
  • Distributed-state-management tools
  • Circuit breakers, timeouts, load-shedding, throttling, rate-limiting

System Scalability

  • Develop distributed-system components supporting horizontal and vertical scaling through distributed-state-management tools.
  • Optimize code and systems for large-scale data processing; implement scalability requirements and review team implementations.
  • Use data-plane platform components for large-scale retrieval, storage, and processing.
  • Implement performance and load testing; design, develop, test, deploy, and maintain production software and cloud services.

Reliability, Correctness, and Availability

  • Build fault-tolerant components that withstand in-service updates through redundancy, replication, and automatic failover.
  • Apply recovery-oriented computing with retries, circuit breakers, and timeouts for network unreliability.
  • Implement tests, alarm configurations, fault-injection, and brown-out scenarios to detect failures and validate correctness.
  • Use standard data-replication and synchronization techniques to preserve integrity and availability.
  • Draft and execute runbooks and operational procedures; build dashboards, telemetry, and alerting for component health.

Operations, Security, and Change Management

  • Diagnose, debug, and resolve component issues; implement strategies to prevent interruptions and avoid customer maintenance windows.
  • Implement automation scripts, tooling, and IaC for troubleshooting and cloud-infrastructure management.
  • Participate in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
  • Apply multi-tenant security measures including encryption and access controls; implement remediation plans.
  • Maintain compliance with industry standards and regulations; follow change-management plans for patching, updates, and rollbacks.

Collaboration and Delivery

  • Partner with product managers, architects, and engineering teams to translate requirements into technical solutions.
  • Create technical and operational documentation; drive projects forward and mentor junior team members.

Similar Jobs