Principal Software Engineer, Core Infrastructure
Job Description
Oracle is seeking an onsite Principal Software Engineer to lead architecture and development of scalable, elastic, fault-tolerant distributed systems.
Responsibilities
- Lead development and initiate architecture of scalable, elastic distributed system components.
- Define and enforce scalability requirements for owned components; optimize code and data paths for high-throughput, hyper-scale workloads.
- Leverage data plane platform components for large-scale retrieval, storage, and processing.
- Design fault-tolerant, in-service-upgradable systems using redundancy, replication, and failover.
- Apply partition policies and operational techniques such as load-shedding, throttling, and rate-limiting to handle network unreliability while meeting SLOs.
- Establish KPIs and telemetry; build proactive dashboards and alerts.
- Design complex validation approaches including fault injection and brownouts to ensure correctness and durability.
- Proactively diagnose and resolve production issues, mentor peers, and support operational readiness.
- Implement robust security controls, execute remediation, and maintain compliance documentation.
- Develop IaC and automation to support safe patching, updates, and rollbacks aligned to change-management plans.
- Build distributed-system components enabling horizontal and vertical scaling using distributed-state-management tools.
- Implement performance and load testing; design, develop, test, deploy, and maintain production software and cloud services.
- Build fault-tolerant components that withstand in-service updates through redundancy, replication, and automatic failover.
- Apply recovery-oriented computing patterns including retries, circuit breakers, and timeouts for network unreliability.
- Implement tests, alarm configurations, fault-injection, and brown-out scenarios to detect failures and validate correctness.
- Use standard data-replication and synchronization techniques to preserve integrity and availability.
- Draft and execute runbooks and operational procedures; build dashboards, telemetry, and alerting for component health.
- Diagnose, debug, and resolve component issues; develop strategies to prevent interruptions and avoid customer maintenance windows.
- Design, implement, and maintain automation scripts, tooling, and IaC for troubleshooting and cloud-infrastructure management.
- Participate in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
- Apply advanced multi-tenant security measures, including encryption and access controls; implement remediation plans.
- Ensure compliance with relevant industry standards and regulations; keep documentation current and follow change-management plans for patching, updates, and rollbacks.
- Partner with product managers, architects, and engineering teams to translate requirements into technical solutions.
- Create technical and operational documentation and mentor junior team members.
Requirements
- 8 years of software-development experience, or a Bachelor’s degree in specified fields (Computer Science, Computer Engineering, Software Engineering, Electrical/Electronics Engineering, Computer/Information Systems, Information Technology, Telecommunications, Mathematics, Physics, or related) plus 4 years software-development experience, or a Master’s degree plus 2 years software-development experience, or a Doctorate in those fields.
- Demonstrated ability or knowledge in distributed systems; prototyping; computer-science programming; software engineering; web development; innovation; cross-functional collaboration; information vulnerabilities; operating systems; API development and integration; applied algorithm engineering; and source control.
- Proficiency in Java, Python, Go, C++, C#, or a similar language.
- Strong skills in data structures, algorithms, object-oriented design, software engineering, problem-solving, communication, and collaboration.
- Experience building, testing, and debugging production-quality software; familiarity with REST APIs, databases, cloud applications, and agile methodologies.
- Cloud-platform experience (AWS, Azure, Google Cloud, or Oracle Cloud); system-level testing and automation; delivering and operating large-scale distributed systems.
Technologies
- Java, Python, Go, C++, C#
- REST APIs
- AWS, Azure, Google Cloud, Oracle Cloud
- Encryption, access controls
- Data plane platforms
- Distributed-state-management tools
- Circuit breakers, timeouts, load-shedding, throttling, rate-limiting
System Scalability
- Develop distributed-system components supporting horizontal and vertical scaling through distributed-state-management tools.
- Optimize code and systems for large-scale data processing; implement scalability requirements and review team implementations.
- Use data-plane platform components for large-scale retrieval, storage, and processing.
- Implement performance and load testing; design, develop, test, deploy, and maintain production software and cloud services.
Reliability, Correctness, and Availability
- Build fault-tolerant components that withstand in-service updates through redundancy, replication, and automatic failover.
- Apply recovery-oriented computing with retries, circuit breakers, and timeouts for network unreliability.
- Implement tests, alarm configurations, fault-injection, and brown-out scenarios to detect failures and validate correctness.
- Use standard data-replication and synchronization techniques to preserve integrity and availability.
- Draft and execute runbooks and operational procedures; build dashboards, telemetry, and alerting for component health.
Operations, Security, and Change Management
- Diagnose, debug, and resolve component issues; implement strategies to prevent interruptions and avoid customer maintenance windows.
- Implement automation scripts, tooling, and IaC for troubleshooting and cloud-infrastructure management.
- Participate in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
- Apply multi-tenant security measures including encryption and access controls; implement remediation plans.
- Maintain compliance with industry standards and regulations; follow change-management plans for patching, updates, and rollbacks.
Collaboration and Delivery
- Partner with product managers, architects, and engineering teams to translate requirements into technical solutions.
- Create technical and operational documentation; drive projects forward and mentor junior team members.
Similar Jobs
J