Post Job Free
Sign in

Senior Site Reliability Engineer

Company:
Hackajob
Location:
Reston, VA, 20191
Posted:
August 15, 2026
Apply

Description:

Job Title

Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning

You will provide cloud operations for Oracle National Security Realms. You'll be part of a dynamic team with a broad knowledge of how Oracle's cloud platform works. You'll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers.

Note - this role is not a Monday to Friday core hours role – it will involve working a 24/7 shift rotation with on-call duties, including nights, weekends and public holidays.

Responsibilities

Escalation points for junior site reliability engineers during complex or high-impact incidents.

Manage and execute complex manual Change Management tickets, by working closely with the service teams to ensure safety and minimal disruption to services.

Support the on-boarding of new services and tools, ensuring they are operationally ready and properly integrated.

Provide mentorship and training to SREs, helping build team capability and confidence.

Create and maintain clear, useful documentation for operational processes and system support.

Identify areas of manual work and drive automation to reduce toil and improve efficiency.

Automate tasks to enable continuous delivery and ensure continuous availability with minimal human overhead

Recognize unsafe or inefficient practices and work with teams to design safer, more effective solutions.

Complete change requests to enable new functionality and maintain realm compliance

Ensure timely resolution of incidents, service requests, and change requests

Collaborate with global service and engineering teams

Define and drive change management, continuous integration, and deployment best practices

Help create and maintain real-world production architectures, scalability, and system design

Use a methodical approach to troubleshoot, large, complex, interconnected systems

Technologies Used

Linux and Unix operating systems

Docker, Kubernetes, and Terraform

Scripting languages such as Bash, shells, Perl, or Python

Citizenship/location requirements - i.e. US Citizenship, U.S. Citizenship and possess and maintain TS/SCI w/Poly security clearance, reside in Austin, TX or Reston, VA

Technology related bachelor's degree and/or equivalent work experience

A desire to learn and keep up with modern technologies

Proficient with writing services/task automation in any modern development language (e.g. Python, Bash, Ruby, Perl, JavaScript, or Java)

Familiarity with core protocols and OSI model (DNS, DHCP, HTTP, TCP/IP)

Deep knowledge of Linux or Unix OS internals and host-based networking

Familiarity with configuration management solutions such as Chef, Puppet, etc

Experience with devising, managing, and extending monitoring solutions for large scale environments.

Knowledge of cloud computing concepts

Experience working in a mission-critical environment (Operations, Technical Support, NOC etc)

Proficient with communication skills (writing, organization, learning exchange)

Experience executing tasks under change management procedures

Experience resolving auto-cut and manual alarms following runbooks

A focus on customer satisfaction

Specific experience working with deployment of AI infrastructure to include clustered GPUs, LLM deployment and maintenance, and understanding of model integration for customer solutions

Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only

Oracle US offers a comprehensive benefits package which includes the following:

1. Medical, dental, and vision insurance, including expert medical opinion

2. Short term disability and long term disability

3. Life insurance and AD&D

4. Supplemental life insurance (Employee/Spouse/Child)

5. Health care and dependent care Flexible Spending Accounts

6. Pre-tax commuter and parking benefits

7. 401(k) Savings and Investment Plan with company match

8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

9. 11 paid holidays

10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

11. Paid parental leave

12. Adoption assistance

13. Employee Stock Purchase Plan

14. Financial planning and group legal

15. Voluntary benefits including auto, homeowner and pet insurance

Apply