Post Job Free
Sign in

Site Reliability Engineer

Company:
Runloop AI
Location:
San Francisco, CA
Posted:
July 28, 2026
Apply

Description:

Site Reliability Engineer

Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the co-founder of Google Wallet and 100% of its founding team, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

We're looking for a skilled and passionate Site Reliability Engineer to join our team. As an SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform, the foundation our users build their work on. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of operations with a software engineering mindset.

Responsibilities

Design and maintain our production infrastructure on cloud platforms like AWS, GCP, or Azure

Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users' code

Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind

Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment

Participate in an on-call rotation to support our production systems

Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing

Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems

Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement

Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience

Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes

Qualifications

Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience

5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations

Strong programming skills in languages like Python or Go

Deep expertise in containerization technologies such as Docker and Kubernetes

Experience with cloud infrastructure and tools like Terraform and/or Pulumi

Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog

A solid understanding of networking, security, and Linux systems administration

Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)

Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity

Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems

Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance

Bonus Points

Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery

Benefits

Competitive salary and equity

Comprehensive health, dental, and vision insurance for employee and dependents

Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering

Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks

Location

Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Join Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.

Apply