Post Job Free
Sign in

Sr. Site Reliability Engineer

Company:
Mobile Programming
Location:
Mohali, Punjab, India
Posted:
September 09, 2026
Apply

Description:

We are looking for a Senior Site Reliability Engineer to join our global SRE team. This is a hands-on individual contributor role for an experienced engineer who can take ownership of high-severity incidents and drive reliability improvements across a fast-moving fintech infrastructure. You will serve in the on-call rotation, work closely with global engineering teams to maintain system health, and improve the resilience of our platforms.

Responsibilities:

Serve in the on-call rotation and act as incident lead for high-severity incidents.

Drive incident triage, coordinate cross-team response, communicate status to stakeholders, and own incidents through resolution.

Independently triage and resolve production alerts across AWS infrastructure and application layers.

Own and evolve Datadog monitoring standards, including monitors, dashboards, and SLO/SLI frameworks.

Drive signal-to-noise improvements to reduce alert fatigue.

Create and maintain runbooks and SOPs for failure scenarios.

Lead post-mortems with structured root cause analysis and actionable follow-ups.

Lead reliability initiatives to reduce operational toil and inefficiency.

Identify systemic weaknesses and implement automation and process improvements.

Lead RCA investigations for infrastructure and application-level failures in AWS environments.

Produce clear, action-oriented incident reports.

Mentor junior and mid-level on-call engineers.

Monitor CI/CD pipelines during deployments and initiate rollbacks when required.

Requirements:

6-8 years of hands-on experience in SRE, DevOps, or platform engineering.

Proven experience leading high-severity production incidents.

Strong AWS expertise, including Amazon ECS, IAM, VPC, ALB/NLB, RDS, S3 MSK, ElastiCache, Lambda, and CloudWatch.

Proficiency with Terraform for Infrastructure as Code.

Strong experience with Datadog, including monitors, dashboards, SLOs/SLIs, and error budgets.

Strong Linux command-line skills.

Ability to write Python automation scripts for operational tasks.

Working knowledge of GitHub Actions or GitLab CI for ECS-based deployments.

Strong understanding of reliability concepts and production troubleshooting.

Excellent written and verbal communication skills.

Strong ownership and structured troubleshooting mindset.

Architecture & Reliability:

Understanding of architectural patterns such as microservices, pub/sub, and load balancing

Ability to evaluate reliability trade-offs and failure modes.

Experience driving reliability improvements and operational best practices.

Preferred Skills:

Basic understanding of financial markets and market data, including equities, options, and market data feeds.

Database query skills.

Familiarity with RDS or Cassandra performance metrics.

Experience mentoring engineers.

Experience establishing operational best practices.

Strong documentation and handover practices.

Ability to learn unfamiliar systems independently through documentation and runbooks.

Apply