Software Developer
We are hiring a Software Developer who will build and support reliable, high capacity, and well-performing systems in support of our mission to protect and improve the client. An ever-watchful eye on reliability, security, performance, cost, and operational excellence. We call this work Site Reliability Engineering. As a Site Reliability Engineer within a small team, you will collaborate in a DevOps model with product development teams; designing, deploying, and managing automation tools that increase predictability as well as time to market while reducing cost. If you love to seek automation over repetitive tasks, are well-versed in AWS services and a desire to continually learn, have complex distributed system experience, and like engineering software solutions to solve cloud-related problems, then you will thrive in this position.
Our stack Code: Node, PHP, Python, Javascript, YAML, Bash
RDBMS: PostGreSQL, MySQL
Cache: ElastiCache (Redis/memcached), DynamoDB
Containers: ECS, EKS & Docker
Cloud: Amazon AWS
Telemetry: New Relic, CloudWatch
Build: Jenkins, CircleCI, GitHub Enterprise
Web: Apache httpd
Infrastructure as Code: Terraform (Preferred), CloudFormation
Your contributions
Cloud Engineering
Hands-on design, analysis, development and troubleshooting of highly-distributed large-scale production systems and event-driven, cloud-based services
Ensure repeatability, traceability, and transparency of our infrastructure automation (infrastructure-as-code, monitoring-as-code)
Participate in continual learning of the AWS ecosystem, game day scenarios, and professional conferences
Collaborative solutioning of enterprise applications with development teams utilizing our software stack
Actively monitor AWS Cost, and utilize optimizer to maximize ROI while maintaining Service Level Objectives
Observability Engineering
Ownership of reliability, uptime, system security, cost, operations, capacity, resiliency and performance thereof
Define, monitor and report on service level indicators for applications workloads
Support on-call rotations for operational duties that have not been addressed with automation, with an eye for correcting issues that result in on-call alarms
Maintain telemetry that improve the visibility to our applications' performance and business metrics and keep operational workload in-check
Develop, communicate, collaborate, and monitor standard processes to promote the long-term health and sustainability of operational development tasks.
DevSecOps
Support healthy software development practices, including complying with agile software development methodology, building standards for code reviews, work packaging, and continuous delivery
Partner with CyberSecurity and develop plans and automation to respond to new risks and vulnerabilities
Qualifications
Experience as a software engineer, with practical experience developing, debugging, and deploying enterprise applications
Experience with infrastructure automation technologies, preferably Terraform
Experience in container/container-fleet-orchestration technologies, preferably EKS or ECS
Versatility with troubleshooting diverse sets of hosting technologies: web server platforms, application platforms, operating systems, network components, virtualization technologies, storage, and database platforms.
Experience with continuous-deployment based software development lifecycles (e.g. CI/CD)
Experience with application caching strategies and high concurrency workloads
Strong communication, problem solving, root cause analysis and systems engineering skills
Ability to design and manage escalation response plans from monitoring, react, respond, remediate and retrospect in culturally aligned (proactive, customer focused, collaborative, data-driven) ways.
Demonstrated expertise building and managing highly scaled production infrastructure in the cloud
BS Degree in Computer Science (or related technical field and/or equivalent industry experience)