Post Job Free
Sign in

Principal DevOps & ML Infrastructure Strategist

Location:
Brentwood, CA
Salary:
Let's talk, I deserve every dollar I earn.
Posted:
August 31, 2026

Contact this candidate

Resume:

MOHAMED ABUAITA 510-***-**** ********@*****.***

Seek an opportunity to help map the strategy of technology, and to play a strategic role shaping the future of enterprise. I hope to expand the role ML and AI in the tech stack and to expand on my work with Sagemaker, Flink, Spark, and other Data Transformation Technologies. QUALIFICATIONS

• More than 15 years in service engineering and network operations, managing 24x7 customer facing applications.

• Solid understanding of monitoring, logging, and alerting practices (e.g., Prometheus, Grafana, ELK, Datadog).

• Experience with full life cycle system management; availability management, capacity management, configuration management, change management, and incident reporting.

• Experience with Data Center Operations, Network Operations, 24/7 Support/NOC, and Remote Site Support Processes.

• Experience managing 3rd party vendors and providers, including outsourcing. Experience in negotiating & managing Systems contracts and vendor management.

• Experience dealing with enterprise application security, compliance, and reporting.

• Experience hosting high availability applications and meeting availability and performance SLA’s.

Technologies: ML technologies, CI/CD, release management, everything AWS, Terraform, Docker, Kubernetes, helm, chef, git, sql and nosql, redis, kafka/zoo, haproxy, CDN, serverless technology, nginx, apache, datadog, F5, cisco gear, and storage management. I program in Bash and Python.

EXPERIENCE

Clario Jun 2018 – Mar 2026

Principal DevOps Engineer

Member of the seed group for an Informatics Startup, we are fostering an agile Bring to Market environment by automating the build and deployment of new releases, as well as, the launch and scale of new AWS environments.

• Using LLM, the team scanned large data sets for patient privacy information as well as image traces of critical diagnosis.

• Managed the production services across12 data centers globally, including budgeting, vendor managements, third party relations, reporting, and initiative planning and implementation.

• Used Kinesis and SageMaker to ingest streams of patient data from more than 1500 national hospitals in 12 Data Centers.and for making it available to research scientists in the US, EU, APAC, and the middle east.

• Custodian of more then 8 Peta Bytes of patient cases.

• Integrated and moved data between AWS and Azure.

• Extensively used technologies including EKS, ECS, Python, Lua, OpenResty, Terraform,Helm, and ArgoCD. Built and managed bare-metal Kubernetes cluster and configured Rook, Ceph, Object Storage, Calico, Istio, and Vault.

• Observability technologies including Prometheus, Grafana, and Loki.

• Created a JWT web service to authenticate research scientists and allow them to download scientific images. The service is based on nginx, openresty/lua, and redis.

• Architect and developer of python-based technology to transfer cases/images from different regions in the world for centralized processing and distribution

• The technology team set off to foster agile, bring it to market and championed a widespread use of scrums, kanban, retro, git integrations, and deployment pipeline. Zoosk Jan 2015 – June 2018

Manager of DevOps

At Zoosk, the devops team is responsible for evolving, maintaining, and the upkeep of ZCloud, Zoosk’s AWS cloud environment, which is made up of 44 different AMI builds, more than 200 cloudformation stacks, more than 55 service tiers, more than 50 docker containers, more than 200 QA instances, and more than 300 production instances.

• Team converted all infrastructure tiers to be automatable – allowing global rollouts in case security or RPM update triggers it.

• Team migrated service units such as docker hub, enterprise github, Jenkins, and legacy services from collocated center to AWS.

• Leading initiative to migrate legacy apps to Kubernetes in order to orchestrate and secure deployments. Early adapter of service mesh using Istio.

• Team integrated Akamai and CloudFront to improve performance, and adapted Cequence and Aqua to enhance security.

Ultimate Gaming May 2012 – Jan 2015

Director of Infrastructure

Head of department responsible for the operations of online, real-money, gaming. Managed a team of professionals:

• Team built data centers in Las Vegas and Atlantic City with OpEx exceeding $1.5 million. Both centers completed ahead of schedule and featured cutting edge technologies; UCS, EMC, 100% fabric, BigIP ltm.

• Designed and configured a mesh interwork of 3 branches and 3 datacenters.

• Team unitized and automated deployment of subsystems, utilizing Chef, and git to manage unit deployments.

• The team successfully fostered an agile deployment and release management cycle with close to %100 CD adoption across all units.

• Built a dashboard, using python, that tracked player attempted deposits and geolocation. The app was used by Customer Services to identify players who maybe having problems depositing money, and those not allowed to play because of erroneous geolocation.

• Designed compliance reports required by the governing bodies in NV and NJ.

• Saved more than $1.3 million in contract negotiations.

• Built two new branch offices and managed two office moves. Responsible for internal IT.

• Managed a mixed environment with more than 300 nodes in AWS and another 200 servers in co-located facilities.

SendMe Jun 2008 – Feb 2012

Director of Network Operations/IT

Responsible for the operations of a consumer-oriented, mobile technology company. Managed the production systems of three platforms; Sendmemobile.com, Solow.com, and mbuzzy.com.

• Member of a team that netted the company $200 million in the first 2 years of operation.

• Created a dashboard in python to monitor and report on carrier (AT&T, VZN, TMO, and Sprint) network availability and performance. The tool was used by CS and marketing

• Architect and managed state of the art data center; with technologies from IBM, HP, EMC, F5, and Cisco.

• Used EC2 and S3 elastic services to augment a co-located center capacity and handle 10-15 million daily YouTube impressions.

• Used Chef and Jenkins to automate and manage builds, and code deployment.

• Implemented change management and configuration control with great success resulting in better than %99.99 service uptime.

Presto Services Aug 2005 – Jun 2008

Director of Network Operations

Managed the security, network, server, and storage infrastructure for a consumer-oriented; high availability, production environment. Presto sold intelligent HP printers that can be placed in homes and used to print emails, pictures, and events. They used regular phone lines.

• Managed a team of 5 with responsibility of an e-commerce service capable of serving a multi-million user base. Handled all major credit cards.

• Built a dashboard that tracked in real time financial transactions and printer interaction.

• Led efforts to launch the service and achieved production readiness ahead of schedule.

• Maintained better than %99.999 uptime through thorough monitoring, responsible coverage, and incident management.

• Designed the monitoring framework which successfully intercepted every occurrence of component failure. The framework used Nagios as well as scripting in Selenium.

• Positioned operations as one of the most respected teams in the organization. Adobe Jul 2003 - Aug 2005

Internet Application Infrastructure Architect

Worked with the engineering group to enable users to store and share photos.

• We built an infrastructure that can accommodate a multi-million user base, and over 120 million e-commerce transactions per year.

• Had budgetary responsibility for the network, server, storage, and data center facilities. CapEx for the first year exceeded $1 million.

• Installed a network capable of handling 2 million visits per day, and file uploads/downloads averaging 1 megabyte per file and 10 files per user per day.

• Relied on concepts such as grid computing, and horizontal expansion to integrate technologies such as Linux GFS, iSCSI, and MySQL Cluster.

• Mitigated threats with packet filtering, layer 4-7 analysis, IPS, and HW hardening.

• Worked with engineering to build a cost-effective file storing nomenclature that can scale indefinitely.

BIZ360. Oct 2001 - Jul 2003

Operations Manager for an Application Service Provider Responsible for a SaaS application and corporate IT.

• Led effort to migrate BIZ360 hosted environment from 3rd party managed services to locally managed data center. Finished project 3 weeks ahead of time and $90k under budget.

• We improved SLA availability from 98% to above 99.99% by implementing Change Management and Risk Controls.

• Outsourced and managed a NOC team based in the Philippines.

• We achieved 3 times application performance enhancement through fine-tuning and scalable design.

• Instituted smart response to security events using Snort, Site-scope, and Python. EDUCATION

Bachelor of Sciences in Systems Engineering, San Jose State University CERTIFICATES

CCNA, CCNP, and CCDP.



Contact this candidate