Daren Jacobs
Lumberton, NJ • ******@*****.*** • 848-***-**** • linkedin.com/in/darenjacobs
SUMMARY
Cloud infrastructure and platform engineering leader with deep database and DevOps expertise across AWS, GCP, and Kubernetes. Currently leads the team provisioning and operating single-tenant enterprise SaaS environments at a 99.99% uptime SLA, and architected the AI-driven automation platform behind its support and SRE workflows. Track record of standardizing infrastructure with Terraform and GitOps, cutting environment build-out from a week to hours, reducing per-environment cloud spend by up to 77%, and converting manual operational work into automated, auditable tooling.
SKILLS
CI/CD Pipelines Configuration Management Nginx / HAProxy / ELB / NLB AWS / GCP / AZURE
Redis / Memcached Aurora / MySQL RDS Prometheus / Grafana / ELK IAM / KMS / VPC
Docker / Kubernetes Bash / Python Scripting Terraform / Ansible Jenkins / GitLab CI / ArgoCD
ClickHouse / Cassandra Helm / EKS / GitOps GitHub Actions AI / LLM Automation
EXPERIENCE
Comet - New York, NY (Remote) September 2024 – Present
Deployment Team Lead
• Lead the deployment team responsible for provisioning and operating single-tenant SaaS (STSaaS) environments for 15+ enterprise customers including Bayer, Netflix, BMW, Porsche, and Mercedes across five AWS regions and GCP, maintaining a 99.99% uptime SLA.
• Architected an AI-powered (Claude/Anthropic) investigation and automation platform that unifies customer support, SRE, and DevOps workflows, scaling it to 65+ automated workflows with 10+ engineering contributors and covering read-only diagnostics across Kubernetes, ArgoCD, RDS/MySQL, ClickHouse, and Redis.
• Codified the full single-tenant environment lifecycle (provisioning through decommissioning) as gated, phase-by-phase workflows across Terraform (EKS, VPC, IAM), Helm, and ArgoCD GitOps with automated post-deploy smoke tests, cutting new enterprise environment build-out from roughly one week to ~3 hours (90%+ reduction) while making it repeatable, drift-free, and safely reversible.
• Automated daily health checks across all customer deployments and built alert-triage tooling for Alertmanager and Pingdom that separates signal from noise, correlates alerts against live cluster and database state, and auto-posts prioritized action items to Jira and Slack.
• Operationalized the health-check automation into a proactive reliability practice, catching OOMKills, memory pressure, stalled pods, and ClickHouse/Aurora/Redis anomalies before they reached end users, and driving right-sizing remediations (backend and frontend memory limits) through GitOps pull requests that reinforce the 99.99% uptime SLA.
• Engineered a load-replay harness that records real production ingestion traffic from ClickHouse (read-only and scrubbed) and replays it at controlled rate and amplification into test environments with a production-parity Grafana/Prometheus stack, reproducing CPU, memory, and database saturation before fixes ship.
• Diagnosed production JVM failures through headless heap- and thread-dump analysis (Eclipse MAT, OQL), isolating out-of-memory root causes, deadlocks, and thread- and connection-pool exhaustion.
• Led database performance and cost optimization across ClickHouse, Aurora/MySQL, and Redis including a backup-retention fix that reclaimed ~1.2 TiB of storage for a single customer, plus ongoing slow-query, replication, and cardinality (HLL) tuning.
• Drove FinOps and rightsizing using AWS Trusted Advisor and DoiT Cloud Analytics, cutting monthly infrastructure spend 67–77% on individual enterprise environments (~$22K to ~$5K and ~$18K to ~$6K per month) for roughly $350K in annualized savings, and automated recurring cost-anomaly detection to Slack.
• Migrated legacy customer Terraform onto a consolidated, versioned terraform-aws-comet-stsaas module, cutting configuration drift and simplifying multi-region infrastructure management.
• Enforced least-privilege, read-only operational guardrails and automated secret redaction across the tooling to protect production customer data while enabling safe self-service diagnostics.
Qualcomm - San Diego, CA (Remote) December 2020 – September 2024
Senior Staff Engineer (DevOps)
•Led CI/CD solutions implementation with Jenkins, ArgoCD to OpenShift and GKE K8s platforms, enhancing deployment efficiency.
•Achieved a 60% reduction in deployment time and a 40% decrease in maintenance efforts through Terraform and Ansible for infrastructure as code.
•Collaborated with 20+ development teams, reducing time-to-market by 40% and earning a promotion.
•Improved data availability by 20% and reduced latency by 15% through integration of Vertica, Apache NiFi, MongoDB, MySQL, and ActiveMQ.
•Integrated cloud-based security mechanisms into build infrastructure, including IAM and KMS.
•Developed and managed production database alerts, conducted root-cause investigations, and led FinOps for data storage systems.
•Implemented Prometheus, Grafana, and ELK Stack for monitoring, logging, and observability, significantly enhancing system performance and visibility.
Tellic - New York, NY November 2018 – December 2020
Senior DevOps Engineer
•Pioneered comprehensive automation strategy with Python and machine-learning, boosting productivity by 50% and reducing manual errors by 75%.
•Deployed GCP based CI/CD pipeline with SonarQube and Locust, enhancing testing efficiency.
•Automated ETL processes in GCP, reducing data processing time by 50% and saving 20 hours per week.
•Led project with data scientists, increasing online sales conversion rate by 10% through a personalized recommendation engine.
•Co-designed NLP application with Graph database and Apache Solr, improving search accuracy by 35% and reducing query response time by 30%.
Federal Home Loan Bank of New York - Jersey City, NJ February 2018 – November 2018
Senior Cloud Operations Engineer
•Enhanced code quality with SonarQube, focusing on vulnerability and bug resolution.
•Utilized Python for AWS infrastructure building and resource utilization analysis, reducing quarterly AWS bills by 23%.
•Increased system availability and scalability by 50% through microservices transition, saving significant costs.
•Achieved a 30% increase in operational efficiency by orchestrating a comprehensive cloud migration strategy.
•Developed CI/CD pipelines in AWS with Packer, Terraform, and Ansible, automating deployment processes.
Upsight - San Francisco, CA (Remote) October 2016 – February 2018
Senior DevOps Engineer
•Developed tool for rapid AWS EC2 VM deployment, reducing deployment time to 50 seconds and saving over $20k/month.
•Led MySQL database replication, backups, and data migrations, ensuring data integrity.
•Designed SaltStack configuration management solution, improving efficiency by over 50%.
•Created queries and shell scripts for cloud services metrics, enhancing C-level data accessibility.
Kaltura - New York, NY October 2012 – October 2016
Systems Engineer
•Revolutionized Kaltura application deployment with Ansible, increasing deployment speed by 200%.
•Executed flawless migration of customer media content to SaaS platform, achieving 100% data accuracy.
•Spearheaded integration and administration of customer data centers globally, earning industry accolades.
•Served as primary support liaison for on-premises customers, enhancing customer support experience.
EDUCATION & CERTIFICATIONS
University of Phoenix • Bachelor of Science • Information Technology • May 2012
AWS Certified DevOps Engineer, Professional • 2020–2023
AWS Certified Security, Specialty • 2020–2023
AWS Certified Solutions Architect, Associate • 2020–2023