Lead Infrastructure Engineer
Join Corporate Technology at JPMorganChase, where infrastructure engineers are empowered to build resilient, scalable platforms that power critical business intelligence capabilities across the firm. Here, your work directly enables data-driven decisions at enterprise scale—and you'll have the tools, autonomy, and collaborative environment to make a lasting impact.
As a Lead Infrastructure Engineer at JPMorganChase within Corporate Technology, you will design, build, and operate the hybrid infrastructure supporting enterprise business intelligence platforms across on-premises and cloud environments. You will drive platform reliability, automation, and AI-enabled operational efficiencies that reduce toil, accelerate incident response, and elevate the quality of operational knowledge management. Your contributions will directly influence platform resilience, engineering standards, and the adoption of intelligent automation across the team.
Job Responsibilities
Engineer and sustain hybrid infrastructure environments supporting enterprise analytics platforms across development, test, and production tiers, including compute, storage, networking, load balancing, DNS, certificates, and OS configurations
Provision and manage infrastructure using Terraform, building reusable modules and standardized patterns that enable consistent, repeatable deployments across environments
Automate recurring operational tasks—including environment builds, validation checks, certificate rotations, and health checks—to reduce manual toil and improve platform reliability
Drive platform resiliency through high availability design, backup and restore procedures, disaster recovery planning and testing, and environment standardization
Implement and maintain monitoring and alerting for infrastructure and platform dependencies, including system metrics, service health, log signals, and network performance
Lead incident response, root-cause analysis, and corrective action tracking, producing structured documentation that reduces repeat incidents and improves mean time to resolution
Identify and implement AI and large language model-assisted efficiencies across operational workflows, including incident summarization, runbook generation, alert correlation, and change-plan validation
Partner with platform administrators and application and data teams to align infrastructure capacity, job scheduling, concurrency, and performance baselines with platform requirements
Produce and maintain operational documentation, including architecture diagrams, standard operating procedures, troubleshooting runbooks, and support checklists
Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.
Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.
Required Qualifications, Capabilities, and Skills
Formal training or certification on infrastructure engineering concepts and 5+ years applied experience
Demonstrated experience engineering and supporting production infrastructure in hybrid on-premises and cloud environments
Strong Linux administration and troubleshooting skills, including service health, system logs, resource contention, process management, and connectivity diagnostics
Hands-on experience with Terraform, including module development, state management, environment separation, and repeatable infrastructure provisioning
Experience with IT service management processes and tooling, including incident, change, and problem management workflows
Proven ability to drive automation and reliability improvements with measurable operational outcomes
Strong production ownership mindset with structured root-cause analysis skills and calm, effective execution under pressure
Clear written and verbal communication skills, with the ability to engage both technical teams and business stakeholders
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
Preferred Qualifications, Capabilities, and Skills
Experience supporting infrastructure for enterprise analytics or business intelligence platforms, including scaling concepts, scheduling and extract window management, and dependency troubleshooting
Demonstrated experience applying AI or large language model tooling to operational workflows—such as triage, documentation generation, or automation—with an evaluation and continuous improvement mindset
Scripting and automation experience using Python or Bash for operational tooling and workflow enhancement
Familiarity with observability tooling, including log aggregation, metrics platforms, and alert tuning
Knowledge of security fundamentals relevant to infrastructure operations, including certificate lifecycle management, secrets management, least privilege principles, and vulnerability remediation
Understanding of network fundamentals applicable to enterprise platforms, including load balancers, reverse proxies, firewall rules, DNS, and TLS