Lead Infrastructure EngineerJoin Corporate Technology at JPMorganChase, where infrastructure engineers are empowered to build resilient, scalable platforms that power critical business intelligence capabilities across the firm.
Here, your work directly enables data-driven decisions at enterprise scale—and you'll have the tools, autonomy, and collaborative environment to make a lasting impact.As a Lead Infrastructure Engineer at JPMorganChase within Corporate Technology, you will design, build, and operate the hybrid infrastructure supporting enterprise business intelligence platforms across on-premises and cloud environments.
You will drive platform reliability, automation, and AI-enabled operational efficiencies that reduce toil, accelerate incident response, and elevate the quality of operational knowledge management.
Your contributions will directly influence platform resilience, engineering standards, and the adoption of intelligent automation across the team.Job ResponsibilitiesEngineer and sustain hybrid infrastructure environments supporting enterprise analytics platforms across development, test, and production tiers, including compute, storage, networking, load balancing, DNS, certificates, and OS configurationsProvision and manage infrastructure using Terraform, building reusable modules and standardized patterns that enable consistent, repeatable deployments across environmentsAutomate recurring operational tasks—including environment builds, validation checks, certificate rotations, and health checks—to reduce manual toil and improve platform reliabilityDrive platform resiliency through high availability design, backup and restore procedures, disaster recovery planning and testing, and environment standardizationImplement and maintain monitoring and alerting for infrastructure and platform dependencies, including system metrics, service health, log signals, and network performanceLead incident response, root-cause analysis, and corrective action tracking, producing structured documentation that reduces repeat incidents and improves mean time to resolutionIdentify and implement AI and large language model-assisted efficiencies across operational workflows, including incident summarization, runbook generation, alert correlation, and change-plan validationPartner with platform administrators and application and data teams to align infrastructure capacity, job scheduling, concurrency, and performance baselines with platform requirementsProduce and maintain operational documentation, including architecture diagrams, standard operating procedures, troubleshooting runbooks, and support checklistsUses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.Required Qualifications, Capabilities, and SkillsFormal training or certification on infrastructure engineering concepts and 5+ years applied experienceDemonstrated experience engineering and supporting production infrastructure in hybrid on-premises and cloud environmentsStrong Linux administration and troubleshooting skills, including service health, system logs, resource contention, process management, and connectivity diagnosticsHands-on experience with Terraform, including module development, state management, environment separation, and repeatable infrastructure provisioningExperience with IT service management processes and tooling, including incident, change, and problem management workflowsProven ability to drive automation and reliability improvements with measurable operational outcomesStrong production ownership mindset with structured root-cause analysis skills and calm, effective execution under pressureClear written and verbal communication skills, with the ability to engage both technical teams and business stakeholdersDemonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.Preferred Qualifications, Capabilities, and SkillsExperience supporting infrastructure for enterprise analytics or business intelligence platforms, including scaling concepts, scheduling and extract window management, and dependency troubleshootingDemonstrated experience applying AI or large language model tooling to operational workflows—such as triage, documentation generation, or automation—with an evaluation and continuous improvement mindsetScripting and automation experience using Python or Bash for operational tooling and workflow enhancementFamiliarity with observability tooling, including log aggregation, metrics platforms, and alert tuningKnowledge of security fundamentals relevant to infrastructure operations, including certificate lifecycle management, secrets management, least privilege principles, and vulnerability remediationUnderstanding of network fundamentals applicable to enterprise platforms, including load balancers, reverse proxies, firewall rules, DNS, and TLS