Leads Site Reliability Engineering and systems operations for critical consumer-facing platforms. Responsibilities include improving reliability, resilience, observability, scalability, and operational readiness; establishing SLOs, SLIs, and error budgets; leading major incident response and root-cause analysis; developing monitoring, automation, self-healing, and recovery capabilities; managing vendor dependencies; and mentoring engineering teams across a regulated financial-services environment.
About this role:
Wells Fargo is seeking a Lead Site Reliability Engineer (SRE) / Lead Systems Operations Engineer within the Consumer Technology (CT) organization. This role will provide technical leadership for operational excellence, platform reliability, resiliency, observability, and support readiness across critical consumer-facing applications and platforms.
The Lead SRE will serve as a senior technical leader responsible for driving reliability engineering practices, reducing operational risk, improving service availability, and enabling scalable platform operations. This role will partner closely with Application Development, Platform Engineering, Infrastructure teams, Shared Services, and External Vendors to ensure highly resilient, supportable, and observable solutions.
The ideal candidate combines deep technical expertise with strong operational leadership and will play a critical role in advancing Site Reliability Engineering practices across the organization.
In this role, you will support:
Reliability Engineering & Platform Stability
Incident Management & Operational Excellence
Observability & Monitoring
Automation & Engineering Excellence
Vendor & Dependency Management
Technical Leadership
10 Sep 2026
*Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements.
Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates, including women, persons with disabilities, aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo .
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.
Wells Fargo is seeking a Lead Site Reliability Engineer (SRE) / Lead Systems Operations Engineer within the Consumer Technology (CT) organization. This role will provide technical leadership for operational excellence, platform reliability, resiliency, observability, and support readiness across critical consumer-facing applications and platforms.
The Lead SRE will serve as a senior technical leader responsible for driving reliability engineering practices, reducing operational risk, improving service availability, and enabling scalable platform operations. This role will partner closely with Application Development, Platform Engineering, Infrastructure teams, Shared Services, and External Vendors to ensure highly resilient, supportable, and observable solutions.
The ideal candidate combines deep technical expertise with strong operational leadership and will play a critical role in advancing Site Reliability Engineering practices across the organization.
In this role, you will support:
Reliability Engineering & Platform Stability
- Lead reliability initiatives across critical business platforms and customer journeys.
- Establish and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budget practices.
- Improve platform resilience through automation, self-healing capabilities, capacity planning, and fault-tolerant designs.
- Identify and eliminate single points of failure across applications, infrastructure, vendor integrations, and customer flows.
- Champion engineering solutions that improve availability, scalability, recoverability, and operational maturity.
Incident Management & Operational Excellence
- Serve as a technical lead during major production incidents, providing coordination, technical direction, and recovery leadership.
- Drive improvements in Mean Time to Detect (MTTD), Mean Time to Diagnose (MTTDiag), and Mean Time to Recover (MTTR).
- Lead Root Cause Analysis (RCA) efforts and ensure corrective actions are implemented and tracked to completion.
- Identify recurring operational patterns and develop preventive solutions to reduce production incidents.
- Develop and maintain incident playbooks, recovery procedures, and operational readiness standards.
Observability & Monitoring
- Lead enterprise observability initiatives leveraging Splunk, Grafana, GCP Monitoring, AppDynamics, and related platforms.
- Define monitoring standards, alerting strategies, dashboards, and customer journey observability solutions.
- Partner with application and infrastructure teams to improve telemetry, logging, tracing, and synthetic monitoring capabilities.
- Develop actionable operational metrics and executive-level reliability reporting.
Automation & Engineering Excellence
- Drive automation strategies that reduce manual effort and improve operational consistency.
- Design and implement self-service operational capabilities and automated recovery solutions.
- Utilize AI-assisted tools and engineering practices to improve incident detection, diagnosis, and remediation workflows.
- Promote Infrastructure-as-Code (IaC), CI/CD best practices, and platform engineering principles.
Vendor & Dependency Management
- Partner with internal and external service providers to improve reliability, support responsiveness, and recovery performance.
- Evaluate vendor operational performance and contribute to service improvement initiatives.
- Establish and monitor operational readiness expectations for critical vendor dependencies.
- Drive resilience planning and support strategies for third-party integrations.
Technical Leadership
- Provide technical leadership and mentorship to SREs, Systems Operations Engineers, and Platform Support Engineers.
- Lead technical reviews, operational readiness assessments, and production support governance activities.
- Influence architecture decisions to ensure supportability, resiliency, observability, and operational sustainability.
- Collaborate with engineering leaders to establish and mature Site Reliability Engineering practices across Consumer Technology.
- 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 5+ years of Site Reliability Engineering, Platform Engineering, Production Support, or equivalent experience demonstrated through work experience, military experience, training, or education.
- 5+ years supporting mission-critical production applications in large enterprise environments.
- 3+ years leading major incident management, operational support, or reliability engineering initiatives.
- 3+ years of experience with observability and monitoring platforms such as Splunk, Grafana, AppDynamics, Dynatrace, GCP Monitoring, or similar technologies.
- 2+ years of experience driving automation, operational improvements, and reliability initiatives.
- 3+ years of experience supporting distributed systems, cloud-based platforms, infrastructure, networking, and application architectures.
- 1+ year of experience supporting highly regulated or customer-facing financial services platforms
- Consumer Technology, Credit Card, Payments, Lending, or Digital Banking experience.
- Experience implementing Site Reliability Engineering (SRE) principles, SLOs, SLIs, and Error Budgets.
- Experience with DevOps, CI/CD, Infrastructure-as-Code, and cloud-native architectures.
- Experience with AI-assisted engineering, incident management automation, or observability platforms.
- Strong executive communication and stakeholder management skills.
- Experience leading cross-functional technical teams without direct authority.
- Experience supporting vendor governance and third-party operational readiness initiatives.
- Relocation assistance is not provided for this position
- Visa sponsorship is not available for this position
- Position requires onsite presence at one of the posted Wells Fargo locations.
- 401 W. Las Collinas Blvd, Irving, Texas
- 300 S. Brevard St. Charlotte, North Carolina
- 2600 S. Price Rd. Chandler, Arizona
10 Sep 2026
*Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements.
Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates, including women, persons with disabilities, aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo .
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.
Wells Fargo Charlotte, North Carolina, USA Office
355 W Martin Luther King, Jr BLVD, Charlotte, NC, United States, 28202
Similar Jobs at Wells Fargo
Fintech • Financial Services
Lead day-to-day Redis and OpenShift platform operations, including cluster builds, maintenance, upgrades, monitoring, troubleshooting, incident response, and decommissioning. Develop automation and remediation workflows using Python, Bash, GitOps, and AI-assisted tools. Build observability solutions, improve reliability and resilience, enforce security and compliance controls, conduct root-cause analysis, and drive operational improvements across engineering, SRE, security, and development teams. Participate in on-call rotations and work onsite.
Top Skills:
Ai AgentsAnsibleBashElasticGitGitopsGrafanaJIRAKubernetesLinuxMcpOpenshiftPrometheusPythonRedisSdlcSplunk
Fintech • Financial Services
Leads systems operations and site reliability engineering for modernized, cloud-native payment platforms. Responsibilities include defining non-functional requirements, capacity and performance testing, observability, SLOs, resilience validation, production readiness, cutovers, runbooks, on-call support, chaos engineering, root-cause analysis, and service stabilization. The role collaborates with engineering, architecture, and service operations teams to build resilient, scalable, and supportable event-driven payment systems.
Top Skills:
BlazemeterChaos MonkeyCi/CdCloud-Native ArchitectureDistributed TracingEvent-Driven ArchitectureKubernetesLoggingMetricsMongoDBRedisResilience4JSlo ToolingSpring BootSpring Webflux
Fintech • Financial Services
Leads systems operations and reliability initiatives across platform, application, and engineering teams. Responsibilities include infrastructure planning, production issue resolution, technical change decisions, observability, incident and change management, automation, cloud-native architecture, Kubernetes and OpenShift operations, resiliency engineering, and technical debt remediation. The role drives self-service and operational toil reduction through scripting, infrastructure automation, SRE practices, and Generative AI or agent development, with occasional on-call coverage.
Top Skills:
Ai AgentsAnsibleAPIsAppdynamicsAzureBashBigpandaElkGCPGenerative AiGitGrafanaKubernetesPowershellPrometheusPythonRed Hat Enterprise LinuxRed Hat Openshift Container PlatformSplunk ObservabilityTerraformThousandeyes
What you need to know about the Charlotte Tech Scene
Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.
Key Facts About Charlotte Tech
- Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
- Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
- Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
- Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

