GE Vernova Logo

GE Vernova

System Reliability Engineering Lead

Posted 4 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
152K-228K Annually
Expert/Leader
Remote
Hiring Remotely in USA
152K-228K Annually
Expert/Leader
Leads system reliability engineering for a global grid software SaaS portfolio. Owns cloud infrastructure, platform standardization, SLOs, release governance, incident command, disaster recovery, FinOps, and capacity planning. Serves as final production deployment authority and customer-facing reliability lead. Builds a reliability enablement center, mentors a distributed SRE team, and drives automation, progressive delivery, observability, compliance, and operational excellence across critical utility applications.
The summary above was generated by AI
Job Description SummaryAs the Tech Lead for System Reliability Engineering within the GridOS SaaS Products organization, you will be the hands-on technical authority on production stability for our global grid software SaaS portfolio. You will bridge the gap between architectural design and real-world operations, driving a culture of high reliability and engineering excellence across a distributed team spanning three geographies. You are the "Gatekeeper" for production environments — owning the Change Management process, holding final authority to approve or halt deployments based on system health, and accountable for meeting SLA/SLO targets for critical infrastructure applications serving major North American utility customers.
This is a player-coach role. You will architect and build alongside your team while setting technical direction, mentoring engineers, and serving as the primary customer-facing SRE point of contact. You will own FinOps for the SaaS platform, driving cloud cost optimization and capacity planning as the customer base scales.

Job DescriptionDay 0 — Strategic Provisioning and Design

Standardized Cloud Infra Provisioning

Architect and implement standardized, secure cloud infrastructure provisioning. Drive extreme automation to reduce account provisioning timelines and accelerate customer onboarding to the SaaS platform.

The Golden Path

Define and build the standardized "Middle-Mile" software delivery platform (IDP) using Backstage, ArgoCD, and GitHub Actions. Eliminate bespoke deployment methodologies and establish a single, repeatable path to production.

Follow-the-Sun Architecture

Design and operate the global handover protocols and 24/7 operational coverage model across US, India, and Mexico time zones. Ensure seamless support continuity without graveyard shifts.

Reliability Targets

Establish and own enterprise-wide Service Level Objectives (SLOs) and Service Level Indicators (SLIs) aligned with critical user journeys for global utility customers. Define error budgets and enforce them.

Day 1 — Release Governance and Deployment

Final Approval Authority

Serve as the final technical authority for all production releases. Enforce rigorous change control and validate that all security and performance quality gates are met before any deployment proceeds.

Progressive Delivery

Implement and operate advanced deployment strategies including Canary and Blue/Green rollouts. Build and verify automated rollback capabilities. Hands-on with deployment tooling and pipeline configuration.

SRE Center for Enablement (C4E)

Build and mature the C4E to provide coaching, standardized templates, and repeatable reliability patterns that uplift practices across all product teams. Act as the go-to technical resource for reliability engineering across the organization.

Day 2 — Operational Excellence and Optimization

Incident Command

Serve as the Lead Incident Commander for high-severity (Sev1/Sev2) events. Lead the technical direction, communication, and containment efforts. Available for P1 escalations around the clock.

Blameless Culture

Own the post-incident lifecycle. Facilitate blameless Root Cause Analysis (RCA) to ensure systemic fixes replace recurring operational risks. Build a team culture where incidents drive improvement, not blame.

Business Continuity

Architect and validate end-to-end Backup and Disaster Recovery (DR) strategies, including cross-region failover and automated recovery testing. Hands-on with DR runbook development and execution.

FinOps and Capacity Planning

Own financial operations for the SaaS platform. Drive cloud cost optimization through reserved instances, right-sizing, and waste elimination. Perform long-term capacity planning based on customer growth trajectory and application scaling requirements.

Customer Engagement and Team Leadership

Customer-Facing Accountability

Serve as the primary SRE point of contact for North American utility customers. Own customer satisfaction and NPS for SaaS reliability. Participate in customer-facing reviews, incident communications, and service health reporting. Must meet customer-mandated background check requirements for access to critical infrastructure data and environments.

Player-Coach Team Leadership

Lead a distributed team of 8 SRE engineers across Hyderabad Technical Center and Querétaro, scaling with SaaS application and customer growth. Set technical direction, assign tasks, own team deliverables, and drive day-to-day execution. Mentor engineers on SRE practices, cloud architecture, and operational discipline. Provide performance feedback to the people leader of record. Foster a culture of high performance and continuous learning.



Required Qualifications

Technical Qualifications

  • Cloud Ecosystem: Deep expertise in AWS core services (EC2, EKS, RDS, S3, IAM) and management tools (CloudTrail, CloudWatch)
  • Orchestration: Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures
  • Continuous Delivery: Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows
  • Automation: Proficiency in Infrastructure as Code (IaC) using Terraform and configuration management via Ansible
  • Observability: Hands-on experience with Prometheus, Grafana, observability platforms (Splunk or Datadog), and OpenTelemetry standard to build comprehensive telemetry pipelines
  • FinOps: Demonstrated experience in cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms

Experience and Leadership

  • Overall Experience: 12+ years in software engineering, cloud operations, or infrastructure roles
  • Domain Depth: 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications
  • Technical Leadership: Proven track record of leading distributed engineering teams as a player-coach — setting technical direction while remaining hands-on with architecture, automation, and incident response
  • Operational Discipline: Exceptional troubleshooting skills under pressure and a "Fire Marshal" mindset toward investigation and proactive inspection
  • Customer Engagement: Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
  • Background Check: Must be able to pass customer-mandated background screening for access to critical infrastructure environments


Desired Characteristics

Regulated Environments

  • Practical knowledge of NERC CIP compliance standards in a SaaS context
  • Experience with SOC2, ISO 27001, or IEC 62443 compliance frameworks
  • Familiarity with operating in highly regulated industries such as utilities, financial services, or critical national infrastructure

Certifications

  • AWS Certification: DevOps Engineer — Professional or Solutions Architect — Associate/Professional
  • CKA: Certified Kubernetes Administrator
  • SRE Practitioner Certification
  • AWS FinOps Practitioner or equivalent cloud financial management certification

Additional Information

About the SRE Team: The Grid Software SRE function is a newly established capability supporting the organization's SaaS transformation. The team currently supports Field Damage Assessment (FDA) and Distributed Dynamic Line Rating (DDLR) applications, with the portfolio expanding as GE Vernova Grid Software scales from its initial SaaS customers to a target of 20+ customers by end of 2027. The team operates a follow-the-sun model with engineers based in Hyderabad Technical Center (India) and Querétaro (Mexico).

Why US-Based: North American utility customers operating critical national infrastructure require that production environments and customer data be managed by US-based personnel who have completed customer-mandated background screening. This role exists to meet that requirement while providing hands-on technical leadership to the global SRE team.


Work Schedule

General shift, US business hours. On-call availability required for P1/Sev1 incidents. Follow-the-sun handoff protocols with Hyderabad and Querétaro teams.

Travel Requirements

Up to 10% — customer sites and team locations as needed (estimated 2–4 trips per year)


Additional Information

GE Vernova offers a great work environment, professional development, challenging careers, and competitive compensation. GE Vernova is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other characteristics protected by law.

GE Vernova will only employ those who are legally authorized to work in the United States for this opening. Any offer of employment is conditioned upon the successful completion of a drug screen (as applicable).

Relocation Assistance Provided: No

#LI-Remote - This is a remote position

For candidates applying to a U.S. based position, the pay range for this position is between $151,800.00 and $227,700.00. The Company pays a geographic differential of 110%, 120% or 130% of salary in certain areas. The specific pay offered may be influenced by a variety of factors, including the candidate’s experience, education, and skill set.

Bonus eligibility: discretionary annual bonus.

This posting is expected to remain open for at least seven days after it was posted on September 25, 2026.

Available benefits include medical, dental, vision, and prescription drug coverage; access to Health Coach from GE Vernova, a 24/7 nurse-based resource; and access to the Employee Assistance Program, providing 24/7 confidential assessment, counseling and referral services. Retirement benefits include the GE Vernova Retirement Savings Plan, a tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions, as well as access to Fidelity resources and financial planning consultants. Other benefits include tuition assistance, adoption assistance, paid parental leave, disability benefits, life insurance, 12 paid holidays, and permissive time off.

GE Vernova Inc. or its affiliates (collectively or individually, “GE Vernova”) sponsor certain employee benefit plans or programs GE Vernova reserves the right to terminate, amend, suspend, replace, or modify its benefit plans and programs at any time and for any reason, in its sole discretion. No individual has a vested right to any benefit under a GE Vernova welfare benefit plan or program. This document does not create a contract of employment with any individual.

Similar Jobs

16 Minutes Ago
Remote or Hybrid
172K-301K Annually
Senior level
172K-301K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Architects and operates production-grade agentic AI systems for enterprise identity security. Responsibilities include designing multi-agent orchestration, tool use, planning, memory, retrieval-augmented generation, model integration, evaluation, safety guardrails, observability, and human-in-the-loop controls. The role provides technical leadership through architecture ownership, code reviews, mentoring, and establishing scalable production AI practices.
Top Skills: Anthropic SdkC++Distributed SystemsGoGoogle Ai SdkHybrid SearchJavaMlopsOpenai SdkPythonRagSemantic SearchVector Stores
60K-70K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Participate in a claims specialist development program focused on investigating insurance claims, reviewing medical records, evaluating damages, determining liability and severity, and negotiating settlements with claimants or attorneys. The role includes comprehensive training, mentoring, independent claims inventory management, and opportunities for career development. Candidates should have 0–2 years of experience, strong analytical and negotiation skills, effective communication, attention to detail, and permanent U.S. work authorization.
Entry level
Healthtech • Social Impact • Telehealth
Provide evidence-based psychotherapy to older adults through telehealth. Responsibilities include developing personalized treatment plans, documenting care in an EMR with AI assistance, working independently within a collaborative care team, and participating in paid case consultations. The role requires an independently licensed therapist in Tennessee with a relevant graduate degree and offers flexible scheduling, remote work, clinical support, and part-time or full-time opportunities.
Top Skills: Ai Documentation ToolsElectronic Health Records (Ehr)EmrTelehealth Platforms

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account