CRC Insurance Services Logo

CRC Insurance Services

Site Reliability Engineer

Reposted 29 Days Ago
Be an Early Applicant
In-Office
Charlotte, NC, USA
Senior level
In-Office
Charlotte, NC, USA
Senior level
Lead SRE role owning enterprise observability and monitoring strategy across cloud and applications. Architect Dynatrace-to-ServiceNow pipelines, alert correlation, paging integration, CMDB/service mapping, and automation for self-healing. Drive SRE practices (SLIs/SLOs/error budgets), mentor teams, reduce alert noise, improve MTTR, and standardize monitoring and incident workflows.
The summary above was generated by AI

The position is described below. If you want to apply, click the Apply button at the top or bottom of this page. You'll be required to create an account or sign in to an existing one.

If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).

Regular or Temporary:

Regular

Language Fluency:  English (Required)

Work Shift:

1st Shift (United States of America)

Please review the following job description:

Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)
We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.
This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.

Key Responsibilities

Strategic Leadership & Decision-Making

  • Define and own the enterprise monitoring and SRE observability strategy

  • Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture

  • Evaluate and recommend tooling, integration patterns, and platform direction

  • Drive decisions on alerting philosophy, noise reduction, and signal quality improvement

Platform Ownership & Architecture

  • Architect and standardize end-to-end monitoring and SRE pipelines:

    • Dynatrace → ServiceNow incident lifecycle

    • Alert correlation, deduplication, and prioritization

    • Integration with paging systems (PagerDuty, SMS, voice, Teams)

  • Establish best practices for:

    • Event ingestion and enrichment

    • Incident routing and automated assignment

    • Integration with CMDB and service mapping

Site Reliability Engineering (SRE) Leadership

  • Lead adoption of SRE principles, including:

    • SLIs, SLOs, and error budgets

    • Reliability engineering practices across services

    • Proactive monitoring and resilience design

  • Champion a shift from reactive operations to proactive reliability engineering

  • Influence application and platform teams to build observable, resilient systems by design

Automation & Self-Healing Enablement

  • Drive development of automated remediation and self-healing capabilities

  • Leverage Dynatrace workflows, Azure services, and automation frameworks to:

    • Reduce manual incident handling

    • Eliminate repeatable operational tasks

    • Minimize unnecessary paging

ServiceNow & Observability Integration Leadership

  • Own integration between Dynatrace and ServiceNow ITSM/ITOM, including:

    • Incident, Event Management, and CMDB alignment

    • Service mapping and dependency visibility

    • Governance for application/service tagging

  • Define standards for:

    • Automated incident creation and resolution

    • Priority assignment and routing logic

    • Monitoring-to-ITSM data synchronization

Team Leadership & Cross-Functional Influence

  • Provide technical leadership and mentorship across SRE, platform, and application teams

  • Act as a central point of coordination between engineering, cloud, and ITSM teams

  • Lead workshops and working sessions to:

    • Drive monitoring standardization

    • Align teams on reliability practices

    • Influence upstream architectural decisions

Operational Excellence

  • Establish KPIs and drive improvement in:

    • Incident response and resolution times

    • Alert quality and paging effectiveness

    • Monitoring coverage across critical services

  • Provide leadership with clear visibility into service health and reliability trends

Required Qualifications

  • 7+ years in Site Reliability Engineering, monitoring, or production engineering

  • Proven experience in a technical leadership or lead engineer role

  • Deep hands-on experience with:

    • Dynatrace (or equivalent observability platforms)

    • Microsoft Azure (IaaS, PaaS, networking, identity)

    • ServiceNow ITSM / ITOM (incident, event management, CMDB)

  • Demonstrated ability to:

    • Design and lead enterprise monitoring/SRE architectures

    • Drive platform and tooling decisions

    • Integrate observability, ITSM, and paging solutions

Preferred Qualifications

  • Experience leading SRE or observability transformation initiatives

  • Strong expertise with Dynatrace–ServiceNow integrations

  • Experience modernizing or consolidating paging/on-call tooling

  • Familiarity with:

    • Azure-based SRE tooling or AI-assisted operations

    • Automation frameworks (GitHub Actions, Runbooks, etc.)

    • Infrastructure as Code (Terraform, ARM, Bicep)

Success Metrics

  • Reduction in alert noise and unnecessary paging

  • Improved incident routing accuracy and MTTR

  • Increased adoption of self-healing and automated workflows

  • Strong alignment between monitoring, CMDB, and service ownership

  • Enterprise-wide adoption of SRE and monitoring standards

General Description of Available Benefits for Eligible Employees of CRC Group: At CRC Group, we're committed to supporting every aspect of teammates' well-being – physical, emotional, financial, social, and professional. Our best-in-class benefits program is designed to care for the whole you, offering a wide range of coverage and support. Eligible full-time teammates enjoy access to medical, dental, vision, life, disability, and AD&D insurance; tax-advantaged savings accounts; and a 401(k) plan with company match. CRC Group also offers generous paid time off programs, including company holidays, vacation and sick days, new parent leave, and more. Eligible positions may also qualify for restricted stock units and/or a deferred compensation plan.

CRC Group supports a diverse workforce and is an Equal Opportunity Employer that does not discriminate against individuals on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status or other classification protected by law. CRC Group is a Drug Free Workplace.

EEO is the Law   Pay Transparency Nondiscrimination Provision   E-Verify

HQ

CRC Insurance Services Charlotte, North Carolina, USA Office

Charlotte, NC, United States

Similar Jobs

23 Days Ago
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
5 Days Ago
Hybrid
Charlotte, NC, USA
119K-187K Annually
Senior level
119K-187K Annually
Senior level
Fintech • Financial Services
Lead SRE responsible for platform stability, resiliency, performance, and security. Define SLIs/SLOs, drive observability and automation, lead incident response and toil reduction, collaborate across teams, and enforce reliability standards for cloud and on‑prem environments.
Top Skills: AngularAppdynamicsAWSAzureGlassboxGrafanaJavaJavaScriptJSONKubernetesLinuxNode.jsOpenshift Container PlatformOpentelemetryPcfPksPrometheusPythonRubyShell ScriptingSplunkSplunk ObservabilityVMwareWindows
3 Days Ago
In-Office
Huntersville, NC, USA
100K-150K Annually
Senior level
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Operate and improve the reliability, availability, performance, and operational excellence of large-scale distributed systems. Build automation and tooling, manage Linux and Kubernetes environments, design CI/CD pipelines, implement observability, troubleshoot production systems, and reduce operational toil. Preferred work includes defining SLOs and error budgets, chaos engineering, cloud operations, capacity planning, performance engineering, load testing, and service mesh implementation.
Top Skills: AWSAzureChaos MonkeyCi/CdConsulDockerEfkElkGCPGoGrafanaGremlinIstioJavaKubernetesLinkerdLinuxLitmusOpentelemetryPrometheusPython

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account