Improve AWS production infrastructure reliability, observability, performance, and operational maturity. Build Terraform infrastructure, enhance CI/CD, automate operational work, manage incident response and on-call operations, lead postmortems, improve application resilience, support capacity planning and database reliability, and collaborate on security hardening and compliance. Mentor engineers and promote reliability practices across the organization.
Senior Site Reliability Engineer
Improve the reliability, performance, and operational maturity of a platform that supports the future of educational fundraising.
CONTRACT-TO-HIRE REMOTE - UNITED STATES SEATTLE / WEST COAST PREFERRED
About the role
Our client is looking for a hands-on Senior Site Reliability Engineer to improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, this role will focus on strengthening our production environment: improving observability, automating infrastructure and operational work, enhancing incident response, and partnering with product engineers to build resilient systems. This is a high-impact role for someone who understands both infrastructure and application development. You will work across the stack, contribute code when appropriate, and help ensure our systems remain secure, scalable, and dependable as we grow.
What you'll do
• Operate, maintain, and improve our production infrastructure in AWS.
• Build and maintain infrastructure as code using Terraform.
• Improve monitoring, alerting, dashboards, and service-level indicators using New Relic or comparable observability platforms.
• Reduce alert noise and build systems that identify problems before customers are affected.
• Participate in the 24/7 on-call rotation and help coordinate the response to production incidents.
• Lead blameless postmortems and ensure corrective actions result in durable improvements.
• Partner with product engineers to diagnose performance and reliability issues throughout the application stack.
• Improve application resilience through appropriate use of timeouts, retries, queuing, backpressure, and idempotency.
• Improve CI/CD pipelines and deployment practices using platforms such as GitHub Actions, GitLab CI, or CircleCI.
• Automate repetitive operational work and reduce engineering toil.
• Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation.
• Support capacity planning, performance testing, database reliability, and production-readiness reviews.
• Collaborate with Security and Engineering teams on infrastructure hardening, access controls, logging, and compliance-related operational practices.
• Mentor engineers and promote effective reliability practices across the Engineering organization
What we're looking for
• 10+ years of overall software engineering, infrastructure, or systems experience, including at least 5 years in an SRE, Platform Engineering, DevOps, or production operations role.
• Previous professional software development experience and the ability to read, debug, and contribute to application code. • Strong, hands-on experience operating production workloads in AWS. • Experience building and maintaining infrastructure with Terraform or a similar infrastructure-as-code tool.
• Strong observability skills using New Relic, Datadog, or another modern monitoring platform. • Experience with incident response, on-call operations, postmortems, and production troubleshooting.
• Experience building or maintaining CI/CD pipelines.
• Working knowledge of networking, Linux, distributed systems, and relational databases.
• Strong judgment when balancing immediate operational needs with long-term maintainability.
• Clear communication skills and the ability to collaborate effectively across engineering disciplines.
• A track record of using automation to improve reliability and create leverage for other engineers.
Bonus points
• Experience with Ruby or Ruby on Rails.
• Strong PostgreSQL administration or performance-tuning experience.
• Experience operating enterprise SaaS products at scale.
• Familiarity with SLOs, SLIs, error budgets, capacity modeling, and load testing.
• Experience with payments, fintech, or other highly regulated systems.
• Experience supporting SOC 2 or similar security and compliance programs. Role details
• Contract-to-hire.
• Remote within the United States.
• Seattle-area or West Coast candidates are preferred to support occasional in-person collaboration, but exceptional candidates elsewhere should also be considered.
• Participation in a shared on-call rotation is required.
Similar Jobs
51 Minutes Ago
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Leads business development and strategic account growth with pharmaceutical and biotechnology partners across the U.S. West and Central regions. Responsibilities include acquiring customers, building pipelines in Salesforce, developing proposals, advancing precision medicine and genomics opportunities, negotiating and closing projects, and collaborating with scientific, commercial, and operational teams. The role focuses on NGS-based assays, biomarker discovery, clinical trial solutions, and companion diagnostics, with approximately 30% domestic and international travel.
Top Skills:
Biomarker DiscoveryClinical Trial AssaysCompanion Diagnostics (Cdx)CtdnaGenomic AssaysLiquid BiopsyNext-Generation Sequencing (Ngs)SalesforceSophia Ddm
Other • Utilities
Senior enterprise sales leader responsible for winning Fortune 1000 new logos and expanding strategic accounts. Build C‑suite relationships, architect tailored technology solutions, lead complex multi-stakeholder deals, manage pipeline and forecasting via CRM, and coordinate cross-functional execution for global, high-value clients.
Top Skills:
CRM
Machine Learning • Payments • Security • Software • Financial Services
Provides enterprise technology support through phone, ticketing, and potentially chat channels. Troubleshoots hardware, software, network, access, and application issues; manages incidents and escalations; documents resolutions in ServiceNow or similar ITSM tools; analyzes trends and operations data; supports knowledge management, training, process improvement, and automation; and coaches junior analysts while maintaining customer service, security, and confidentiality standards.
Top Skills:
Call Center TechnologiesHardware InfrastructureItsmNetworkingServicenow
What you need to know about the Charlotte Tech Scene
Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.
Key Facts About Charlotte Tech
- Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
- Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
- Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
- Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus



