Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.
Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Role Summary
As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and
operational maturity of the Filevine platform. You’ll design automation that reduces toil,
strengthen observability, support reliable deployments at scale, and solve production challenges
that keep Filevine running for legal teams across the country.
This role is built for engineers who apply software engineering principles to infrastructure problems, thrive on complex
technical challenges, and are energized by taking ownership of the systems they build and
operate — growing into deeper expertise as their platform knowledge expands
technical challenges, and are energized by taking ownership of the systems they build and
operate — growing into deeper expertise as their platform knowledge expands
Responsibilities
- Design, build, and maintain the monitoring, logging, distributed tracing, dashboards, and
alerting that give teams meaningful visibility into production health. - Build automation, tooling, and CI/CD improvements that increase engineering efficiency,
reduce toil, and support reliable deployments at scale. - Design, implement, and maintain reliable systems for building, deploying, testing, and
operating Filevine products — proactively identifying and resolving reliability,
performance, scalability, and security risks before they impact customers. - Participate in a shared 24/7 on-call rotation, using operational insights to drive automation
and long-term reliability improvements; continuously improve runbooks, documentation,
and engineering standards. - Take ownership of technical initiatives from design through implementation, develop deep
expertise in critical areas of the Filevine platform, and communicate clearly with technical
and business stakeholders.
What we are looking for
- 4+ years of hands-on experience in software engineering, cloud infrastructure, platform
engineering, DevOps, or related technical roles, including at least 2 years in a Site Reliability
Engineering or reliability-focused role. - Working knowledge of distributed systems and how applications, infrastructure, and cloud
services interact in production; demonstrated ability to troubleshoot production issues,
perform root cause analysis, and drive long-term reliability improvements. - Proficiency with Python, Bash, or similar scripting languages; experience building
production tooling, automation, or CI/CD pipelines and deployment automation. - Hands-on experience operating Kubernetes-based workloads and cloud infrastructure in
AWS or a comparable platform, including compute, container orchestration, networking,
IAM, object storage, and cloud-native monitoring. - Experience with Infrastructure as Code tools such as Terraform, Pulumi, or AWS
CloudFormation, and familiarity with modern observability practices including monitoring,
logging, alerting, distributed tracing, and incident response. - Experience using AI-assisted engineering tools to improve productivity, accelerate
troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for
building reliable systems through continuous improvement. - Strong written and verbal communication skills; Bachelor’s degree in Computer Science,
Information Systems, or a related field, equivalent industry certifications, or comparable
professional experience.
Compensation Information: $147,000 - $168,000
The base salary range represents the low and high end of the salary range for this position. The total compensation package for this position will be determined by each individual’s location, qualifications, education, work experience, skills and performance. We believe in the importance of pay equity - the range listed is just one component of Filevine’s total compensation package for employees. This position is also eligible for a paid time off policy, as well as a comprehensive benefits package.
Filevine is an Equal Opportunity Employer. Qualifications for employment, promotion and other terms and conditions of employment are based upon the ability to perform the job. Equal-employment opportunities are provided to all applicants and employees without regard to race, creed, religion, color, age, national origin, sex, disability, veteran status, or other legally protected class. Filevine is committed to providing reasonable accommodations for qualified individuals with disabilities. If you need assistance or accommodation due to disability, or if you have concerns related to Filevine’s equal employment opportunities, you may contact us at [email protected]
Cool Company Benefits:
- A dynamic, rapidly growing company, focused on helping organizations thrive
- Medical, Dental, & Vision Insurance (for full-time employees)
- Competitive & Fair Pay
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
Privacy Policy Notice
Filevine will handle your personal information according to what’s outlined in our Privacy Policy.
Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.
Similar Jobs
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills:
AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills:
AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills:
Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
What you need to know about the Charlotte Tech Scene
Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.
Key Facts About Charlotte Tech
- Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
- Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
- Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
- Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus


.png)
