Metasys Logo

Metasys

Site Reliability Engineer Internship

Posted 4 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
The summary above was generated by AI
Overview: Reliability and Operational Excellence

The Site Reliability Engineer (SRE) is responsible for the ultimate stability, performance, and scalability of our entire integrated supply chain e-commerce platform. You will apply software engineering principles to operations, ensuring the high availability and resilience of the customer-facing e-commerce storefront, internal SaaS tools (WMS, OMS), and specialized AI agent services.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will be the champion of uptime, performance, and automated operations for systems handling the critical MES → WMS → OMS flow.

  • Availability & SLO Management: Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for core business processes and all application layers. Manage the platform's overall Service Level Agreement (SLA).

  • Observability & Alerting: Architect, maintain, and optimize the comprehensive observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry). Develop high-fidelity alerting and ensure distributed tracing across the NestJS modular monolith and associated data stores (PostgreSQL, Redis).

  • Incident Response & Review: Own the incident response workflow, ensuring rapid triage, mitigation, and root cause analysis. Conduct thorough post-incident reviews to drive continuous improvement and eliminate recurring toil.

  • Scalability & Capacity Planning: Optimize auto-scaling policies for all services running on Docker containers. Conduct capacity planning based on business projections, especially for peak e-commerce and manufacturing load.

  • Disaster Recovery (DR): Design, implement, and regularly test Disaster Recovery procedures, including backup and restoration workflows for PostgreSQL 15 using tools like pgBackRest.

  • Automation: Eliminate operational toil through automation, managing infrastructure-as-code (Terraform) and CI/CD pipelines (Makefile).

Required Technologies & Tools

Candidates must possess deep experience in cloud operations, observability, and infrastructure automation:

  • Observability Stack: Prometheus, Grafana, Loki, Tempo, OpenTelemetry (mandatory).

  • Infrastructure & Platform: Terraform, Docker, Traefik, Oracle Cloud Free VMs (or equivalent public cloud).

  • Data & Resilience: PostgreSQL (Deep knowledge), Redis, pgBackRest.

  • Automation: Strong scripting skills (Python/Bash) and experience with CI/CD tools and Makefile.

  • Methodology: Expert knowledge of SRE principles, toil reduction, and error budgeting.

AI Agent Focus

You will be responsible for the operational reliability of the emerging AI layer.

  • Agent Reliability: Implement specialized monitoring and logging for the AI agent services, ensuring LLM integrations and multi-agent systems (built with frameworks like LangChain) meet defined performance and availability SLOs.

  • Resource Optimization: Efficiently manage resource allocation for computationally intensive AI workloads to maintain platform stability and cost-efficiency.

Success Metrics & Career Path

Performance will be measured by:

  • Uptime/Availability: Achieving defined SLAs/SLOs across the platform.

  • MTTR: Reduction in Mean Time To Recover from production incidents.

  • Toil Reduction: Measured percentage reduction in manual, repetitive operational tasks through automation.

Mentorship Structure: Reports to the Head of Technology/CTO, working collaboratively with DevSecOps and development teams to ensure software is designed for reliability.

Similar Jobs

An Hour Ago
Remote
United States
110K-130K Annually
Mid level
110K-130K Annually
Mid level
AdTech • Artificial Intelligence • Big Data • Digital Media • eCommerce • Machine Learning • Marketing Tech
Plan and execute regional and large-scale industry events, design executive-level activations, manage partner marketing and budgets, track lead flow and ROI, collaborate with Sales/Product/Global Marketing, and standardize event and partner marketing processes.
Top Skills: AsanaCRMRoi DashboardsSpreadsheets
4 Hours Ago
Easy Apply
Remote or Hybrid
Easy Apply
166K-196K Annually
Senior level
166K-196K Annually
Senior level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Own and execute the multi-year telematics hardware platform roadmap end-to-end. Drive commercial outcomes for gateways, sensors, cabling and accessories through field validation, launch, lifecycle management, and cross-functional alignment. Engage customers and sales on strategic deals, use market and fleet data to prioritize, and mentor junior PMs while ensuring hardware strategy supports revenue, margin, and company growth.
Top Skills: Can BusCloudEdge AiFccIotJ1939Obd-IiPtcrbSensorsTelematicsVehicle GatewaysWireless
5 Hours Ago
Remote
United States
253K-275K Annually
Senior level
253K-275K Annually
Senior level
Blockchain • Software • Cryptocurrency • Web3
Design, build, test, and deploy smart contracts and decentralized applications. Maintain blockchain integrations and backend services, optimize for security and gas efficiency, contribute to architecture and technical strategy, conduct code reviews, mentor junior engineers, and collaborate with product, frontend, and security teams.
Top Skills: AnchorAvalancheBnb ChainCi/CdCloud InfrastructureDaosDatabasesDefiEthereumEthers.JsFoundryGitGoHardhatNftsNode.jsPolygonPythonRustSmart ContractsSolanaSolidityTruffleTypescriptWallet IntegrationsWeb3.Js

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account