DDN Storage Logo

DDN Storage

Staff Engineer

Posted Yesterday
Be an Early Applicant
Remote or Hybrid
Hiring Remotely in North Carolina, USA
Senior level
Remote or Hybrid
Hiring Remotely in North Carolina, USA
Senior level
Design and build a software-defined storage cluster management platform, focusing on high-performance Go components and production-scale observability. Develop and optimize telemetry pipelines using OpenTelemetry, Prometheus, and VictoriaMetrics; troubleshoot Linux-based clustered storage systems; write tests, conduct code reviews, document standards, support Scrum delivery, and participate in a global on-call rotation.
The summary above was generated by AI

We are looking for a hands-on Staff Software Engineer who will help design and build our Cluster management platform. DDN – Infinia engineers come from diverse backgrounds, and we consider them to be in the highest tier of technology game changers globally. They have created, architected and pioneered numerous advancements in data enabling technologies, and delivered groundbreaking ideas which have shaped and transformed the storage industry.

What will you bring to DDN
  • 8+ years of backend development experience (with a target of matching our senior engineering standards), including deep proficiency in Go for building high-performance, low-overhead system components.

  • Production-Scale Observability Expertise: Hands-on experience implementing and operating telemetry pipelines for software-defined clustered, distributed, or cloud-native solutions.

  • Deep Mastery of the Metrics Stack: Strong technical knowledge of Prometheus (operators, alerting rules, scraping mechanics) and the VictoriaMetrics stack for long-term, high-cardinality storage.

  • Telemetry Industry Standards: Practical experience with the OpenTelemetry (OTel) ecosystem, including custom OTel collector configurations, instrumentation SDKs, and data processing.

  • Systems-Level Troubleshooting: A solid understanding of Linux networking, filesystems, and how clustered storage applications behave under heavy I/O workloads.

  • Collaborative Mindset: Proven ability to work effectively across geographically distributed teams, driving technical clarity through code reviews and clear documentation.

What you have achieved…
  • Built and Maintained Telemetry Pipelines: A proven track record of developing or extending proprietary Go components to efficiently ingest, process, and forward massive streams of metrics, logs, and traces.

  • Optimized Resource Consumption: Experience managing the CPU and memory footprint of monitoring agents to ensure they do not compete with core storage data paths.

  • Strong Testing & Regression Habits: Dedication to writing robust unit and integration tests to ensure telemetry components remain stable during live cluster upgrades.

  • Independent Feature Delivery: A proven ability to take ownership of complex technical initiatives and independently make progress in a fast-paced environment.

  • Concept Visualization: Ability to turn abstract cluster state data into logical, well-structured telemetry frameworks that bring visibility to complex system scenarios.

What will you be doing
  • Execute the Telemetry Architecture: Take on core projects within the observability domain, ensuring seamless integration between our proprietary Go infrastructure and open-source tools.

  • Optimize the Observability Stack: Help design and refine how OpenTelemetry, Prometheus, and VictoriaMetrics handle the massive metrics volume generated by our storage cluster.

  • Drive Code Excellence: Act as a key technical contributor to our Software Defined Storage control plane, writing clean, performant Go code and providing rigorous code reviews.

  • Full Lifecycle Engineering: Participate actively within the Scrum model—from initial design and coding to automated testing, usability reviews, and release.

  • Document and Standardize: Ensure our telemetry frameworks are well-documented, making it easy for other engineering teams to instrument their components properly.

  • Global Support Rotation: Contribute to our global team on-call rotation, leveraging your own observability tools to provide high-level technical support for our distributed footprint.

Similar Jobs

Yesterday
Remote
USA
200K-300K Annually
Senior level
200K-300K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Own production operations for workplace AI platforms, including configuration, integrations, observability, secrets, access controls, model and prompt changes, benchmarking, and incident response. Translate technical findings into governance, risk, compliance, and training materials while collaborating with Security, Legal, HR, and business stakeholders. Investigate failures, implement mitigations, document root causes, and maintain reliable, secure AI services.
Top Skills: Access ControlAi/Llm PlatformsAlertingAPIsInfrastructure As CodeJSONLoggingMetricsMonitoring And ObservabilityPrompt EngineeringRpa/Workflow EnginesSecrets ManagementYaml
2 Days Ago
Remote
United States
242K-288K Annually
Senior level
242K-288K Annually
Senior level
Healthtech • Social Impact • Software • Telehealth
Develop and operate production ML and AI systems powering patient experiences such as matching, ranking, recommendations, onboarding, and personalization. Build reusable ML infrastructure, pipelines, evaluation, observability, and model-serving capabilities. Provide technical leadership across the ML lifecycle, guide architecture, evaluate custom versus foundation-model approaches, and partner with ML and Patient Engineering teams to deliver reliable, scalable AI products.
Top Skills: Artificial IntelligenceExperimentationFeature EngineeringFoundation ModelsGenerative AiMachine LearningMatching SystemsMl InfrastructureMl PipelinesModel EvaluationModel ServingModel TrainingObservabilityPersonalization SystemsPythonRanking SystemsRecommendation SystemsSearch Systems
Yesterday
Remote
US
159K-254K Annually
Senior level
159K-254K Annually
Senior level
Cloud • Fintech • Food • Information Technology • Software • Hospitality
Designs, builds, and ships scalable full-stack features for Toast’s restaurant and retail technology platform. Responsibilities include developing resilient backend services and APIs, responsive web interfaces, distributed systems, and high-throughput enterprise solutions using Java, Kotlin, React, and TypeScript. The role requires solving ambiguous technical problems, ensuring system reliability and observability, applying engineering best practices, collaborating cross-functionally, and mentoring engineers.
Top Skills: DynamoDBGraphQLJavaKotlinPostgresReactRestTypescript

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account