X, The Moonshot Factory Logo

X, The Moonshot Factory

Senior Machine Learning Data Engineer (DataOps), Materra

Reposted One Month Ago
In-Office
Mountain View, CA
166K-244K Annually
Mid level
In-Office
Mountain View, CA
166K-244K Annually
Mid level
Build and consolidate data infrastructure for ML training: design automated ETL/ELT pipelines, implement DataOps practices (validation, monitoring, anomaly detection), integrate annotation workflows, manage dataset versioning and storage, and collaborate with ML engineers and operations to produce high-quality, reproducible training datasets.
The summary above was generated by AI

About the team:

Materra is on a mission to radically reduce global waste and move to a true circular economy. The team has developed technology that identifies waste material at the molecular level—starting with plastics. Materra works with industry partners to improve the way recycling centers process plastics using AI and robotics, to make recycling more affordable and scalable.
About the Role
We are looking for a Senior Machine Learning Data Engineer (DataOps) to build and unify the data infrastructure that powers our model training pipelines. In this role, you will lead the effort to consolidate fragmented data sources into a cohesive, high-quality data foundation.
Your primary focus will be designing automated ingestion pipelines, establishing data quality validation frameworks, and managing dataset versioning to support our machine learning training loops. You will bridge the gap between operations, remote annotation teams, and machine learning engineers to ensure our models are trained on reliable, well-structured data.
Key Responsibilities

  • Architect and build automated ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) data pipelines to aggregate, clean, and harmonize data from disparate sources, databases, and operational ingestion flows.
  • Implement DataOps practices, including data quality monitoring, automated schema validation, and anomaly detection to catch corrupt or mislabeled data early.
  • Standardize and integrate third-party annotation workflows and remote labeling feeds into unified datasets ready for model training.
  • Design and maintain dataset versioning and storage systems to allow reproducible machine learning experiments and seamless data retrieval.
  • Collaborate with machine learning engineers and operations teams to translate raw material, form factor, and sensor metadata into structured training features.

Requirements

  • Education: Degree in Computer Science, Data Engineering, Software Engineering, or a related technical field.
  • Data Engineering & Architecture: 5+ years  experience building scalable data pipelines, managing relational and non-relational databases, and unifying fragmented data storage systems.
  • Modern Python Proficiency: Expertise in Python and data manipulation libraries (e.g., Pandas, NumPy, or SQL).
  • Data Quality & DataOps: Practical experience implementing automated data validation, quality control frameworks, and dataset versioning practices.
  • ML Data Lifecycle Understanding: Hands-on experience structuring datasets specifically for machine learning workflows, including handling annotations, metadata tracking, and training set curation.

Preferred Skills

  • Google Cloud Ecosystem: Hands-on experience with Google Cloud platform tools (e.g., BigQuery, Cloud Storage, Dataflow, Dataproc, Vertex AI Data Pipelines).
  • Workflow Orchestration: Experience managing pipelines using Google Cloud Composer or equivalent orchestration frameworks (e.g., Apache Airflow, Prefect, Dagster).
  • Multimodal / Unstructured Data: Experience handling mixed data types, including image datasets, sensor metadata, and unstructured physical property records.
  • Annotation Platform Integration: Familiarity with data labeling platforms, human-in-the-loop workflows, or integrating third-party annotation APIs.
  • Validation & Versioning Tooling: Exposure to data quality and ML versioning tools (e.g., Great Expectations, DVC, or TFX/Data Validation).

The US base salary range for this full-time position is $164,000 - $261,000 + bonus + equity + benefits. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.
Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.

Similar Jobs

3 Minutes Ago
In-Office or Remote
United States
85K-193K Annually
Senior level
85K-193K Annually
Senior level
Automotive
Design, build, and operate a global observability platform across hybrid cloud and on-premises environments. Responsibilities include developing monitoring pipelines, infrastructure-as-code, reliability automation, SLI/SLO frameworks, performance optimization, incident response, root-cause analysis, and production troubleshooting. The role partners with engineering teams to improve system resilience, reduce toil, integrate AI/ML for anomaly detection, and establish observability best practices. It also provides technical mentorship and guidance.
Top Skills: Ai/MlAmazon Web ServicesCC++Ci/CdDatadogDockerDynatraceElkGoGoogle Cloud PlatformJ2EeJavaKafkaKubernetesMicroservicesAzureNagiosNew RelicNoSQLOpentofuPrometheusPythonRestful ApisScalaSensuSplunkSpring BootSQLTcp/IpTerraform
4 Minutes Ago
In-Office or Remote
United States
85K-193K Annually
Mid level
85K-193K Annually
Mid level
Automotive
Develop and maintain global monitoring and observability platforms using Go, JavaScript, GCP, Kubernetes, OpenTelemetry, PostgreSQL, and Terraform. Improve reliability, scalability, performance, security, and disaster recovery for cloud services. Responsibilities include troubleshooting production systems, capacity planning, automation, on-call support, incident postmortems, code reviews, documentation, and vulnerability assessments.
Top Skills: Document DatabasesDynatraceGoGoogle Cloud PlatformInfrastructure As CodeJavaScriptKubernetesOpentelemetryPostgresRelational DatabasesTerraform
5 Minutes Ago
In-Office
100K-120K Annually
Junior
100K-120K Annually
Junior
Information Technology • Software
Coordinate interview scheduling, feedback, next steps, onsite logistics, executive interviews, and candidate communications across the company. The role ensures a high-quality candidate experience, manages overlapping interview processes, maintains hiring momentum, and identifies opportunities to improve recruiting operations and logistics.

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account