Padmi
BLD Talent logo
BLD Talent

technical recruiting · engineering hiring

Data Infra Engineer

San Francisco Bay Area · Onsite$110k–$175k/yrPosted 7 days ago
DataJuniorFull Time
Apply at BLD Talent

Opens the source posting on jobs.gem.com

Source description

About the role

View original

Company Overview

We’re building intelligent robotic arms that can learn new skills in hours, not months. Backed by Y Combinator and top-tier Silicon Valley investors, we’re turning physical AI into reality, helping industries facing critical labor shortages (manufacturing, logistics, and more) automate back-of-house tasks like packaging, kitting, and assembly.

Our flagship robot combines affordable robotic hardware with cutting-edge imitation learning algorithms, enabling reliable, sample-efficient robots that deliver customer value from day one. We’re already live with pilot partners and scaling fast. The founding team brings experience from Apple, Stanford, and Microsoft, with deep expertise in robotics, embodied AI, and large-scale machine learning.

The Role

We’re looking for a Robotics Data Infrastructure Engineer to own and build the data systems that power our robots in the real world. This is a hands-on founding engineer role with true ownership and freedom — your work will directly impact robots performing customer-critical tasks every day.

You will architect and deploy data pipelines on both AWS and edge devices, manage large-scale multi-modal datasets (images, video, time-series, text, etc.), and build the tooling that connects real-world robot data to training and evaluation workflows. You’ll work across the full robotics software stack, from ingesting sensor data and telemetry, to enabling large-scale policy learning pipelines that drive production robots.

Beyond writing great code, you’ll help drive technical decisions, lead cross-functional efforts, and bridge robotics, machine learning, and product requirements into scalable, reliable systems.

What You’ll Do

  • Build and own our data backbone on AWS: Design and run cloud + edge pipelines using services like IoT Core, S3, ECR, Batch, ECS/EKS, and Step Functions. Your work keeps robot data flowing reliably and cost-efficiently from the field into the lab.
  • Develop on-device data systems: Build robust, fault-tolerant data capture on edge PCs using MCAP/Protobuf, with clean schema contracts, buffering, and resumable uploads to the cloud.
  • Wrangle massive multimodal datasets: Organize and version millions of images, videos, time-series (robot state, force/torque), and annotations. Enforce metadata, retention, and access patterns that scale.
  • Build MLOps and DataOps pipelines: Automate data validation, labeling, augmentation, and model training/evaluation using containerized jobs and orchestrators like Batch, Step Functions, Airflow, or Prefect.
  • Ensure data quality and health: Create ingestion checks, schema validation, deduping, drift detection, and real-time alerting around data freshness and completeness.
  • Build internal tools that unblock others: Develop UIs/CLIs for browsing data, launching jobs, tracking experiments, and debugging robots in the field. Integrate with tools like Foxglove.
  • Work across teams: Partner with hardware, ML, and product to turn raw field data into smarter robots and real customer value—fast.

Qualifications

  • B.S., M.S., Ph.D. in computer science or related fields.
  • Strong programming skills in Python (you write clean, efficient, production-ready code)
  • Strong experience in AWS
  • Systems engineering skills (networking, concurrency, performance)
  • At least 2 years of full-time work experience for candidates with a B.S. in related fields. 1 year of experience for M.S. or Ph.D. candidates.

Compensation

The base pay range for this role is $110,000 – $175,000 per year.

Ready to apply?

Powered by

Gem Logo

First name *

Last name *

Email *

LinkedIn URL

Resume *

Click to upload or drag and drop here

Apply

Req ID: R21