Source description
About the role
Define and execute a comprehensive data strategy that spans AI model training, simulation, and production deployment across safety-critical autonomous systems.
Own the end-to-end data pipeline — from raw collection and labeling through curation, versioning, and delivery — ensuring the reliability and scale that training and simulation workflows demand.
Build and maintain data flywheels that continuously improve model performance by closing the loop between deployed system behavior and future training iterations.
Collaborate closely with teams provisioning and operating large-scale GPU/TPU training clusters to align data delivery with compute capacity and training schedules.
Drive the design and integration of data pipelines that feed data-driven, physics-based, and high-fidelity simulators, ensuring simulated environments are realistic enough to support confident AI model validation.
Partner with safety, validation, and certification teams to establish data quality standards and traceability practices that satisfy regulatory requirements in aviation and/or automotive domains.
Lead, mentor, and grow a team of data and infrastructure engineers, setting technical direction and fostering a culture of rigor, ownership, and continuous improvement.
Define and track KPIs for data pipeline health, simulation fidelity, and model readiness, using these metrics to prioritize investments and communicate progress to senior leadership.
More at Merlin Labs
Related open roles
Test Data Engineer
United States · Onsite
Flight Test Instrumentation Engineer
United States · Onsite
Information Systems Security Engineer
Boston · Hybrid
Lead Systems Engineer, Flight Autonomy
Boston · Hybrid
Senior Systems Engineer, Flight Autonomy
United States · Hybrid
Director, Compute Platform
Boston · Hybrid
