Director of Machine Learning Infrastructure
Apply or review the details on the original posting.
Apply for this roleOriginal posting on LinkedIn
Metric are exclusively partnered with one of the world's leading L4 autonomous driving organizations who have taken a foundation model approach to autonomous driving & are also leveraging world models for hyper-realistic simulation, enabling them to generalize and scale their autonomous driving service across different vehicle platforms and new driving environments.
They are seeking to hire a Director of Machine Learning Infrastructure who will be responsible for defining the strategy for, and implementing, an enterprise-wide centralized ML Infrastructure platform spanning model training, simulation & real-world model deployment with the aim being to increase efficiency and enable cost optimization.
This is a rare opportunity to join a mature, stable organization but with the mandate to architect a greenfield ,centralized ML platform, establish and scale your own team & influence company-wide AI strategy.
Responsibilities
Define and execute company-wide ML infrastructure spanning training, world-model based simulation, real-world deployment.
Partner with Developer Experience/Platform Team to develop centralized scheduling, intelligent queuing, and fair-share resource allocation that maximizes GPU utilization, reduces idle capacity, and ensures predictable access for ML teams.
Build ML platform observability capabilities; telemetry standards and self-service dashboards to provide Research Teams with visibility into model performance, infrastructure health, GPU utilization, and cost-per-workload.
Establish and scale a world-class centralized ML Infra Engineering Team
Requirements
10+ years of experience within ML Infra Engineering, with specific experience in a technical leadership capacity
Demonstrated success delivering organization-wide infrastructure transformation to standardize and modernize ML platforms
Technical expertise across modern ML platform technologies, with hands-on experience designing scalable training infrastructure, workload orchestration, experiment tracking, and platform observability to be used by a world-class Research Team.