Padmi
Institute of Foundation Models logo
Institute of Foundation Models

foundation models · large language models

Distributed Machine Learning Engineer

San Francisco Bay Area · Onsite$150k–$450k/yrPosted 16 months ago
Machine learningUnspecifiedFull Time
Apply at Institute of Foundation Models

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization

Design and implement performance benchmarks and testing methodologies to evaluate application performance

Build tools to automate workload analysis, workload optimization, and other critical workflows

Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization

Support the team to develop appropriate kernels and systems for new model architectures and algorithms

Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.

Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).

Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.

Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC and deep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.

Perform all other duties as reasonably directed by the line manager that are commensurate with these functional objectives.

More at Institute of Foundation Models

Related open roles

View all roles