Padmi
Institute of Foundation Models logo
Institute of Foundation Models

foundation models · large language models

Inference Optimization Intern – Performance Modeling

United States · OnsitePosted 26 days ago
Machine learningInternIntern Fall
Apply at Institute of Foundation Models

Opens the source posting on jobs.lever.co

Source description

About the role

View original

This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs. Responsibilities include:

Develop analytical performance models for GPU kernels and inference workloads.

Build and validate a simulator to estimate theoretical hardware performance limits.

Compare measured kernel performance against architectural peak throughput.

Identify performance bottlenecks in compute, memory, communication, and scheduling.

Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.

Investigate PTX and SASS code generation to understand low-level execution behavior.

Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.

Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.

Design profiling methodologies for Hopper and Blackwell architectures.

Document findings and provide actionable recommendations for performance improvements.

More at Institute of Foundation Models

Related open roles

View all roles