Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Principal AI Accelerator Tools Development Engineer

San Francisco Bay Area · Seattle · Portland · OnsitePosted 4 days ago
HardwareStaff+Full TimeH-1B track record
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

AI Workload & Stress Tool Development Design and develop scalable stress, performance, and validation frameworks for MAIA AI accelerator platforms. Build workload generation infrastructure capable of exercising compute, memory, interconnect, networking, storage, and system-level resources. Develop reusable stress tools using PyTorch, Triton, Python, C++, and custom MAIA SDKs. Create synthetic and production-inspired workloads that model training and inference behaviors observed in large-scale AI deployments. Build automated infrastructure for workload deployment, orchestration, telemetry collection, and result analysis. Develop and optimize kernels targeting custom AI accelerators. Analyze execution behavior across the hardware-software stack and identify bottlenecks impacting utilization and performance. Collaborate with compiler and runtime teams to improve workload efficiency and hardware utilization. Develop tooling that integrates with MAIA compiler pipelines, SDKs, runtime environments, and performance analysis tools. Build automation around model compilation, kernel validation, regression testing, and workload portability. Design workload suites for platform bring-up, qualification, and reliability testing. Build comprehensive regression infrastructure supporting silicon, firmware, system software, and platform releases. Develop automated validation tools capable of identifying correctness, performance, thermal, power, and stability issues. Enable platform readiness through scalable validation methodologies and continuous regression testing. Develop benchmarking methodologies and performance dashboards. Developer Productivity & Automation Improve developer productivity through automation, CI/CD integration, diagnostics, and debugging infrastructure. Build reusable tooling for workload generation, failure triage, telemetry analysis, and reporting. Develop dashboards and automated workflows for large-scale validation environments. Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 7+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience 8+ years of experience developing and optimizing AI training and inference workloads for GPUs, AI accelerators, or HPC platforms, including distributed AI systems, compute-intensive kernel development, and performance-focused software development using frameworks such as C++, PyTorch and Triton. 8+ years of experience analyzing and optimizing workloads on AI accelerator, GPU, or HPC platforms, including performance profiling, bottleneck analysis, and workload optimization, with knowledge of accelerator architectures, memory hierarchies, interconnects, runtime systems, and distributed AI infrastructure. 8+ years of experience developing and optimizing GPU or AI accelerator kernels; building automated stress, validation, benchmarking, and reliability frameworks; and driving performance analysis and root-cause resolution across hardware, software, and distributed system environments. These requirements include but are not limited to the following specialized security screenings: Experience with AI compiler technologies and kernel generation frameworks, including LLVM, MLIR, Triton Compiler, or similar compiler toolchains. Experience training, optimizing, or deploying large-scale AI models, including LLM training and inference workloads. Experience with custom AI accelerator SDKs, collective communication libraries, and large-scale distributed computing environments. Experience supporting silicon bring-up, platform qualification, post-silicon validation, or hardware/software integration activities and cloud-scale validation infrastructure

More at Microsoft

Related open roles

View all roles