Source description
About the role
AI Workload & Stress Tool Development Design and develop scalable stress, performance, and validation frameworks for MAIA AI accelerator platforms. Build workload generation infrastructure capable of exercising compute, memory, interconnect, networking, storage, and system-level resources. Develop reusable stress tools using PyTorch, Triton, Python, C++, and custom MAIA SDKs. Create synthetic and production-inspired workloads that model training and inference behaviors observed in large-scale AI deployments. Build automated infrastructure for workload deployment, orchestration, telemetry collection, and result analysis. Develop and optimize kernels targeting custom AI accelerators. Analyze execution behavior across the hardware-software stack and identify bottlenecks impacting utilization and performance. Collaborate with compiler and runtime teams to improve workload efficiency and hardware utilization. Develop tooling that integrates with MAIA compiler pipelines, SDKs, runtime environments, and performance analysis tools. Build automation around model compilation, kernel validation, regression testing, and workload portability. Design workload suites for platform bring-up, qualification, and reliability testing. Build comprehensive regression infrastructure supporting silicon, firmware, system software, and platform releases. Develop automated validation tools capable of identifying correctness, performance, thermal, power, and stability issues. Enable platform readiness through scalable validation methodologies and continuous regression testing. Develop benchmarking methodologies and performance dashboards. Developer Productivity & Automation Improve developer productivity through automation, CI/CD integration, diagnostics, and debugging infrastructure. Build reusable tooling for workload generation, failure triage, telemetry analysis, and reporting. Develop dashboards and automated workflows for large-scale validation environments. Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 7+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience 8+ years of experience developing and optimizing AI training and inference workloads for GPUs, AI accelerators, or HPC platforms, including distributed AI systems, compute-intensive kernel development, and performance-focused software development using frameworks such as C++, PyTorch and Triton. 8+ years of experience analyzing and optimizing workloads on AI accelerator, GPU, or HPC platforms, including performance profiling, bottleneck analysis, and workload optimization, with knowledge of accelerator architectures, memory hierarchies, interconnects, runtime systems, and distributed AI infrastructure. 8+ years of experience developing and optimizing GPU or AI accelerator kernels; building automated stress, validation, benchmarking, and reliability frameworks; and driving performance analysis and root-cause resolution across hardware, software, and distributed system environments. These requirements include but are not limited to the following specialized security screenings: Experience with AI compiler technologies and kernel generation frameworks, including LLVM, MLIR, Triton Compiler, or similar compiler toolchains. Experience training, optimizing, or deploying large-scale AI models, including LLM training and inference workloads. Experience with custom AI accelerator SDKs, collective communication libraries, and large-scale distributed computing environments. Experience supporting silicon bring-up, platform qualification, post-silicon validation, or hardware/software integration activities and cloud-scale validation infrastructure
More at Microsoft
Related open roles
Fabric Interconnect Design Verification Engineer
San Francisco Bay Area · Seattle · Austin · Onsite
Senior Firmware Engineer
San Francisco Bay Area · Seattle · Portland · Onsite
Principal Signal Integrity Simulation Engineer
Seattle · Onsite
Firmware Engineer II
San Francisco Bay Area · Seattle · Austin · Onsite
Senior Silicon Test Engineer
San Francisco Bay Area · Seattle · Austin · Onsite
Critical Environment Electrical Engineer
Phoenix · Onsite