On-device ML Performance Engineer, Graphics, Games and Machine Learning
Apple
- Location
- Onsite (Cupertino, California)
- Employment
- Full-time
- Level
- Senior Level
Posted 1 day ago
About the Role
Join Apple's On-device ML Performance team to optimize machine learning models for Apple Silicon hardware. You will analyze and enhance performance, power, and memory usage across iPhones and Macs, ensuring full hardware utilization for cutting-edge AI experiences.
Skills
Machine Learning
Computer Architecture
Performance Optimization
Python
C++
PyTorch
GPU
CPU
SoC
Inference
Quantization
Shell Scripting
Linux
macOS
Metal Performance Shaders
CoreML
Full job details
The On-Device Machine Learning team at Apple is responsible for enabling the Research to Production lifecycle of cutting edge machine learning models that power magical user experiences on Apple’s hardware and software platforms. Apple is the best place to do on-device machine learning, and this team sits at the heart of that discipline, interfacing with research, SW engineering, HW engineering, and products.
The On-device ML Performance team has the responsibility to analyze latency, memory, power and numerical correctness of the latest machine learning models running on Apple SoC’s, and to make Apple’s ML software stack take full advantage of the capabilities in Apple’s ML accelerators. The work from this cross functional team enables model developers’ decisions to optimize performance via advanced techniques such as different model authoring techniques, quantization, sparsity, performance and accuracy tradeoffs. The work of this team impacts all new Apple HW and ML Inference on them.
Our group is looking for an On-device ML Performance Engineer, with technical expertise in computer architecture, performance, memory, power, ML model architectures, ML frameworks such as PyTorch, and on-device ML inference. The role entails deep analysis of ML models and their architecture, the implementation of the models in the ML SW stack for optimum performance, power and memory usage, and debug involving the performance and power consumption of CPU, GPU, and Apple Neural Engine.
As an engineer in this role, you will be primarily focused on analyzing and optimizing the performance of the latest ML models on the latest iPhones and Mac’s. You will work with models created by the most popular ML frameworks (PyTorch, MLX, etc) and will analyze the inference of those models on device to ensure the stack achieves full machine performance on Apple Silicon. The role also includes scripting, coding, model import and conversions, and generation of utilities and debug tools to extract, analyze, and report performance and power related metrics for Apple HW. The ideal candidate will have a passion for ML model architectures and ML inference, deep knowledge of GPU and CPU, computer architecture, compilers, and has a natural inclination toward innovation and exploration.
Experience with ML inference, quantization, performance and accuracy Familiarity and experience with the most popular ML architectures (e.g. LLM’s, Diffusion models, CNN’s) A passion to explore and learn about the latest advances in ML model design and architecture, particularly as related to model implementation on HW and on-device inference Familiarity with Operating Systems, embedded systems, and CPU/GPU/SoC/Memory HW architectures Highly proficient in Python/C++ and shell scripting Familiarity with Linux or macOS Exceptional clarity in verbal and written communication, including the ability to summarize, present and lead discussions in larger groups
Masters or PhDs in Computer Science or relevant disciplines. Experience with Apple’s CoreML, MPS Graph, Metal Performance Shader’s or MLX frameworks Experience with any ML authoring framework (PyTorch, TensorFlow, JAX, etc.) Experience with implementation of high performance compute kernels for CPU, GPU or AI Accelerators Experience with Apple’s App development framework such as Xcode, Swift, Objective-C Experience with any on-device ML stack, such as TFLite, ONNX, ExecuTorch, etc. Experience with any compiler stack (MLIR/LLVM/TVM etc.)
Description
As an engineer in this role, you will be primarily focused on analyzing and optimizing the performance of the latest ML models on the latest iPhones and Mac’s. You will work with models created by the most popular ML frameworks (PyTorch, MLX, etc) and will analyze the inference of those models on device to ensure the stack achieves full machine performance on Apple Silicon. The role also includes scripting, coding, model import and conversions, and generation of utilities and debug tools to extract, analyze, and report performance and power related metrics for Apple HW. The ideal candidate will have a passion for ML model architectures and ML inference, deep knowledge of GPU and CPU, computer architecture, compilers, and has a natural inclination toward innovation and exploration.
Minimum Qualifications
Experience with ML inference, quantization, performance and accuracy Familiarity and experience with the most popular ML architectures (e.g. LLM’s, Diffusion models, CNN’s) A passion to explore and learn about the latest advances in ML model design and architecture, particularly as related to model implementation on HW and on-device inference Familiarity with Operating Systems, embedded systems, and CPU/GPU/SoC/Memory HW architectures Highly proficient in Python/C++ and shell scripting Familiarity with Linux or macOS Exceptional clarity in verbal and written communication, including the ability to summarize, present and lead discussions in larger groups
Preferred Qualifications
Masters or PhDs in Computer Science or relevant disciplines. Experience with Apple’s CoreML, MPS Graph, Metal Performance Shader’s or MLX frameworks Experience with any ML authoring framework (PyTorch, TensorFlow, JAX, etc.) Experience with implementation of high performance compute kernels for CPU, GPU or AI Accelerators Experience with Apple’s App development framework such as Xcode, Swift, Objective-C Experience with any on-device ML stack, such as TFLite, ONNX, ExecuTorch, etc. Experience with any compiler stack (MLIR/LLVM/TVM etc.)