Research Engineer - ML Infrastructure
Описание от работодателя
You will develop core frameworks for model training and evaluation that make models performant, resource-efficient, and reliable at scale. You will build and optimize the training stack, profile GPU-cluster workloads, remove bottlenecks, and ensure fault-tolerant, deterministic execution for long-running jobs.