Senior HPC DevOps Engineer
Коротко
Engineers HPC DevOps solutions on AWS with GPU, MPI, NCCL, GPUDirect, Apptainer, SLURM, Lustre, and NVIDIA Nsight.
Описание от работодателя
Required skills: HPC, AWS, GPU, MPI, NCCL, GPUDirect, Apptainer, SLURM, Lustre, NVIDIA Nsight
We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform.
Responsibilities
Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction
Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads
Maintain automated pipelines for infrastructure provisioning and platform service deployments
Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas
Collaborate with developer experience teams to improve documentation
Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus
Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines
Ensure compute availability through capacity planning and reservation management
Deploy containerized environments tuned for HPC and GPU pass-through
Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions
Requirements
5+ years of experience in HPC or DevOps engineering roles
Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect
Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia
Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management
Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot
Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre
Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)
Proficiency in English at a B2+ level
We offer
We gather like-minded people:
Top tech minds driving innovation in AI, cloud and digital platform modernization
Supportive team and agile, startup-like culture
Hybrid by design mode and opportunity to work remotely within Poland
Chance to work abroad for up to 60 days annually
Business-driven relocation opportunities
We provide growth opportunities:
Career development programs
Thought leadership, mentoring, soft skills and well-being programs
Certification (Anthropic, Gemini, GCP, Azure, AWS)
English classes
We cover it all:
Stable pay
Participation in the Employee Stock Purchase Plan with a 15% discount
Benefits package (health insurance, multisport, shopping vouchers)
Referral bonuses up to $2,000
Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
Corporate, social and well-being events
Please, note:
Benefits listed above are available to employees only
We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
We will reach out to selected candidates exclusively
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.