Overview:

About Us

Core42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.


The opportunity

We are seeking a highly skilled Senior Engineer – HPC Operations to oversee the daily operations and support of high-performance computing clusters designed to power large-scale AI and ML workloads. This role ensures stable, secure, and high-performing infrastructure leveraging technologies such as Slurm, Kubernetes, and modern MLOps platforms. The ideal candidate will bring deep technical expertise in HPC and a strong operational mindset to drive continuous improvement and automation across globally distributed environments. Responsibilities will extend to collaborating with multidisciplinary teams, leading complex projects, implementing cutting-edge technologies, and providing mentorship to operations engineers.

Responsibilities:

 

Your key responsibilities

 

  • Lead the daily operational support of HPC infrastructure including compute, storage, networking, and scheduler components (Slurm, Kubernetes, etc.).
  • Lead efforts to maximize the efficiency and performance of HPC systems, ensuring optimal resource utilization and minimal downtime.
  • Act as the primary technical escalation point for L2 support teams and ensure prompt resolution of incidents and service requests.
  • Monitor system health, performance, and utilization using advanced tools (e.g., Prometheus, Grafana, DCGM).
  • Manage user environments for AI/ML workloads including container orchestration (e.g., Docker, Kubernetes) and workflow tools (e.g., MLflow, Kubeflow).
  • Implement and manage job scheduling policies, priorities, and partitions within Slurm and/or Kubernetes environments to ensure fairness and efficiency.
  • Lead root cause analysis (RCA) of operational issues and contribute to post-mortem documentation and continuous improvement efforts.
  • Provide mentorship and guidance to junior engineers and participate in on-call rotation if required.
  • Ensure compliance with security and operational policies; assist in audits and documentation for change and incident management processes.

Qualifications:

 

What we’re looking for

(a) Required skills / qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related technical field.
  • 5+ years of experience in HPC operations, systems engineering, or DevOps roles.
  • Advanced knowledge and expertise in configuring, optimizing, and maintaining complex HPC environments, including hardware, software, and storage systems.
  • Hands-on experience managing Slurm clusters and/or Kubernetes-based environments for AI/ML workloads.
  • Expert knowledge of GPU resource management, workload schedulers, and performance tuning for AI/ML workloads.
  • Experience with monitoring and observability frameworks such as Prometheus, Grafana, and DCGM.
  • Strong scripting and automation skills (Python, Bash, Ansible, Terraform).
  • In-depth understanding of Linux (RHEL/CentOS/Ubuntu), networking concepts (RDMA, InfiniBand, RoCE), and storage technologies (NFS, Lustre, Ceph).

 

Compensation

The U.S. base salary range for this full-time role is US$106,400 to US$159,600 per year, with bonus and benefits on top. Salary ranges are determined by role, level, and location. The range listed represents the minimum and maximum target salary for new hires across all U.S. locations. Actual compensation within this range will depend on factors such as work location, job-related skills, experience, and relevant education or training

 

What Working at Core42 Offers

With a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative, and collaborative environment. At Core42, we are grounded in trust, accountability, and high performance. We are united by our values: Grit, Passion, and Impact—driving resilience, excellence, and meaningful progress across everything we do.

 

Core42 is committed to building a diverse and inclusive workplace. As an equal opportunity employer, Core42 does not discriminate based on race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or any other legally protected status. In compliance with the Americans with Disabilities Act (ADA), we provide reasonable accommodations to qualified individuals with disabilities throughout the application and employment process. If you need assistance or a reasonable accommodation, please contact reasonableaccommodations@core42.com, including the role you are applying for and the accommodation required.