HPC System Engineer

Apply now »

Date: 7 Sept 2026

Location: Abu Dhabi, AE

Company: EDGE Group PJSC

ABOUT EDGE

At EDGE bold ideas are engineered into technologies that protect, improve and save lives.

 

Headquartered in the United Arab Emirates, EDGE is a leading advanced technology group working at the forefront of defence and emerging technologies. Spanning more than 35 specialised companies and multiple centres of excellence, we are purpose-built to move fast. Free from heavy legacy processes, we give our people the autonomy, accountability, and agility to bring breakthrough technologies from concept to reality.

 

Based in Abu Dhabi, a globally connected hub at the crossroads of Europe, Asia, Africa, and the Middle East, EDGE is home to a truly multicultural community where bold ideas thrive, and the future is shaped.

 

Together, we are shaping the future.

ABOUT THE ROLE

We are seeking a highly skilled HPC Systems Administrator to manage, optimize, and support our high-performance computing environment used for Computational Fluid Dynamics (CFD) and Finite Element Analysis (FEA) workloads. This role is responsible for the full lifecycle of HPC operations — from infrastructure and scheduler management to user support, performance optimization, and long-term capacity planning.
The ideal candidate has strong Linux administration experience, deep knowledge of HPC schedulers, and hands-on familiarity with engineering simulation tools.

RESPONSIBILITIES

A. Infrastructure Management
• Maintain and administer compute nodes, login nodes, heterogeneous nodes, NAS storage servers, and high-speed interconnects (InfiniBand).
• Manage and maintain HPC-related databases, ensuring timely backups and audit compliance.
• Monitor hardware health including CPU temperatures, memory errors, and disk failures.
• Ensure high availability, reliability, and minimal downtime across the HPC environment.
B. Scheduler & Resource Management
• Configure and tune job schedulers (Slurm / PBS / LSF / Grid Engine).
• Implement and maintain fair-share scheduling, job priority rules, QoS limits, and preemption policies.
• Manage node reservations for large-scale CFD/FEA workloads.
• Provide best-practice recommendations for CPU/GPU architecture selection for engineering simulations.
• Prevent resource misuse, including node hogging, queue congestion, and starvation of small jobs.
C. User Access, Security & Compliance
• Manage user accounts, permissions, and storage quotas.
• Enforce secure SSH access, MFA, and other security controls.
• Ensure compliance with IT governance, data-security, and audit requirements.
• Apply OS patches, security updates, and vulnerability fixes.
D. Software Stack Management
• Install, update, and manage licenses for engineering solvers such as ANSYS, Siemens, NASTRAN, and others.
• Maintain version control and apply service pack updates as required.
• Manage environment modules (Lmod, Environment Modules).
• Optimize compilers, MPI libraries, math libraries, and GPU/graphics-intensive drivers for performance.
RESTRICTED
E. Performance Monitoring & Optimization
• Monitor node utilization, queue wait times, job failures, and overall cluster performance.
• Identify and resolve network congestion, I/O bottlenecks, and memory pressure issues.
• Conduct performance tuning and benchmarking for CFD/FEA workloads.
• Recommend improvements to enhance throughput and efficiency.
F. Troubleshooting & User Support
• Diagnose and resolve job crashes, memory leaks, solver errors, and environment issues.
• Assist users with HPC batch scripting and job optimization.
• Provide training sessions and best-practice guidance for HPC users.
• Maintain documentation, onboarding guides, and knowledge-base resources.
G. Storage & Data Lifecycle Management
• Enforce storage quotas, purge policies, and data-retention rules.
• Manage backup, archival systems, and RAID configurations.
• Ensure smooth operation of parallel file systems (Lustre, GPFS, BeeGFS).
• Support efficient data workflows for large CFD/FEA datasets.
H. Capacity Planning & Future Growth
• Plan for future expansion including increased core counts, GPU adoption, memory-heavy nodes, and faster interconnects.
• Evaluate new hardware technologies and benchmark workloads before procurement.
• Provide input into long-term HPC strategy and infrastructure roadmap.

BASIC QUALIFICATIONS

  • Bachelor’s or Master’s degree in Computer Science, Engineering, Physics, or related field.
    • 3–7 years of experience administering Linux-based HPC clusters.
    • Strong knowledge of HPC schedulers (Slurm/PBS/LSF).
    • Experience with MPI, OpenMP, and distributed computing.
    • Familiarity with CFD/FEA solvers and engineering workflows.
    • Proficiency in scripting (Bash, Python).
    • Experience with high-speed networking and parallel file systems.

PREFERRED QUALIFICATIONS

  • Experience with GPU-accelerated workloads and CUDA-enabled solvers.
    • Knowledge of container technologies (Singularity/Apptainer).
    • Experience with monitoring tools (Grafana, Prometheus, XDMoD).
    • Background supporting engineering or scientific computing environments.
    • Understanding of performance profiling and benchmarking.

WHY CHOOSE EDGE

Launched in November 2019, the UAE's EDGE is one of the world's leading advanced technology groups, established to develop agile, bold and disruptive solutions for defence and beyond, and to be a catalyst for change and transformation. It is dedicated to bringing breakthrough innovations, products, and services to market with greater speed and efficiency, to position the UAE as a leading global hub for future industries, and to creating clear paths within the sector for the next generation of highly skilled talent to thrive.

At EDGE, the real investment goes beyond compensation. Through our dedicated learning academy and digital learning platforms, employees have access to extensive opportunities to develop new skills, deepen their expertise, and advance their careers. Combined with the freedom to innovate, strong career support, and the opportunity to collaborate with world-class talent from across the globe, EDGE is a place where you can continue to grow, keep learning, and turn ambitious ideas into meaningful outcomes.

Working at EDGE also comes with a package that genuinely reflects how much we value our people. Salaries are highly competitive and tax-free and, depending on your role and seniority, benefits can include family visas, annual flight tickets, medical insurance for you and your dependents, and education allowances for your children.

CANDIDATE PRIVACY & EQUAL OPPORTUNITY STATEMENT

If you feel your skills, experience and potential are a strong match for a role, we encourage you to apply even if you don’t meet every requirement. We are genuinely interested in what you can bring to EDGE.

Any information you share as part of your EDGE candidate profile or job application will be handled in accordance with applicable data protection and privacy regulations.

EDGE is an equal opportunity employer committed to building a diverse, inclusive, and high-performing workforce. We believe every individual deserves to be treated with fairness, respect and dignity, and our employment decisions reflect that. We do not discriminate based on race, colour, nationality, ethnicity, religion, gender, marital status, age, disability, pregnancy or parental status, or any other characteristic protected by applicable law.

As a global organisation, effective collaboration across international teams is central to how we work, so English proficiency is required for all roles unless otherwise stated in the job posting.

Please note that EDGE does not accept unsolicited resumes from recruitment agencies. Any resumes in the absence of a formally executed agreement in place will be deemed the property of EDGE, and no placement fees or compensation will be payable in respect of such submissions.


Job Segment: Computer Science, Linux, Technology

Apply now »