user avatar

HPC Infrastructure & Cluster Engineer

Marathon TS Inc

Posted today

Job Requirements

Top Secret/SCI CI Polygraph
Career Level not specified
$150,000 - $190,000

Job Description

HPC Infrastructure & Cluster Engineer
Springfield, VA
$150k-$190k

Marathon TS is seeking a TS/SCI-cleared Infrastructure Engineer to manage and optimize a high-performance compute environment supporting intensive AI/ML workloads. This role is heavily focused on Linux cluster administration, GPU infrastructure, workload scheduling, high-speed networking, storage, and containerized environments.
Key Responsibilities:
  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.


Basic Qualifications:
  • Clearance: Active TS/SCI with the ability to obtain CI Poly.
  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
    • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
    • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
    • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
    • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.


Preferred Qualifications:
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.


US Citizenship Required
TS/SCI Required

Marathon TS is committed to the development of a creative, diverse and inclusive work environment. In order to provide equal employment and advancement opportunities to all individuals, employment decisions at Marathon TS will be based on merit, qualifications, and abilities. Marathon TS does not discriminate against any person because of race, color, creed, religion, sex, national origin, disability, age or any other characteristic protected by law (referred to as "protected status ").

#CJJOBS
group id: 10362312

Similar Jobs


Clearance Level
Top Secret/SCI