Overview In this role you will define and evolve a next gen HPC architecture to enable advanced research, analytics and AI at MSK. You will design scalable, reliable systems and partner with infrastructure teams to support mission critical scientific discovery. You'll optimize compute, storage and networking while exploring cloud enabled or hybrid HPC options. This is a high impact, technical leadership role at a world class cancer research center.
Compensation / Benefits- hybrid work schedule
- onsite a few times per month
- competitive compensation
- location: NJ data center
- benefits and diversity commitments
- exempt status
Responsibilities- Define and evolve the technical vision and architecture for a next generation HPC environment
- Design system level solutions for performance, scalability, reliability, and maintainability
- Support research computing workloads, analytics, simulation, and AI/ML pipelines
- Architect Linux based HPC platforms, clusters, and services with enterprise resilience
- Evaluate and integrate cloud resources to extend or hybridize HPC capabilities
- Optimize compute, storage, networking, and scheduler performance by workload
- Collaborate with networking, facilities, data center, storage, and cloud teams
- Establish standards, documentation, and operational practices for HPC hybrid platforms
- Provide technical leadership and guidance on architecture decisions, tradeoffs, and roadmap planning
Key requirements- Bachelor's degree in Computer Science, Computer Engineering, or related field
- Strong background in HPC environments including cluster design and operations
- Advanced Linux administration experience (Red Hat based)
- Experience with job scheduling/workload management (e.g., Slurm)
- Exposure to cloud platforms, distributed systems, and modern compute architectures; cloud certifications a plus
- Collaborative partnership across infrastructure, engineering, and operations
- Strong analytical and optimization mindset
- Clear technical communication and documentation
- High-performance computing architectures
- Linux system administration (Red Hat based)
- Job scheduling/workload managers (Slurm)