Senior HPC / Kubernetes Infrastructure Engineer
TEN openings here!
Please focus on candidates with deep hands-on Kubernetes infrastructure experience at significant scale. Strong profiles will also demonstrate Python automation, Terraform, GPU/HPC infrastructure, NVIDIA GPU exposure, and production troubleshooting experience.
Role Overview
We are seeking a highly experienced High Performance Computing (HPC) / Kubernetes Infrastructure Engineer to build, operate, troubleshoot, and scale the infrastructure supporting large-scale AI and LLM workloads.
This is a deeply hands-on engineering role focused on the infrastructure underneath AI models, including Kubernetes, GPU compute, automation, cloud/on-prem infrastructure, job scheduling, monitoring, and large-scale distributed environments.
Candidates should have experience working in complex, production-scale infrastructure environments, ideally connecting hundreds or thousands of servers into cohesive compute platforms.
Key Responsibilities
- Build, operate, troubleshoot, and scale large-scale Kubernetes environments supporting AI and compute-intensive workloads.
- Design and support infrastructure utilizing NVIDIA GPUs and GPU-based compute environments.
- Develop Python tooling and automation to integrate and manage infrastructure components.
- Build and maintain infrastructure using Terraform.
- Support GPU/HPC workload scheduling and large-scale compute environments.
- Implement monitoring, troubleshooting, and incident response across complex infrastructure.
- Work with AWS infrastructure, including EC2, S3, EFS, and FSx.
- Support infrastructure across cloud and on-prem environments.
- Work with multi-cloud Kubernetes environments, including platforms such as GCP and CoreWeave.
- Take ownership of complex infrastructure problems and drive solutions with minimal oversight.
Required Qualifications
- 4+ years of hands-on Kubernetes experience at scale.
- Strong experience building and operating large, complex infrastructure environments.
- Strong Python programming experience, particularly for infrastructure tooling and automation.
- Hands-on experience with Terraform / Infrastructure as Code.
- Experience with GPU infrastructure, HPC environments, or GPU/HPC job scheduling.
- Strong production monitoring, troubleshooting, and incident response experience.
- Experience with AWS, particularly EC2, S3, EFS, and/or FSx.
- Demonstrated ability to independently solve complex infrastructure and systems problems.
- Must be a hands-on engineer who has actually built, operated, scaled, and troubleshot production infrastructure.
Highly Preferred
- Direct High Performance Computing (HPC) experience.
- Experience supporting AI, machine learning, or LLM infrastructure.
- Hands-on experience with NVIDIA GPU infrastructure.
- Experience operating environments consisting of hundreds to thousands of servers.
- Multi-cloud Kubernetes experience, particularly GCP and/or CoreWeave.
- Experience with both on-prem and cloud infrastructure.
Ideal Candidate
The ideal candidate comes from a strong HPC, Kubernetes Platform Engineering, AI Infrastructure, GPU Infrastructure, or large-scale distributed systems background. We are looking for engineers who thrive in technically demanding environments, take ownership, and can execute independently without significant hand-holding.
This is not primarily an AI/ML model development role. The focus is on building the infrastructure and systems that enable AI and LLM workloads to operate at scale.
Technical Priority Areas
- Kubernetes at scale, 4+ years: Highest priority
- Python tooling & automation
- Terraform
- GPU / HPC job scheduling
- Monitoring & incident response
- AWS: EC2, S3, EFS, FSx
- Multi-cloud Kubernetes: GCP / CoreWeave