Principal Software Engineer
To support the scaling of AI infrastructure, the full-time Principal Software Engineer will focus on Kubernetes cluster operations, GPU resource scheduling, and implementing monitoring capabilities for production systems, working remotely or onsite in Durham, NC. Key responsibilities Develop custom software for scheduling GPU resources on Kubernetes to support large scalable GPU clusters for AI workloads Implement monitoring and health management capabilities to ensure reliability and scalability of GPU assets Collaborate with cross-functional teams to evaluate system failures and enhance service performance through a defined incident management process Required qualifications 15+ years of experience in software engineering, particularly with large-scale production systems Proficient in Kubernetes APIs and frameworks, with a focus on software development rather than just cluster operations BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience Technical knowledge in systems programming languages such as Go or Python, along with a solid understanding of data structures and algorithms Experience in managing and automating large-scale distributed systems and cluster management systems