We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team
In this role, you will contribute to the design, development, and operation of sophisticated cloud-based AI solutions built on the CNCF ecosystem including Kubernetes, running on cutting-edge hardware from leading vendors
This role focuses mainly on deploying AI infrastructure built on NVIDIA-certified hardware, following architecture and implementation designs produced by our engineering team
You will play a pivotal role in ensuring the reliability, security, and performance of container infrastructure, while mentoring team members and Mirantis customers to deliver high-quality software and services
As a senior engineer, you will work closely with stakeholders to define technical strategies, solve complex challenges, and ensure the seamless integration of cloud and software services
This is an excellent opportunity to make a significant impact while driving innovation in a rapidly evolving cloud ecosystem
Work with geographically distributed international teams on technical challenges and process improvements
Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software
Collaborate with stakeholders to gather and refine technical requirements
Optimize system performance, reliability, and scalability
Troubleshoot, debug, and resolve complex technical issues
Participate in code reviews to maintain high quality standards
Stay up to date with industry trends and best practices in cloud operations and development
Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance
Facilitate knowledge transfer to customers during the delivery phases- A commitment to innovation, continuous learning, and delivering high-quality results
Ability to travel up to 25% if needed, including internationally
Excellent customer-facing communication skills
Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams
Excellent written and spoken English
Experience with high-performance data center processing, networking, and storage
At least 5 years of DevOps or Software Development experience or in a similar role
Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines
Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security
Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight
5+ years of professional experience in DevOps, with a strong focus on cloud and infrastructure technologies, including Kubernetes and/or OpenStack
Exposure to Golang and working knowledge of other programming languages (Python, JavaScript)
Bachelor’s degree in Computer Science or a related field, or equivalent experience
Extensive experience in network and/or storage architecture
Experience with high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle or NVIDIA AI Enterprise
Presence in the open source community including upstream contribution and conference presentations
Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware