We are looking for enthusiastic individuals to join our team as Site Reliability Engineer II, where you’ll play a crucial role in ensuring the reliability and performance of our platform
This is a top-notch opportunity to learn from skilled engineers and contribute to solving complex technical challenges
You will be involved in monitoring systems, responding to incidents, and developing automation to streamline operations
The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems
The team is responsible for ensuring high reliability, optimal system performance, and continuous improvement for both Instacart’s critical internal services and externally visible systems
SRE’s work emphasizes optimizing existing systems, developing robust infrastructure, and automating processes to reduce manual tasks
Joining the SRE team means tackling unique scaling challenges while applying expertise in coding, algorithms, complexity analysis, and large-scale system design
The team thrives on a culture of intellectual curiosity, problem-solving, and collaboration. With members from diverse backgrounds and experiences, SRE fosters a supportive and risk-tolerant environment where individuals can think big, take on meaningful projects, and grow with guidance and mentorship
Participate in monitoring systems and responding to alerts, escalating issues as needed
Assist in incident management, following established protocols and documenting steps
Help maintain and improve documentation for processes and procedures
Contribute to the development and maintenance of automation scripts and tools
Learn and apply best practices in site reliability engineering
Support the deployment of applications and services, ensuring smooth and reliable releases
Collaborate with senior engineers to troubleshoot and tackle technical issues- Are you passionate about technology and ready to dive into the world of Site Reliability Engineering?
We are looking for someone eager to grow their skills, learn new technologies, and contribute to a culture of reliability
Ability to work well in a team environment and communicate effectively
Strong interest in system administration and troubleshooting
Understanding of programming concepts and scripting languages
Excellent problem-solving and analytical skills
Strong sense of ownership on incident handling
2-4 years of Software engineering experience
Eagerness to learn and grow in the field of Site Reliability Engineering
Familiarity with cloud platforms (e.g., AWS, GCP, Azure) is a plus