Site Reliability Engineering Manager
Conifers.ai · District de Tel Aviv
תיאור המשרה
About the role
Conifers is building an AI‑native security platform and needs a Site Reliability Engineering Manager to own the health and reliability of its production system. This hands‑on leader will guide a small SRE team, collaborate with R&D, Product and GTM, and ensure the platform is observable, stable and ready for customers 24/7.
Key responsibilities
- Own end‑to‑end health and reliability of the production platform, beyond basic infrastructure uptime.
- Design, build and continuously improve monitoring, observability, alerting, dashboards and health‑check tooling.
- Define platform‑wide reliability, availability, performance and product‑health metrics and targets.
- Proactively detect production issues, abnormal behavior and reliability risks before they affect customers.
- Lead incident response, root‑cause analysis, post‑mortems and follow‑up actions.
- Partner with R&D to improve system resilience, error handling, scalability and production readiness.
- Work with Product to embed health indicators and operational readiness into new features.
- Collaborate with GTM and customer‑facing teams during production incidents and complex troubleshooting.
- Build automation and internal tools that reduce manual operational work.
- Mentor and lead a small team of engineers while staying deeply hands‑on.
Required profile
- 4+ years of experience in SRE, production engineering, DevOps or related fields.
- Proven technical leadership or engineering‑management experience.
- Strong software‑engineering background with ability to investigate complex, multi‑service problems.
- Demonstrated experience operating highly available, complex production systems.
- Excellent troubleshooting, root‑cause analysis and incident‑management skills.
- Ability to work across Engineering, Product, GTM and customer‑facing teams.
- Strong sense of ownership and a track record of driving issues from detection to resolution.
Required skills
- Observability, monitoring and alerting frameworks.
- Distributed systems design and performance tuning.
- Scalability and reliability engineering principles.
- Automation and scripting/programming proficiency.
- Modern cloud platforms (e.g., AWS, GCP, Azure).
- Container orchestration and Docker/Kubernetes environments.
- Incident management and post‑mortem processes.
Questions fréquentes
מדוע אתם מדווחים על ההצעה הזו?
Explore further
Salaries, guides and searches for ישראל.
Salaries by job title
הגש בקשה ב-30 שניות
הזינו את המייל שלכם כדי להגיש בקשה. חשבון יווצר אוטומטית.
בהמשך, אתם מסכימים לתנאי השימוש שלנו.
כבר יש לכם חשבון? התחברות
מתפרסם לפני יום
תפוגה בעוד חודש מעכשיו
7 צפיות · 0 interested
הגדל את סיכוייך
העלה את קורות החיים שלך: אנו מציעים לך מודעות תואמות לפרופיל שלך.
מנתח את קורות החיים שלך...
Conifers.ai
District de Tel Aviv