Site Reliability Engineering Manager
Conifers.ai · District de Tel Aviv
وصف الوظيفة
About the role
Conifers is building an AI‑native security platform and needs a Site Reliability Engineering Manager to own the health and reliability of its production system. This hands‑on leader will guide a small SRE team, collaborate with R&D, Product and GTM, and ensure the platform is observable, stable and ready for customers 24/7.
Key responsibilities
- Own end‑to‑end health and reliability of the production platform, beyond basic infrastructure uptime.
- Design, build and continuously improve monitoring, observability, alerting, dashboards and health‑check tooling.
- Define platform‑wide reliability, availability, performance and product‑health metrics and targets.
- Proactively detect production issues, abnormal behavior and reliability risks before they affect customers.
- Lead incident response, root‑cause analysis, post‑mortems and follow‑up actions.
- Partner with R&D to improve system resilience, error handling, scalability and production readiness.
- Work with Product to embed health indicators and operational readiness into new features.
- Collaborate with GTM and customer‑facing teams during production incidents and complex troubleshooting.
- Build automation and internal tools that reduce manual operational work.
- Mentor and lead a small team of engineers while staying deeply hands‑on.
Required profile
- 4+ years of experience in SRE, production engineering, DevOps or related fields.
- Proven technical leadership or engineering‑management experience.
- Strong software‑engineering background with ability to investigate complex, multi‑service problems.
- Demonstrated experience operating highly available, complex production systems.
- Excellent troubleshooting, root‑cause analysis and incident‑management skills.
- Ability to work across Engineering, Product, GTM and customer‑facing teams.
- Strong sense of ownership and a track record of driving issues from detection to resolution.
Required skills
- Observability, monitoring and alerting frameworks.
- Distributed systems design and performance tuning.
- Scalability and reliability engineering principles.
- Automation and scripting/programming proficiency.
- Modern cloud platforms (e.g., AWS, GCP, Azure).
- Container orchestration and Docker/Kubernetes environments.
- Incident management and post‑mortem processes.
Questions fréquentes
لماذا تبلغ عن هذا العرض؟
اكتشف المزيد
الرواتب والأدلة وعمليات البحث في Israel.
الرواتب حسب المهنة
قدم طلبك في 30 ثانية
أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.
بالمتابعة، أنت توافق على شروط الاستخدام.
لديك حساب بالفعل؟ تسجيل الدخول
عزز فرصك
حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.
جاري تحليل سيرتك الذاتية...
Conifers.ai
District de Tel Aviv
عروض عمل ذات صلة
-
Technical Support Engineer
Fireberry District de Tel Aviv -
Senior Data Engineer – GenAI Foundation Models
Booking.com District de Tel Aviv -
Senior Machine Learning Engineer – GenAI Applications
Booking.com District de Tel Aviv -
Senior Software Engineer (DevOps)
Akamai Technologies Tel Aviv-Yafo -
IT Coordinator – Academic IT Support
Weizmann Institute of Science Rehovot