Site Reliability Engineering Manager
Conifers.ai · District de Tel Aviv
Job description
About the role
Conifers is building an AI‑native security platform and needs a Site Reliability Engineering Manager to own the health and reliability of its production system. This hands‑on leader will guide a small SRE team, collaborate with R&D, Product and GTM, and ensure the platform is observable, stable and ready for customers 24/7.
Key responsibilities
- Own end‑to‑end health and reliability of the production platform, beyond basic infrastructure uptime.
- Design, build and continuously improve monitoring, observability, alerting, dashboards and health‑check tooling.
- Define platform‑wide reliability, availability, performance and product‑health metrics and targets.
- Proactively detect production issues, abnormal behavior and reliability risks before they affect customers.
- Lead incident response, root‑cause analysis, post‑mortems and follow‑up actions.
- Partner with R&D to improve system resilience, error handling, scalability and production readiness.
- Work with Product to embed health indicators and operational readiness into new features.
- Collaborate with GTM and customer‑facing teams during production incidents and complex troubleshooting.
- Build automation and internal tools that reduce manual operational work.
- Mentor and lead a small team of engineers while staying deeply hands‑on.
Required profile
- 4+ years of experience in SRE, production engineering, DevOps or related fields.
- Proven technical leadership or engineering‑management experience.
- Strong software‑engineering background with ability to investigate complex, multi‑service problems.
- Demonstrated experience operating highly available, complex production systems.
- Excellent troubleshooting, root‑cause analysis and incident‑management skills.
- Ability to work across Engineering, Product, GTM and customer‑facing teams.
- Strong sense of ownership and a track record of driving issues from detection to resolution.
Required skills
- Observability, monitoring and alerting frameworks.
- Distributed systems design and performance tuning.
- Scalability and reliability engineering principles.
- Automation and scripting/programming proficiency.
- Modern cloud platforms (e.g., AWS, GCP, Azure).
- Container orchestration and Docker/Kubernetes environments.
- Incident management and post‑mortem processes.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Israel.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 2 days ago
Expires 1 month from now
11 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Conifers.ai
District de Tel Aviv
Related job offers
-
Senior Information Security Engineer
Fireblocks District de Tel Aviv -
AI Team Leader
Axon Pulse District de Tel Aviv -
Director of Product Management
Tipalti District de Tel Aviv -
Product Manager, Risk & Resilience (Cybersecurity)
mastercard Ramat-Gan -
IT End User Support Specialist
Altera District de Jérusalem