Ensure continuous, reliable operation of company services.
- •We are looking for a Site Reliability Engineer (SRE) to join our team and ensure the continuous, reliable operation of company services.
- •Key Responsibilities Ensure monitoring and uninterrupted operation of company services Write and maintain alerting rules and runbooks Perform triage of incoming incidents and initial diagnosis of issues Build and maintain escalation chains for incident response Perform technical incident resolution activities according to runbooks Requirements Experience with observability tools (Grafana, ELK, VictoriaMetrics) Experience working with Linux Experience working with Kubernetes (k8s) Experience with AWS and Azure cloud platforms Ability to analyze incidents, identify root causes, and propose remediation steps
View original posting →