Ensure continuous, reliable operation of company services.
- •Ensure the continuous, reliable operation of company services through proactive monitoring, incident response, and building resilient observability and escalation practices.
- •Key Responsibilities Ensure monitoring and uninterrupted operation of company services.
- •Write and maintain alerting rules and runbooks.
- •Perform triage of incoming incidents and initial diagnosis of issues.
- •Build and maintain escalation chains for incident response.
- •Collaborate with development and infrastructure teams to identify reliability risks and implement preventive measures.
- •Requirements Experience with observability tools (Grafana, ELK, VictoriaMetrics).
- •Experience working with Linux.
- •Experience working with Kubernetes (k8s).
- •Experience with AWS and Azure cloud platforms.
View original posting →