Skip to content
← Back to job listings

Platform Monitoring & Incident Engineer

Adyen · San Francisco, United States

External listingfull-timeabout 2 months ago

About The Role

Join our team as a Monitoring Engineer, where you'll play a crucial role in ensuring the reliability and performance of our platform. You'll be responsible for monitoring platform performance, coordinating incidents, communicating with customers, and providing feedback to product engineering teams. You'll also lead initiatives to proactively detect issues and improve platform reliability. This is a dynamic role that requires strong communication skills, problem-solving abilities, and experience with monitoring and logging tools.

  • Monitoring platform performance, coordinating and commanding incidents, and communicating with customers.
  • Initiating and leading initiatives across platform offerings to proactively detect issues and increase reliability.
  • Coordinating the mitigation, recovery, and resolution of high-impact incidents, ensuring a rapid and effective response.
  • You have experience with problem management practices - identifying trends across incidents, conducting root cause investigations and driving preventative action
  • Work schedule: The shifts are from 9.00AM - 6.00PM PDT or 11.00AM - 8.00PM CDT with a 6-day workweek at least twice a month (Sunday–Friday or Monday–Saturday)
  • You have solid communication skills and the ability to develop strong working relationships throughout the organization, able to translate technical situations clearly and concisely to a diverse audience via data-visualizing dashboards and written documents
  • You have a natural ability for handling complex situations and multiple responsibilities simultaneously
  • You have experience with observability platforms like Datadog, Dynatrace, Splunk
  • You have a passion for defining and standardizing processes to drive strategic improvement and able to translate complex technical concepts with ease for all non technical audiences
  • You thrive in an environment where collaboration is crucial and where a global approach is key for are you successful implementation of processes and projects
  • You have excellent analytical and problem-solving skills, with the ability to analyze complex systems and spot the root cause of issues
  • You have at least 5 years of experience with incident management, problem management, incident client communication, and platform monitoring operations
  • You have experience with monitoring and logging tools like Prometheus, Grafana, ELK Stack, etc
  • You're a strong team player and thrive in a dynamic environment
  • You're willing to participate in the on-call rotation and work in a fast-paced, dynamic environment

This is an external listing. JobSpring does not represent or verify the employer. Report this listing