Skip to content
← Back to job listings

Senior Systems Engineer (Performance & Reliability)

Graphcore · London, United Kingdom

External listingfull-time21 days ago

About The Role

Join Graphcore, a leading AI technology company, as a Senior Systems Engineer specializing in Performance & Reliability. In this role, you will measure and evaluate large-scale Linux systems, design workloads, build execution systems, and interpret results. You will have the opportunity to specialize in various areas while contributing to the overall understanding of system behavior at scale. Enjoy a flexible work-life balance, private medical insurance, pension plans, income protection, and opportunities for career progression and personal development.

  • Designing and implementing workloads that effectively expose system behavior and performance characteristics.
  • Building and maintaining systems to run experiments across large clusters, ensuring accurate and reliable measurement.
  • Interpreting results and defining what "good enough" looks like, contributing to decision-making regarding system reliability and performance.
  • We’re looking for engineers who are comfortable working where the right answer isn’t obvious, and where careful measurement matters more than output volume
  • Our engineers typically bring significant practical experience and sound engineering judgement. Depth in one area is valued, but the ability to work across boundaries is equally important
  • Experience with automation and CI/CD systems (e.g. GitLab CI, Jenkins, GitHub Actions)
  • Ability to interpret results and communicate findings clearly, with an emphasis on accuracy and usefulness to decision-making
  • Comfortable working in areas where requirements are not fully defined and judgement is required
  • Experience working in Linux-based environments, ideally with distributed or high-performance systems
  • Strong software engineering experience, typically gained across multiple projects or systems over several years
  • Ability to design, implement, and run experiments or tests that produce meaningful results
  • Proficiency in Python
  • Experience working with large-scale or distributed systems (e.g. clusters, cloud platforms, HPC environments)
  • Experience with performance, reliability, or systems-level testing/measurement
  • Familiarity with pytest or similar frameworks for structured test/measurement execution
  • Experience analysing system behaviour under load (compute, network, or ML workloads)
  • Experience working with containerisation, orchestration, or provisioning systems (e.g. Docker, Kubernetes, OpenStack)
  • Proficiency in other applications programming languages (e.g. C++)
  • Exposure to data analysis, statistics, or interpreting variability in results

This is an external listing. JobSpring does not represent or verify the employer. Report this listing