Skip to content
← Back to job listings

Senior Systems Engineer (Performance & Reliability)

Graphcore · Bristol, United Kingdom

External listingfull-time21 days ago

About The Role

Join Graphcore, a leading AI technology company, as a Senior Systems Engineer specializing in Performance & Reliability. In this role, you will measure and evaluate large-scale Linux systems, design workloads, build execution systems, and interpret results. You will have the opportunity to specialize in various areas while contributing to the overall understanding of system behavior at scale. Enjoy a flexible work-life balance, private medical insurance, pension, income protection, and opportunities for progression and development.

  • Designing workloads and building execution systems to measure and evaluate large-scale Linux systems.
  • Expanding measurement coverage from small clusters to full racks, and interpreting results to define what “good enough” looks like.
  • Collaborating with team members to understand system behavior at scale, and contributing to the overall decision-making process.
  • Our engineers typically bring significant practical experience and sound engineering judgement
  • Depth in one area is valued, but the ability to work across boundaries is equally important
  • Experience working in Linux-based environments, ideally with distributed or high-performance systems
  • Comfortable working in areas where requirements are not fully defined and judgement is required
  • Proficiency in Python
  • Ability to interpret results and communicate findings clearly, with an emphasis on accuracy and usefulness to decision-making
  • Experience with automation and CI/CD systems (e.g. GitLab CI, Jenkins, GitHub Actions)
  • Ability to design, implement, and run experiments or tests that produce meaningful results
  • Strong software engineering experience, typically gained across multiple projects or systems over several years
  • Experience working with large-scale or distributed systems (e.g. clusters, cloud platforms, HPC environments)
  • Experience with performance, reliability, or systems-level testing/measurement
  • Familiarity with pytest or similar frameworks for structured test/measurement execution
  • Experience analysing system behaviour under load (compute, network, or ML workloads)
  • Experience working with containerisation, orchestration, or provisioning systems (e.g. Docker, Kubernetes, OpenStack)
  • Proficiency in other applications programming languages (e.g. C++)
  • Exposure to data analysis, statistics, or interpreting variability in results

This is an external listing. JobSpring does not represent or verify the employer. Report this listing