Skip to content
← Back to job listings

Staff Infrastructure Engineer

Replit · United States

External listingfull-time2 months ago

About The Role

Join Replit as a Staff Infrastructure Engineer and play a crucial role in ensuring the reliability, scalability, and performance of our infrastructure that serves millions of developers worldwide. You will bridge the gap between development and operations, implement automation, and establish best practices for efficient scaling and high availability. Your mission will involve proactively finding and analyzing reliability problems, designing and implementing software and systems for improvement, and mentoring the broader engineering team.

  • Assurer la fiabilité, l'évolutivité et la performance de l'infrastructure de Replit, en mettant en œuvre l'automatisation et en établissant des pratiques exemplaires.
  • Identifier et analyser proactivement les problèmes de fiabilité à travers notre pile technologique, puis concevoir et mettre en œuvre des logiciels et des systèmes pour créer des améliorations significatives.
  • Collaborer avec les équipes d'infrastructure et de produit pour optimiser nos déploiements cloud, identifier et résoudre les goulets d'étranglement de performance.
  • 8-10 years of experience in Infrastructure Engineering or similar roles (DevOps, Systems Engineering, Site Reliability Engineering)
  • We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale
  • Strong incident management skills with experience leading incident response and demonstrated critical thinking under pressure
  • Excellent written and verbal communication skills, with an ability to explain technical concepts clearly and simply and a bias toward open, transparent cultural practices
  • You write high-quality, well-tested code
  • Strong interpersonal skills, with experience working with engineers from junior to principal levels
  • Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture
  • Strong programming skills in languages like Python or Go
  • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies
  • Experience with infrastructure as code (e.g., Terraform) and configuration management tools
  • Proven track record of implementing and maintaining monitoring/observability solutions, with strong skills in debugging and performance tuning
  • A willingness to dive into understanding, debugging, and improving any layer of the stack
  • You're passionate about making software creation accessible and empowering the next generation of builders
  • Deep experience with Google Cloud Platform (GCP) services and tools
  • Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.)
  • Experience designing and building reliable systems capable of handling high throughput and low latency
  • Experience with Go and Terraform
  • Familiarity with working in rapid-growth environments
  • Experience writing company-facing blog posts and training materials

This is an external listing. JobSpring does not represent or verify the employer. Report this listing