Skip to content
← Back to job listings

Infrastructure & Reliability Software Engineer

CrewAI · United States

External listingfull-time23 days ago

About The Role

Join CrewAI as an Infrastructure & Reliability Software Engineer, where you'll build and operate the platform infrastructure behind our cloud and enterprise deployments. You'll work across multiple hyperscalers, including AWS, Azure, and GCP, and focus on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. This role is not just a DevOps support position; you'll write code, improve systems, design deployment paths, harden production, and build the internal platform that allows CrewAI to scale.

  • Construire et exploiter l'infrastructure de la plateforme derrière les déploiements cloud et d'entreprise de CrewAI.
  • Écrire du code, améliorer les systèmes, concevoir des chemins de déploiement, renforcer la production et construire la plateforme interne qui permet à CrewAI de se développer.
  • Posséder et améliorer l'infrastructure qui exécute la plateforme CrewAI : AWS, ECS/ECR, Docker, Kubernetes/Helm, réseaux, secrets, bases de données, Redis et services connexes.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers
  • Strong infrastructure/platform engineering experience in production SaaS environments
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar
  • Experience supporting enterprise/self-hosted deployments
  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems
  • Terraform or other IaC experience
  • SRE background: SLOs, incident review, capacity planning, load testing
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability

This is an external listing. JobSpring does not represent or verify the employer. Report this listing