← Back to job listings
BV
DevOps Engineer (Cloud Infrastructure & Live Games)
Big Viking Games · Toronto, Canada
About The Role
Join Big Viking Games as a Senior DevOps Engineer, where you'll design, maintain, secure, and modernize the infrastructure supporting our live-service games and internal development workflows. This hands-on role requires expertise in cloud infrastructure, automation, CI/CD, containers, monitoring, uptime, and production reliability. You'll work closely with various teams to improve our systems and processes, while also managing incident response and cloud spend. Enjoy benefits like extended health and dental insurance, hybrid work options, generous paid time off, and continuous learning opportunities.
- Design, maintain, secure, and modernize the infrastructure that supports live-service games and internal development workflows.
- Drive infrastructure modernization while maintaining uptime for live games, building, maintaining, and improving automation for deployments.
- Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support live games.
- Experience with containerized applications, especially Docker
- Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting
- Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices
- Comfort creating documentation, runbooks, and repeatable operating processes
- Practical ownership mindset with the ability to prioritize, execute, and close loops
- Strong communication skills with both technical and non-technical stakeholders
- Ability to work closely with software engineers to improve build, deploy, and operational workflows
- Experience with CI/CD tools, version control, deployment automation, and modern release workflows
- Experience supporting production systems where uptime, reliability, and performance matter — especially systems that cannot tolerate extended downtime
- Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar
- Experience designing, maintaining, and improving production infrastructure — including comfort with legacy systems that predate modern cloud-native patterns
- Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes
- Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages
- 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role
- Strong hands-on experience with AWS or similar cloud platforms
- Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments
- Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance
- Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products
- Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools
- This role is best suited for someone who wants meaningful ownership over production infrastructure, cloud health, automation, and DevOps modernization inside a live-service gaming company
- Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on
- The ideal candidate is a practical infrastructure engineer who can keep live systems stable while helping modernize how the company builds, deploys, secures, and operates technology
- Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability
- Experience with container orchestration platforms such as Kubernetes, ECS, EKS, or Nomad
- Experience with Redis, Memcached, queues, workers, or event-driven systems
- Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning
- Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms
- Experience using AI tools such as Claude, ChatGPT, Gemini, or similar platforms to improve DevOps workflows, documentation, troubleshooting, and automation
- Experience operating infrastructure that supports AI/ML workflows, API integrations, or automation platforms (<Make.com>, webhook-driven orchestration, MCP servers)
- They are not only focused on tools. They understand uptime, developer experience, production risk, cloud costs, security, release quality, and operational discipline. They recognize that modernizing a decade-old live game requires patience, pragmatism, and the ability to improve systems incrementally without disrupting what's working
- Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools
- They are comfortable operating across a mix of legacy and modern infrastructure, managing credentials and secrets lifecycle across multiple hosting platforms, and ensuring data pipelines are healthy and alerting properly. They can work independently, collaborate with engineers, and create systems that reduce friction instead of adding process for its own sake
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring