Site Reliability Engineer (FedRAMP)
Veeam · United States
About The Role
Join Veeam, a leading provider of data management solutions, as a Government Site Reliability Engineer. In this role, you will support the Veeam Data Cloud, our SaaS platform, specifically focusing on our Government and Sovereign Cloud environment. You will work alongside senior engineers to execute reliability work, respond to incidents, and maintain the operational foundation of the team. Key responsibilities include discovery and documentation, reliability and incident response, observability, infrastructure and delivery, and collaboration. This position offers a range of benefits, including unlimited paid time off, medical coverage, a 401(k) retirement plan, and professional training and education opportunities.
- Participate in incident response, including triage, investigation, mitigation, and postmortems, and help implement and maintain SLIs, SLOs, and error budgets.
- Work with Infrastructure as Code (IaC), CI/CD pipelines, and deployment tooling in compliance-restricted environments, and support testing, canary deployments, and release validation workflows.
- Collaborate with engineering, security, compliance, and operations teams to execute on reliability improvements, and communicate clearly about system behavior, risk, and status.
- Strong programming skills in one or more of: TypeScript/JS, Go, Java, C#, or similar
- Experience with IaC tools (Terraform, Terragrunt, or Pulumi) and container orchestration (Kubernetes)
- Experience with cloud infrastructure on Azure or a comparable cloud provider
- 3+ years in Software Engineering, with at least 1 year in SRE, Platform Engineering, or DevOps working on cloud-hosted services
- Able to read and understand code well enough to investigate system behavior without always having someone walk you through it
- Clear written and verbal communication skills
- Solid understanding of distributed systems fundamentals and networking basics
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK stack)
- Experience with CI/CD tooling such as GitHub Actions, Azure DevOps, GitLab CI, or ArgoCD
- Familiarity with regulated or compliance-oriented environments such as government (FedRAMP, CMMC), financial (PCI-DSS), or healthcare (HIPAA). You understand that compliance shapes what you can and can't do operationally
- Experience in Government or Sovereign Cloud environments (e.g., Azure Government, AWS GovCloud)
- Background in SaaS platforms or multi-tenant systems
- Familiarity with chaos engineering, resilience testing, or load testing
- Exposure to building or improving reliability practices on a team
- Familiar with AI-first development workflows using LLM-powered tools for automation, code generation, or documentation
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring