Contract Site Reliability Engineer (AI Accelerator Infrastructure)
d-Matrix · United States
About The Role
Join d-Matrix as a Contract Site Reliability Engineer (AI Accelerator Infrastructure) for a 6-month contract with potential for full-time conversion. As a core member of the SRE team, you will be responsible for the reliability, automation, and observability of the infrastructure. You will work across colocation, on-premises lab environments, and cloud platforms, owning your systems end-to-end. Your role will involve hands-on infrastructure work, capacity planning, IaC and configuration management, automation, monitoring, incident response, and documentation. You will also support customer-facing environments and collaborate with the DevOps team.
- Assurer la fiabilité et la disponibilité des infrastructures assignées, y compris les serveurs en colocation, les clusters de laboratoire sur site et les environnements cloud.
- Effectuer des travaux d'infrastructure pratiques, y compris le provisionnement des serveurs, la configuration du système d'exploitation, la gestion du stockage et le dépannage matériel.
- Posséder l'IaC et la gestion de la configuration pour vos domaines d'infrastructure, en veillant à ce que tous les provisionnements et changements soient effectués par le code.
- IaC experience with Terraform and/or Ansible — writing and maintaining production configurations, not just running existing playbooks
- Hands-on experience with colocation or on-premises server infrastructure — physical hardware, rack networking, and bare-metal provisioning
- Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking
- Python and/or Bash scripting: production-quality automation, not just one-off scripts
- Prometheus + Grafana or DataDog: building dashboards, writing alert rules, and understanding signal quality
- Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance diagnostics
- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration
- Comfort operating in fast-moving startups: you own your systems, document what you build, and iterate without waiting for perfect requirements
- Incident response experience: structured triage, RCA production, and follow-through on action items
- Experience operating customer-facing infrastructure or platform services with external reliability expectations
- HPC job scheduler experience: Slurm, LSF, or equivalent — operations and troubleshooting
- Cloud infrastructure operations across AWS, Azure, or GCP—including hybrid environments spanning cloud and on-prem
- Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink — configuration and troubleshooting
- Experience with large-scale infrastructure automation: host lifecycle management, fleet auto-healing, or AIOps-driven operations
- Go programming for SRE tooling — health-check services, exporters, or auto-remediation agents
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring