Skip to content
← Back to job listings

Contract Site Reliability Engineer (AI Accelerator Infrastructure)

d-Matrix · United States

External listingfull-time13 days ago

About The Role

Join d-Matrix as a Contract Site Reliability Engineer (AI Accelerator Infrastructure) for a 6-month contract with potential for full-time conversion. As a core member of the SRE team, you will be responsible for the reliability, automation, and observability of the infrastructure. You will work across colocation, on-premises lab environments, and cloud platforms, owning your systems end-to-end. Your role will involve hands-on infrastructure work, capacity planning, IaC and configuration management, automation, monitoring, incident response, and documentation. You will also support customer-facing environments and collaborate with the DevOps team.

  • Assurer la fiabilité et la disponibilité des infrastructures assignées, y compris les serveurs en colocation, les clusters de laboratoire sur site et les environnements cloud.
  • Effectuer des travaux d'infrastructure pratiques, y compris le provisionnement des serveurs, la configuration du système d'exploitation, la gestion du stockage et le dépannage matériel.
  • Posséder l'IaC et la gestion de la configuration pour vos domaines d'infrastructure, en veillant à ce que tous les provisionnements et changements soient effectués par le code.
  • IaC experience with Terraform and/or Ansible — writing and maintaining production configurations, not just running existing playbooks
  • Hands-on experience with colocation or on-premises server infrastructure — physical hardware, rack networking, and bare-metal provisioning
  • Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking
  • Python and/or Bash scripting: production-quality automation, not just one-off scripts
  • Prometheus + Grafana or DataDog: building dashboards, writing alert rules, and understanding signal quality
  • Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance diagnostics
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration
  • Comfort operating in fast-moving startups: you own your systems, document what you build, and iterate without waiting for perfect requirements
  • Incident response experience: structured triage, RCA production, and follow-through on action items
  • Experience operating customer-facing infrastructure or platform services with external reliability expectations
  • HPC job scheduler experience: Slurm, LSF, or equivalent — operations and troubleshooting
  • Cloud infrastructure operations across AWS, Azure, or GCP—including hybrid environments spanning cloud and on-prem
  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink — configuration and troubleshooting
  • Experience with large-scale infrastructure automation: host lifecycle management, fleet auto-healing, or AIOps-driven operations
  • Go programming for SRE tooling — health-check services, exporters, or auto-remediation agents

This is an external listing. JobSpring does not represent or verify the employer. Report this listing