Skip to content
← Back to job listings

DevOps/SRE

Solidus Labs · New York, United States

External listingfull-timeabout 1 month ago

About The Role

Join our DevOps team as a Site Reliability Engineer in New York. You will be responsible for the reliability, stability, and operational support of our production systems. This role involves production ownership, monitoring, incident response, and on-call support. You will work with a modern cloud-native stack and play a key role in keeping systems highly available, secure, and performant.

  • Assurer la fiabilité, la disponibilité et la performance des environnements de production, y compris la gestion des clusters Kubernetes et des environnements AWS.
  • Évoluer l'infrastructure en tant que code en utilisant Terraform et Helm, et soutenir les pipelines CI/CD GitLab.
  • Diriger la réponse aux incidents de bout en bout, y compris le dépannage, l'atténuation et la résolution.
  • Proficiency with Terraform, Helm, and GitLab CI (or similar)
  • Scripting experience with Bash and Python
  • Familiarity with pub/sub systems (SQS, Kafka, or similar)
  • Solid knowledge of AWS (EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, DynamoDB)
  • 3+ years of hands-on DevOps / SRE experience
  • Strong troubleshooting skills across infrastructure, CI/CD, and networking
  • Strong production experience with Docker and Kubernetes
  • Experience with monitoring, logging, and alerting systems
  • Willingness to participate in on-call rotations
  • GitOps workflows and advanced Git usage
  • Experience supporting databases such as Postgres, Snowflake, or ClickHouse
  • Experience with Redis, Airflow, Databricks, Spark/EMR

This is an external listing. JobSpring does not represent or verify the employer. Report this listing