Skip to content
← Back to job listings

Tech Ops Engineer (AI Native)

ConsumerAffairs · United States

External listingfull-time6 days ago

About The Role

Join our team as a Tech Ops Engineer (AI Native) and play a crucial role in driving operational excellence through AI-powered automation. You will be responsible for monitoring and maintaining our infrastructure, providing technical support, executing system upgrades, and optimizing performance. You will also support disaster recovery plans, manage infrastructure lifecycle, and collaborate with various teams to ensure effective communication and alignment on projects. This position offers comprehensive benefits, including health coverage, retirement plans, open PTO, and opportunities for professional growth.

  • Monitor and maintain the organization’s infrastructure, including servers, networks, storage systems, and applications, ensuring optimal performance and uptime.
  • Identify opportunities to automate routine tasks and processes, improving operational efficiency and reducing manual workload, and implement scripts, automation tools, and AI skills to streamline system management and monitoring.
  • Plan and execute system upgrades, patches, and configuration changes, ensuring minimal disruption to business operations, and test and validate updates in development environments before deploying them to production.
  • 5+ years of professional experience in Linux administration, managing AWS resources, developing CI/CD and server orchestration pipelines, scripting and monitoring
  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent work experience
  • 5+ years of experience in system administration, or a similar role
  • Configuring and managing data sources like PostgreSQL/Aurora RDS, OpenSearch, Redis, and message streaming platforms like Kafka
  • Design and maintain log ingestion pipelines (e.g., Vector → OpenSearch, Vector → Kafka), including index retention, document shape optimization, and failure recovery
  • Experience maintaining logging, monitoring, and alerting capabilities using OpenSearch, Vector log pipelines, Prometheus, and Kafka
  • Cloud-based production systems at scale
  • Infrastructure as Code tools, primarily Terraform
  • Creating CI/CD pipelines with Jenkins, Concourse or other CI/CD implementation
  • Monitoring tools, like Datadog or Prometheus
  • You are an expert in:
  • Amazon Web Services (EC2, VPC, EFS, S3, EKS etc.)
  • You have experience with:
  • Scripting for server side automation, auditing, and monitoring
  • Working in a Python and JavaScript-centric codebase and are familiar with their related best-practices
  • Production experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments
  • Exceptional analytical and problem-solving skills
  • A working knowledge of modern software practices and technologies such as Agile methodologies
  • Triage and remediate security vulnerabilities (CVEs) across infrastructure components, including container base images, OS packages, and third-party services
  • Experience around Security and Compliance
  • Promoting and establishing development standard methodologies for AWS infrastructure-as-code
  • The ideal candidate would also have:
  • Experience with High Availability implementations
  • Experience with AWS Well-Architected principles
  • Displays a growth mindset to continually improve; encourages everyone around them to be tenacious and never settle
  • Strong communicator with the ability to collaborate across teams and provide clear, concise technical support
  • Intellectual curiosity, a willingness to learn new skills and the ability to contribute new ideas
  • Learns quickly and using whatever resources to solve new problems
  • Adaptable and flexible, able to manage multiple tasks and prioritize effectively in a fast-paced environment
  • Demonstrates a relentless focus on results with a commitment to deliver
  • Detail-oriented with a focus on maintaining high standards of operational reliability
  • Relentless in their pursuit of success and possessing the willpower to embrace challenges as opportunities
  • Takes decisive action, and confidently changes course if unsuccessful
  • Obsessed with ensuring an exceptional customer experience- for both internal and external customers
  • Constantly seeks feedback to improve; Focuses on solving issues through teamwork, and collaboration
  • Acts with urgency; delivers top results in hours and days instead of weeks and months
  • Stands up for decisions, takes responsibility for results, and shares both good and bad outcomes transparently
  • Proactive and self-motivated, with a passion for continuous learning and improvement in technology operations

This is an external listing. JobSpring does not represent or verify the employer. Report this listing