← Back to job listings
TA
Staff Platform Engineer (Service Infrastructure)
Together AI · San Francisco, United States
About The Role
Join Together AI as a Staff Platform Engineer, where you will drive the service infrastructure strategy for the Product Foundations engineering organization. This hands-on role involves evolving the core infrastructure strategy, improving existing services, and partnering with various teams to ensure reliability and consistency. You will also build reusable service infrastructure primitives and establish durable technical standards. Benefits include competitive health insurance, retirement plans, flexible time off, and more.
- Own the technical direction for service infrastructure within Product Foundations, including Kubernetes, AWS, Terraform, CDNs, ALBs, DNS, IAM, service networking, and related operational patterns.
- Up-level existing Product Foundations services by improving reliability, operability, deployment safety, infrastructure consistency, and production readiness.
- Build and evolve reusable service infrastructure primitives, including Helm charts, Terraform modules, GitHub Actions/GitOps workflows, service scaffolding, runbooks, and documentation.
- AWS experience, ideally including EKS, IAM, VPC networking, load balancing, Route 53, CloudFront, ECR, and related service infrastructure
- Deep production experience with Kubernetes, including EKS, Helm, ArgoCD/Argo Rollouts, ingress, autoscaling, secrets, service identity, networking, and progressive delivery
- Direct experience with observability systems, including metrics, logs, traces, dashboards, alerting, SLOs, and incident response
- Strong Terraform experience, including module design, infrastructure CI/CD, policy enforcement, production applies, and safe self-service workflows
- Experience operating networking and edge infrastructure such as CDNs, ALBs/NLBs, DNS, TLS, ingress/egress controls, and traffic management
- Proficiency in one or more programming languages used for infrastructure tooling and automation, such as Go, Python, TypeScript, or similar
- Proven ability to lead cross-functional technical initiatives across product engineering, infrastructure, networking, and security teams
- 7+ years of professional experience in platform engineering, service infrastructure, SRE, distributed systems, cloud infrastructure, or related roles
- Staff-level judgment: you can define ambiguous problems, make pragmatic tradeoffs, influence without authority, and leave both systems and teams better than you found them
- Strong written communication skills, with experience producing clear design docs, migration plans, operational guidance, and technical standards
- Experience with supply-chain security, image signing, SBOMs, vulnerability management, or compliance automation
- Experience with multi-region, multi-cluster, hybrid-cloud, or cross-provider service networking
- Experience with service mesh or zero-trust infrastructure such as mTLS, SPIFFE/SPIRE, Cilium, Istio, Linkerd, Envoy, or similar
- Experience with OPA, Gatekeeper, Kyverno, Sentinel, or other policy-as-code systems
- Experience embedding infrastructure best practices into product engineering teams at scale
- Experience building internal developer platforms or paved-path service frameworks used by many engineering teams
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring