Skip to content
← Back to job listings

Principal Software Engineer (Machine Learning Infrastructure)

Snap Inc. · Palo Alto, United States

External listingfull-time20 days ago

About The Role

Join Snap Inc. as a Principal Software Engineer specializing in Machine Learning Infrastructure. In this role, you will oversee the technical strategy and architecture of ML Inference Platform services, design and implement critical engineering components, and advocate for best practices in availability, scalability, and operational excellence. You will have the opportunity to influence the entire company with your technical direction. Enjoy a comprehensive benefits package, including generous parental leave, medical coverage, and retirement savings options.

  • Oversee technical strategy and architecture in ML Inference Platform services, ensuring alignment with company goals.
  • Design, implement, and scale critical engineering components and services to support ML inference and deployment.
  • Work across teams to understand product requirements, evaluate trade-offs, and deliver the solutions needed to build innovative products or services.
  • Skilled at solving ambiguous problems
  • Strong collaboration and mentorship skills
  • Ability to proactively learn new concepts and technology and apply them at work
  • Proven track record of operating highly-available systems at scale
  • Excellent programming and software design skills, including debugging, performance analysis, and test design
  • 2+ years of experience with technical leadership or acting as the domain-expert to a technical organization
  • Experience in technical leadership/ownership and setting technical direction for engineering projects
  • Bachelors in technical field such as computer science, mathematics, statistics or equivalent years of experience
  • 10+ years of post-Bachelor’s software development experience; or a Master’s degree in a technical field + 9+ year of post-grad software development experience; or a PhD in a related technical field + 6+ years of post-grad software development experience
  • Experience architecting, designing, and developing large scale distributed systems and high-throughput RPC services
  • Advanced degree in a technical field such as computer science
  • Experience with cross platform development
  • Ability to promote product excellence and collaboration, driving a portfolio of concurrent engineering projects, from short-term critical feature launches to long-term research initiatives
  • Ability to create a compelling vision for the future, communicate clearly, and have a collaborative leadership approach
  • Experience with Tensorflow, PyTorch, Kubernetes, GPU, LLM inference is a plus

This is an external listing. JobSpring does not represent or verify the employer. Report this listing