Staff Software Engineer, HPC
@ ZooxStaff Software Engineer, HPC
About the job
Zoox develops autonomous vehicles and infrastructure, focusing on scalable HPC systems to support AI workloads and software teams. Innovate in robotics, machine learning, and mobility tech in a fast-paced, collaborative environment.
Requirements
- Experience designing large-scale distributed systems
- Experience with Ray.io or similar
- Experience with Kubernetes workloads
- Experience with cloud infrastructure on AWS
- Proficiency with Python
Qualifications
- Background in scalable system operation
- Strong collaboration skills
- Prior experience in reliable infrastructure
- Ability to prioritize technical work
Full job description
In this role, you will:
Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs
Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap for the HPC platform
Lead multi-quarter, cross team initiatives that drive org-wide improvements
Create production-grade APIs, SDKs, and tools that make it easy for engineers across Zoox to run large-scale distributed workloads
Design and improve job scheduling algorithms and auto-scaling policies to maximize reliability and resource availability
Design multi-region orchestration strategies that optimize for data locality, reliability, and performance
Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners across multiple teams
Evaluate new technologies and paradigms that improve Zoox's computational and storage capabilities
Develop capacity planning tools and forecasting models to support Zoox's growing compute needs
Mentor junior engineers, guiding them through their career development
Qualifications
Experience designing and operating large-scale distributed systems in production
Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
Experience with Kubernetes, particularly for heterogeneous workloads
Experience with cloud infrastructure on AWS or similar providers
Track record of shipping and operating reliable, highly available scalable infrastructure
Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
Proficiency with Python
Bonus Qualifications
Exposure to machine learning workloads (training, inference, data generation)
Experience with Kubernetes or SLURM at scale (>10k+ nodes)
Experience with SLURM workload manager and advanced scheduling policies
Background in algorithmic optimization or operations research
Experience building developer tools and platforms used by large engineering organizations
Similar jobs in Columbus, OH
- A
Staff Metrology Engineer
Anduril Industries · Ashville, Ohio, United States
Posted 1 week ago - A
Staff Manufacturing Engineer, Roadrunner
ANDURIL INDUSTRIES · Ashville, OH, United States
Posted 1 day ago - A
Transportation Engineer
AMT Engineering · Columbus, OH
Posted 2 weeks ago - O
Staff RN- Grant, ICU
OhioHealth · Columbus, OH, United States
Posted 3 weeks ago - O
Staff RN- Surgical Urology
OhioHealth · Columbus, OH, United States
Posted 1 week ago - O
Staff RN, Endoscopy
OhioHealth · Columbus, OH, United States
Posted 3 days ago