Engineering Manager, Cloud Infrastructure

@ Replit

Engineering Manager, Cloud Infrastructure

Foster City, California
Posted 1 day ago

About the job

Replit democratizes software creation with an AI-powered platform for building apps. The Cloud Infrastructure Engineering Manager leads foundational tech, including Kubernetes, networking, storage, and team growth, ensuring reliability and scalability.

Requirements

  • Engineering management experience
  • Built cloud platforms or distributed systems
  • Owned production migrations and incidents
  • Reasoned across infrastructure code
  • Led teams with technical expertise

Qualifications

  • Software-oriented infrastructure skills
  • Deep knowledge of Kubernetes and networking
  • Judgment for safe-change management
  • Understanding of internal customer needs
  • Team leadership and coaching ability

Full job description

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the Role

Replit enables people to build software with AI. The infrastructure underneath that experience must make it straightforward to launch services, isolate workloads, and run reliable systems at scale.

We're hiring a hands-on Engineering Manager to lead Cloud Infrastructure: the shared infrastructure as code (IaC), networking, storage, compute, and service mesh platforms that Replit's product and platform teams depend on. You'll lead and grow an existing engineering team building and operating these foundations, including Kubernetes, shared edge networking, service mesh and workload identity.

This is a platform-building role with production accountability. You should be comfortable going deep on a design or incident while developing a team that does not depend on you for every decision.

What You'll Do

  • Own the cloud-platform roadmap. Lead the team's IaC, networking, storage, compute, and service mesh platforms. Translate product, platform, reliability, and security needs into sequenced outcomes, balancing foundational investment, lifecycle work, and delivery commitments against the team's capacity.

  • Make infrastructure repeatable and self-service. Build maintained IaC interfaces for services, cells, connectivity, identities, and shared resources. Enable internal customer teams to provision infrastructure without bespoke coordination or dependence on individual experts.

  • Operate what the team builds. Own platform availability, upgrades, isolation, recovery, and incident remediation. Maintain clear SLOs, sustainable on-call coverage, and primary and backup owners for critical systems.

  • Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated infrastructure changes.

  • Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, set clear expectations, manage performance, and hire against agreed needs. Delegate meaningful ownership as the team grows.

What You'll Bring

  • Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor.

  • Software-oriented infrastructure depth. You have built and operated cloud platforms or distributed systems and can reason across infrastructure code, Kubernetes, networking, service identity, and stateful dependencies.

  • Safe-change and production judgment. You have owned consequential migrations and incidents, can explain failure modes and rollback limits, and know when simplifying a system is better than adding another platform.

  • Platform-product and engineering judgment. You understand internal customers, create interfaces other teams adopt, and make clear tradeoffs among reliability, developer autonomy, engineering effort, and workload efficiency.

Nice to Have

  • Experience with multi-tenant, cellular, regional, or dedicated enterprise infrastructure.

  • Familiarity with GCP/GKE, Terraform or similar IaC systems, Cloudflare, Envoy/Istio, SPIFFE/SPIRE, and managed data services.

  • Experience with large-fleet rightsizing, infrastructure consolidation, or migrating CI compute without disrupting developer workflows.

  • A track record using AI tools to increase engineering output while preserving production safeguards.

Full-Time Employee Benefits Include:

💰 Competitive Salary & Equity

💹 401(k) Program with a 4% match (US Only)

⚕️ Health, Dental, Vision and Life Insurance

🩼 Short Term and Long Term Disability

🚼 Paid Parental, Medical, Caregiver Leave

🏝 Flexible Time Off (FTO) + Holidays

🚗 Commuter Benefits (In-Office & US Only)

📱 Monthly Wellness Stipend

🧑‍💻 Autonomous Work Environment

🖥 In Office Set-Up Reimbursement (In-Office Only)

🚀 Quarterly Team Gatherings

☕ In Office Amenities (In-Office Only)

Want to learn more about what we are up to?

Interviewing + Culture at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Show full description