Site Reliability Engineer (SRE) / DevOps Engineer
@ E-SpaceSite Reliability Engineer (SRE) / DevOps Engineer
This job is still taking applications, but it's been up a while.
About the job
E-Space aims to create sustainable, accessible space connectivity via low Earth orbit satellites. We disrupt traditional space tech, focusing on reliable software systems, cloud infrastructure, and innovation to advance satellite capabilities and space sustainability.
Requirements
- 5+ years experience in SRE or DevOps
- Proven AWS mission-critical system design
- Automation scripting in Python Bash
- Experience with Kubernetes, Terraform
- Strong CI/CD pipeline skills
Qualifications
- Experience with RDS, Aurora
- Knowledge of cloud security and compliance
- Experience with monitoring tools
- Deep understanding of infrastructure as code
- Passion for space and satellite tech
Full job description
What you will be doing:
- Design, deploy, and maintain highly-scalable, highly-available software systems in AWS
- Architect and manage containerized applications on Amazon EKS with focus on reliability and performance
- Build and maintain Infrastructure as Code using Terraform for AWS cloud resources
- Develop and optimize CI/CD pipelines for automated testing, deployment, and rollback capabilities
- Implement comprehensive monitoring, alerting, and observability solutions using CloudWatch, Prometheus, and Grafana
- Ensure system reliability through SLI/SLO definition, error budgets, and incident response procedures
- Collaborate directly with engineering teams to optimize application deployment and operations
- Manage deployments and scaling strategies to support mission-critical operations
- Automate and enforce cloud security, governance, and compliance controls
- Participate in on-call rotation and lead incident response for production level systems
What you bring to this role:
- 5+ years of experience in SRE, DevOps, or Platform Engineering roles
- Proven experience designing and operating mission-critical, highly-available systems within AWS
- Advanced proficiency in Infrastructure as Code using Terraform (OpenTofu)
- Deep experience with Kubernetes, EKS, Helm, and container orchestration
- Strong CI/CD pipeline development and management experience (Bitbucket preferred)
- Proficiency in Python and Bash scripting for automation
- Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack)
- Knowledge of capacity planning and performance optimization
- Experience with database operations and scaling (RDS, Aurora, or similar)
Extra bonus points for the following:
- AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA), or equivalent expertise
- Experience with incident management and post-mortem processes
- Experience with GitOps workflows and tools (ArgoCD, Flux)
- Knowledge of service mesh technologies (Istio, Linkerd)
- Experience with chaos engineering and disaster recovery planning
- Experience with Zero Trust Networking (ZTNA) or VPN solutions
- Background in aerospace, defense, or other mission-critical industries
- Strong intellectual curiosity and commitment to continuous learning
- Exceptional attention to detail and an ownership mentality
Additional Requirements
Similar jobs in Columbus, OH
- L
Site Reliability Engineer
Leidos · Columbus, OH, United States
Posted 3 weeks ago - L
Senior Security Engineer
Leidos · Columbus, OH, United States
Posted 2 weeks ago - J
HVAC Controls Engineer – Building Automation
Jobot · Columbus, OH, United States
Posted 1 day ago - A
Site Strategy Manager, Data Center Sites - Remote (U.S.)
AECOM · Columbus, OH, United States
Posted 3 weeks ago - S
Seeking experienced babysitter for toddler with own car & reliability!
Sittercity · Columbus, OH, United States
Posted 3 days ago - A
Roadway Engineer
AMT Engineering · Columbus, OH
Posted 3 weeks ago