Senior Principal SRE Engineering
@ 100 Eli Lilly and CompanySenior Principal SRE Engineering
About the job
Lilly is dedicated to creating medicines that make life better. The Senior Principal SRE leads reliability practices, standards, and incident response, ensuring scalable, secure operations in regulated environments, supporting global health initiatives.
Requirements
- 7+ years engineering experience
- 6+ years as SRE or similar
- Experience with observability tools
- Knowledge of infrastructure-as-code
- Experience in regulated environments
Qualifications
- Bachelor's in Computer Science or related
- Hands-on experience with SLO/SLI frameworks
- Experience with incident management
- Knowledge of cloud platforms
- Leadership in reliability practices
Full job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
About the technology organization
Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable.
Within Technology at Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise.
About the Team
Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business.
The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms.
The Reliability Engineering team is the engineering-first function that owns the stability, observability, and operational quality of a multi-application production estate. It operates in close partnership with the engineering team that builds the agentic automation platform, and is in active transition from human-executed operations to engineering-led, agent-assisted reliability.
Role Summary:
As the Senior Principal SRE Engineering Lead, you are the senior-most engineering authority for the reliability of the supported production estate. You own the bar for what reliable means in this organization: service-level objectives, error-budget governance, observability standards, and the engineering practices that protect both production and the team's engineering time.
This role is the judgment layer between an agent's recommendation and a production change. You combine engineering rigor with operational pragmatism, and you decide which patterns surfaced from production warrant durable engineering investment. You hold the SRE engineering bar within the function — the cross-pillar reference architecture is owned by the Senior Architect, but the practice, standards, and engineering judgment of reliability at this site are yours.
You are a senior individual contributor. You do not manage people. You partner with the Reliability Leader, the Senior Principal Tech Shift Lead in the same pillar, the Senior Architect, and the senior engineering principals in the agentic automation team. Success is measured by organizational impact, sustained reliability outcomes, and the ability to scale reliability through systems and people, not heroics.
What you'll be doing:
1) Reliability strategy and SLO governance
- Define service-level objectives and indicators across the supported production estate, tiered by application risk and business impact.
- Govern error-budget burn: when to slow change, when to invest in durable fixes, when to accept the budget.
- Drive the adoption of reliability reporting and the disciplined operating cadence that makes SLOs real, not decorative.
- Hold the SLO and error-budget conversation with product and application owners — including the conversation about what their service needs to change to meet the bar.
2) Engineering standards, observability, and self-healing
- Establish observability and instrumentation standards as the contract every supported application must meet — and hold the line on them.
- Set the bar for infrastructure-as-code, continuous-delivery hardening, and deployment safety across the supported estate.
- Define the engineering work that flows from incidents and root-cause analyses into durable production change, and the architectural patterns (self-healing runbooks, graceful degradation, circuit breakers) that reduce repeat failure.
3) Incident learning, durable fixes, and partnership with agentic automation
- Drive blameless postmortem culture; ensure root-cause analyses produce engineering work, not just narrative.
- Govern which patterns from production warrant durable engineering investment, against the error-budget regime.
- Partner with the production operations team in the same pillar on which patterns from the field warrant engineering attention; partner with the agentic automation engineering team on which fixes become safe agent-assisted remediations, and on the confidence thresholds, guardrails, and human-in-the-loop boundaries that make those remediations safe in production.
- Lead high-severity incident response as incident commander when escalation reaches this seat, and coach the team to handle the rest.
4) Compliance, security, and regulated-environment readiness
- Ensure reliability practices comply with Lilly standards and applicable regulatory requirements, and engineer them so that audit evidence falls out of normal operation.
- Promote secure operational practices, auditability, and validated-environment-friendly engineering as part of the standards the team holds.
- Act as a trusted technical leader in regulated and validated environments.
5) Technical leadership and talent development
- Set the engineering bar through standards, expectations, and role modeling.
- Mentor senior reliability engineers in the pillar; build the bench for sustained team growth and develop the next layer of principal engineers.
- Influence engineering, product, and platform leaders through credibility and outcomes rather than authority.
- Contribute to the evolution of enterprise-wide reliability practices in partnership with the Senior Architect and peer technical leaders.
How you will succeed
At the senior-most engineering individual-contributor level for reliability, success is defined by breadth of impact and sustained outcomes:
- Be recognized as the senior reliability authority for your area.
- Demonstrate measurable, sustained improvements such as: reduced major incidents, fewer recurring failures, improved time-to-recovery, and a credible error-budget regime.
- Influence decisions across multiple teams and leaders through expertise and trust.
- Scale reliability through systems, standards, and people, not heroics.
Your Basic Qualification:
- Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Information Science, Network Engineering, Systems Engineering, Computer Information Systems (CIS), Management Information Systems (MIS), Cloud Computing, Data Science
- 7+ years of progressive engineering experience, with at least 6 years as a Site Reliability Engineer, Production Engineer, or equivalent, including a tour as the senior-most reliability engineer for a multi-application production estate — not a single product.
- Hands-on experience authoring and validating runbooks: safe execution order, rollback steps, and exception handling for real remediation procedures. Demonstrated hands-on experience designing, implementing, and operating enterprise-scale SRE platforms, including observability solutions (Splunk, Datadog, New Relic, or Grafana/Prometheus), infrastructure-as-code with Terraform, CI/CD pipeline hardening, Kubernetes-based container platforms, and production workloads hosted on AWS, Azure, or GCP.
- Experience designing self-healing patterns (circuit breakers, graceful degradation, automated remediation) and validating them before they're trusted in production. Production reliability experience in a regulated or audited environment (GxP, SOX, HIPAA, PCI, or equivalent), including familiarity with change-control discipline, audit evidence, and validated-system constraints.
- Hands-on ownership experience of an SLO/SLI framework and error-budget policy across a multi-application estate, including defining SLIs, negotiating SLOs with product owners, and operationalizing burn-rate alerting and error-budget governance.
- Experience leading high-severity incident response as incident commander, running blameless postmortems, and converting findings into durable engineering work — measured by reduced recurrence rather than narrative quality.
- Demonstrated technical leadership experience at scale: mentoring senior ICs, influencing engineering and product leaders without org-chart authority, and writing the standards and engineering documents that set the bar for a discipline.
What You Should Bring:
- Hands-on experience designing self-healing automation and running chaos engineering or resilience-testing programs (AWS Fault Injection Service, Gremlin, LitmusChaos, or equivalent) tied to measurable reliability gains.
- Deep AWS fluency across reliability-relevant services (EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, Route 53), and familiarity with AWS Well-Architected Reliability Pillar.
- Experience with AIOps or agent-assisted operations, including designing the guardrails, confidence thresholds, and human-in-the-loop boundaries that make automated remediation safe in production.
- Track record of saying no to a deploy because the error budget was burned, and making the call stick.
- Prior experience building a reliability practice from a small founding team.
- Experience operating in pharma, healthcare, financial services, or other regulated industries.
Leadership expectations
- Acts as the judgment layer between aspiration and engineering reality.
- Combines engineering rigor with operational pragmatism, and knows when each is the right answer.
- Leads through what they build and how they write, not through org-chart authority.
- Comfortable telling a product owner that their application does not meet the reliability bar yet, and showing them how to get there.
- Treats mentorship of senior reliability engineers as a first-class outcome of the role.
Additional information
Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays.
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.
Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.
Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).
Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is
$129,000 - $231,000Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.
#WeAreLilly
Similar jobs in Indianapolis, IN
- 1
Principal SRE Engineer
100 Eli Lilly and Company · US, Indianapolis IN
Posted 1 day ago - 1
Senior Principal Agentic Engineer
100 Eli Lilly and Company · US, Indianapolis IN
Posted 1 day ago - I
Principal Packaging Engineer
Ingredion Incorporated (NA-US) · Indianapolis, IN
Posted 2 weeks ago - 1
Senior Engineering Advisor (SDD)
100 Eli Lilly and Company · US, Indianapolis IN
Posted 1 week ago - C
Senior Manufacturing Engineer
CFO Trane U.S. Inc.-CFO · Noblesville, Indiana
Posted 4 weeks ago - L
Senior Software Engineer
LE0021 Trimedx Holdings LLC · Indianapolis, IN
Posted 1 week ago