Senior Principal Agentic Engineer

@ 100 Eli Lilly and Company
100 Eli Lilly and Company100elilillyandcompany.com

Senior Principal Agentic Engineer

Indianapolis, IN
Posted 1 day ago

About the job

Lilly develops medicines to improve lives worldwide, building digital platforms supporting discovery, development, and delivery of innovative healthcare solutions. The role focuses on architecting agentic AI platforms, improving reliability, and scaling automation in a regulated environment.

Requirements

  • 7+ years in enterprise automation or cloud-native platforms
  • Experience with agentic AI solutions in production
  • Kubernetes, Terraform, cloud AI services
  • Hands-on Python and possibly Go development
  • CI/CD for AI workloads

Qualifications

  • Bachelor's in Computer Science or related field
  • Proven leadership in platform engineering
  • Experience in regulated industries
  • Strong security and observability skills
  • Familiarity with AI/ML operational tools

Full job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. 


About the technology organization 

Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable. 

Within Technology at Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise. 

About the Team 

Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business. 

The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms. 

The Reliability Engineering team is the engineering-first function that owns the stability, observability, and operational quality of a multi-application production estate. It operates in close partnership with the engineering team that builds the agentic automation platform, and is in active transition from human-executed operations to engineering-led, agent-assisted reliability. 

Role summary 

You own the agentic resolution platform end to end — both the cloud-native substrate it runs on and the intelligent agents that run on top of it. From Kubernetes, CI/CD, and observability through agent runtime, LLM patterns, retrieval, evaluation, and human-in-the-loop boundaries, you write the production code, design the architecture, and set the engineering bar that lets automation deflect routine operational work before it ever reaches a human. 

This is a senior individual-contributor role that is both strategic and deeply hands-on. You ship production platform services and agentic capabilities, make the build-vs-buy calls on tooling and model access, and partner with the senior architect on the patterns that scale. Success is measured by rising deflection, sustained toil reduction, platform availability and developer velocity, and the ability to scale agentic intelligence through systems and people — not individual effort. 

What you'll be doing 

1) Agentic resolution platform — intake to action 

  • Architect the agentic resolution platform end to end: intake, classification, action execution, verification, and human-in-the-loop fallback. 
  • Build the agent runtime and orchestration layer: agent state and memory, tool integration, multi-agent coordination patterns, and confidence-thresholded handoffs to humans. 
  • Define agent-decision observability, audit-ready posture, and the data contracts that let the platform consume durable fixes from upstream engineering teams. 

2) LLM application patterns, knowledge, and evaluation rigor 

  • Design production LLM patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, multi-model routing, and hybrid retrieval over the knowledge corpus. 
  • Own the knowledge-base strategy as a compounding deflection lever — every resolved incident becomes training data and structured retrieval input for future automation. 
  • Establish evaluation and guardrail frameworks for non-deterministic systems: automated evals, quality scoring, drift detection, and feedback loops that compound agent quality over time. 

3) Cloud-native platform — build and operate 

  • Architect and operate Kubernetes (EKS or equivalent) at scale for container and serverless workloads supporting agentic and LLM inference traffic patterns. 
  • Write production platform services and internal tooling (Python or Go) that automate provisioning, deployment, and operational workflows — not just infrastructure configuration. 
  • Define and maintain infrastructure as code (Terraform) integrated with a major cloud's AI stack (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI), with secrets management and audit-ready posture for regulated environments. 

4) CI/CD, observability, and developer experience 

  • Build and maintain CI/CD pipelines tuned for agentic and AI workloads: model and agent versioning, canary rollouts, evaluation gates, and rollback. 
  • Own the observability stack (Prometheus/Grafana/OpenTelemetry plus enterprise tooling) and instrument platform health, agent-decision telemetry, and model-inference metrics. 
  • Establish SLOs, SLIs, and reliability standards for both platform and agentic system health; design self-service patterns and golden-path templates that accelerate delivery. 

5) Cross-team partnership, security, and talent development 

  • Partner with the Reliability Engineering team on which production patterns become agent-assisted automations, and with senior architects on the patterns that scale. 
  • Own security posture: network policies, pod security, secrets rotation, vulnerability scanning, and access controls in a regulated pharmaceutical environment. 
  • Set the engineering bar through code quality standards, architectural reviews, and role modeling; mentor senior engineers in agentic AI, LLM application patterns, and platform engineering. Influence engineering leaders to adopt automation-friendly patterns at the source, not just downstream. 

How You Will Succeed:

  • Be recognized as the senior technical authority for agentic platform engineering in your area. 
  • Demonstrate measurable, sustained improvements: rising deflection rates, reduced toil, fewer recurring incidents, faster resolution, platform uptime, and deployment velocity. 
  • Ship production platform services and agentic capabilities that tangibly move the deflection and developer-experience numbers. 
  • Scale agentic intelligence through systems, standards, and people — not individual heroics. 

Your Basic Qualifications:

  • Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Information Science, Network Engineering, Systems Engineering, Computer Information Systems (CIS), Management Information Systems (MIS), Cloud Computing, Data Science
  • 7+ years of progressive technology experience with substantial hands-on architecture and delivery of automation, AIOps, agentic systems, or cloud-native platforms at enterprise scale. 
  • Demonstrated experience designing, building, and deploying agentic AI solutions in production environments, including agent runtimes, multi-agent orchestration, tool integration, and memory/state management, with hands-on implementation using at least one major framework (LangGraph, LangChain, LlamaIndex, or MCP).
  • Production experience with LLM application patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, and multi-model routing. 
  • Demonstrated experience designing, deploying, and supporting cloud-native solutions in production environments, including Kubernetes (EKS or equivalent), containerized workloads, infrastructure as code using Terraform, and implementation of AI services on at least one major cloud platform (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI).
  • Demonstrated experience designing, developing, and deploying production platform services and engineering tools, with hands-on software development expertise in Python (including asynchronous programming, packaging, and performance optimization) and the ability to deliver scalable, maintainable solutions beyond infrastructure configuration. Experience with Go is a plus.
  • Strong CI/CD engineering experience for AI and agentic workloads: model and agent deployment, versioning, canary/blue-green strategies, evaluation gates, and rollback. 
  • Demonstrated experience implementing evaluation and guardrail frameworks for AI and agentic systems, including automated evaluations, drift detection, human-in-the-loop verification, and feedback loops in production environments.
  • Demonstrated experience implementing security controls for cloud-native or production environments, including network policies, secrets management, vulnerability scanning, and access controls within regulated environments.

What You Should Bring:

  • Working knowledge of operational ML: anomaly detection, event correlation, alert-noise reduction, and model/agent monitoring in production. 
  • Experience with AI observability tooling (e.g., LangSmith, Weights & Biases, custom telemetry pipelines). 
  • Hands-on experience with ServiceNow integration patterns and chat-based intake platforms (e.g., Microsoft Teams). 
  • Experience with vector databases and hybrid retrieval architectures. 
  • Experience building internal developer platforms (IDP), self-service provisioning, or platform-as-a-product tooling. 
  • Backend development skills beyond infrastructure tooling: REST/gRPC API design, database design, message queues, distributed systems patterns. 
  • Demonstrated ownership of measurable deflection or operational-toil-reduction outcomes in a production environment — with the numbers to show it. 

  • Experience with policy-as-code, automated compliance checks, or shift-left security tooling. 
  • Observability for AI systems: metrics, logs, traces, and agent-decision telemetry using Prometheus/Grafana/OpenTelemetry, plus at least one enterprise stack (Splunk, Datadog, or Dynatrace). 

  • AWS cloud platform fluency, including EKS, Bedrock, SageMaker, and cost-optimization tooling. 
  • Familiarity with FinOps practices and cloud cost governance. 
  • Experience operating in highly regulated industries (life sciences, financial services, healthcare). 
  • Prior experience standing up a new engineering capability from a small founding team. 
  • Frontend experience (React, TypeScript) for internal dashboards and developer portals. 

Leadership expectations :

  • Acts with enterprise-first mindset, beyond individual products or teams. 
  • Drives accountability, clarity, and engineering rigor across the team. 
  • Builds trust through consistency, technical depth, and follow-through. 
  • Raises the capability of the organization, not just personal output. 
  • Leads through what they build and how they write, not through org-chart authority. 

Additional information 

Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.


Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.


Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).


Actual compensation will depend on a candidate’s education, experience, skills, and geographic location.  The anticipated wage for this position is

$129,000 - $231,000

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Show full description