Senior Site Reliability Engineer – AI & Automation.

@ Neshent Technologies
Neshent Technologiesneshenttechnologies.com

Senior Site Reliability Engineer – AI & Automation.

Medley, Florida
Posted today

About the job

Our team designs AI agents, automation, and cloud solutions to improve reliability and operational efficiency. We support modern systems and mentor engineers in SRE, DevOps, and observability practices.

Requirements

  • 5+ years SRE or DevOps experience
  • Strong Python skills for automation
  • Experience with AI agents and LLMs
  • Hands-on cloud infrastructure experience
  • Experience with observability tools

Qualifications

  • Effective communicator and mentorer
  • Able to work in 24x7 environment

Full job description

Primary Responsibilities

*Design and implement AI agents, LLM-based solutions, and operational automation for SRE and DevOps environments.
*Support production systems, perform incident response, troubleshooting, and 24x7 operational triage.
*Develop automation and integrations using Python.
*Build and maintain monitoring and observability solutions using Splunk, Datadog, New Relic, or AppDynamics.
*Design and manage cloud infrastructure across AWS, Azure, and/or GCP.
*Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, and/or AWS CDK.
*Apply AI/ML, prompt engineering, and agentic AI concepts to operational use cases.
*Develop runbook automation, anomaly detection, and intelligent incident management capabilities.
*Collaborate with engineering, operations, and business stakeholders to improve system reliability and operational processes.
*Mentor engineers and help establish SRE, DevOps, automation, and observability best practices.

Required Skills

*5+ years of SRE, DevOps, or Site Reliability Engineering experience.
*Strong Python development skills for automation, tooling, and integrations.
*Experience with AI agents, LLMs, and agentic AI frameworks.
*Experience with Splunk, Datadog, New Relic, AppDynamics, or similar observability platforms.
*Hands-on experience with AWS, Azure, and/or GCP.
*Experience with Terraform, CloudFormation, and/or CDK.
*Strong production support and incident response experience.
*Knowledge of AI/ML, prompt engineering, and operational AI automation.
*Strong communication, stakeholder management, and mentoring skills.
*Ability to work effectively in a 24x7 operational environment.

Preferred Skills

*Experience with Adobe Experience Manager (AEM).
*Experience supporting CMS platforms and digital applications.
*Knowledge of incident management, runbook automation, and anomaly detection.
*Familiarity with Atlassian Rovo, AI operational tooling, and modern observability platforms.

Show full description