ML-Ops / Platform Engineer
@ Long Finch TechnologiesML-Ops / Platform Engineer
About the job
The company specializes in building scalable, secure, and automated ML/GenAI solutions across cloud platforms, focusing on deploying enterprise-grade AI infrastructure and workloads.
Requirements
- MLOps experience
- Cloud platforms (AWS/Azure)
- Kubernetes and Docker
- Terraform and CI/CD pipelines
- Knowledge of GenAI/LLM
Qualifications
- Degree in computer science
- Strong scripting skills
- Experience with cloud infrastructure
Full job description
Must have skills: MLOps, AWS/Azure, Kubernetes, Docker, Python, Terraform, CI/CD, GenAI/LLM, RAG, Bedrock/Azure OpenAI, Vector DB, Observability
Responsibilities:
· Designed, deployed, and operated enterprise-grade MLOps/GenAI platforms across AWS and Azure, leveraging AWS Bedrock, SageMaker, Azure OpenAI, Azure AI Foundry, model serving, embeddings, RAG, vector databases, AI gateways, guardrails, and agentic frameworks.
· Built and automated cloud-native ML infrastructure using Terraform, Kubernetes (EKS/AKS/OpenShift), Docker, GitHub Actions, Azure DevOps, Jenkins, GitOps, and ArgoCD, enabling scalable model deployment, CI/CD, versioning, and release management.
· Implemented secure and highly available ML/AI platforms using AWS IAM, Azure IAM/RBAC, VPC/VNet, Key Vault, Secrets Manager, API Gateway, load balancers, ingress controllers, service mesh, autoscaling, and multi-account/subscription architectures.
· Developed and operationalized ML/GenAI workloads using Python, REST APIs, microservices, MongoDB, PostgreSQL, Redis, and vector databases, implementing model evaluation, prompt engineering, state management, caching, and high-throughput inference capabilities.
· Monitored, troubleshot, and optimized production ML/AI workloads using observability, logging, monitoring, SRE practices, performance tuning, resiliency, disaster recovery, and cost optimization, while collaborating with application, platform, infrastructure, and security teams.