35 Backend Engineer- Inference Services jobs in Remote, California

Scroll to load more...

GenAI Engineer – LLM Infrastructure & Inference Services

2T Consulting

2T Consulting2tconsulting.com

GenAI Engineer – LLM Infrastructure & Inference Services

Monte Vista, California
Posted today

About the job

A company focused on building enterprise GenAI platforms, deploying large language models, and optimizing inference services across cloud and GPU infrastructure.

Requirements

  • Expertise in LLM infrastructure
  • Experience with model deployment
  • Knowledge of GPU infrastructure
  • Proficiency in Python and Kubernetes
  • Experience with cloud platforms

Qualifications

  • Bachelor's degree in relevant field
  • Strong problem-solving skills
  • Experience in AI and machine learning
  • Ability to optimize model performance
  • Excellent collaboration skills

Full job description

We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments.

Key Responsibilities

  • Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.
  • Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.
  • Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.
  • Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.
  • Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.
  • Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.
  • Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.
  • Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.
  • Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.

Core Technologies

  • LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve
  • AI/GenAI: RAG, LangChain, LangGraph, Vector Databases
  • Development: Python, FastAPI, Microservices
  • Infrastructure: Kubernetes, Docker, GPU Infrastructure
  • Cloud: AWS, Azure, GCP
  • MLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance

Show full description
Ask Duey about this job

Duey AI may make mistakes. Please double-check key information.