GenAI Engineer – LLM Infrastructure & Inference Services
@ 2T ConsultingGenAI Engineer – LLM Infrastructure & Inference Services
About the job
Our company develops advanced AI platforms focusing on LLM infrastructure, deployment, and high-performance inference. We build scalable enterprise GenAI solutions leveraging GPU and cloud tech, aiming to optimize AI model serving, deployment, and management for diverse applications.
Requirements
- Experience with GPU infrastructure
- Knowledge of inference services
- Proficiency in Python and APIs
- Cloud platform experience
- Familiarity with MLOps practices
Qualifications
- Bachelor's in Computer Science
- Experience in AI model deployment
- Strong coding and scripting skills
- Ability to collaborate with teams
Full job description
We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments.
Key Responsibilities
- Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.
- Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.
- Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.
- Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.
- Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.
- Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.
- Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.
- Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.
- Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.
Core Technologies
- LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve
- AI/GenAI: RAG, LangChain, LangGraph, Vector Databases
- Development: Python, FastAPI, Microservices
- Infrastructure: Kubernetes, Docker, GPU Infrastructure
- Cloud: AWS, Azure, GCP
- MLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance
Similar jobs in Woodland Acres, California
- S
Principal Engineer
SYSTEMS PLANNING AND ANALYSIS, INC. · Dahlgren, VA, United States
Posted 4 days ago - K
Acoustics Engineer
KMS Solutions, LLC · Waldorf, MD, United States
Posted 5 days ago - N
GENERAL ENGINEER
Naval Sea Systems Command · Waldorf, MD, United States
Posted 3 days ago - N
ENGINEER/SCIENTIST
Naval Sea Systems Command · Dahlgren, VA, United States
Posted 3 days ago - M
Licensed Engineer
My Maritime Career · Piney Point, MD, 20674, US
Posted today - B
Wireless Engineer
BOOZ, ALLEN & HAMILTON, INC. · California, MD, United States
Posted 4 days ago