More jobs:
LLM Engineer
Job in
Cincinnati, Hamilton County, Ohio, 45202, USA
Listed on 2026-08-31
Listing for:
TALENT Software Services
Full Time
position Listed on 2026-08-31
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
LLM Engineer
Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities.
Build secure, reliable, reusable, and enterprise-ready AI capabilities.
Support:
Agentic AI workflows, AI for SDLC, Knowledge retrieval, Model evaluation, Private AI hosting, Agent Ops.
Work closely with:
Principal AI Architect, AI Engineering Lead, Platform Engineers, Security teams, Enterprise Architecture, Product Owners, Domain teams.
Must-Have Technical Skills
LLM & Generative AI- Large Language Models (LLMs)
- Small Language Models (SLMs)
- Prompt Engineering
- Context Engineering
- Retrieval-Augmented Generation (RAG)
- Embeddings
- Semantic Search
- Agentic AI Patterns
- Multi-Agent Workflows
- Tool Calling
- Function Calling
- Model Evaluation
- LLM Observability
- Fine-Tuning
- Supervised Fine-Tuning
- LoRA
- QLoRA
- Quantization
- Distillation
- Model Compression
- Synthetic Data Generation
- Model Benchmarking
- Model Selection
- Model Routing
- Private LLM Hosting
- On-Prem Model Deployment
- GPU-Based Inference
- Model Serving APIs
- High-Availability Inference
- Autoscaling
- Load Balancing
- Caching
- Batch and Real-Time Inference
- Kubernetes
- Docker
- Kubeflow
- KServe
- Ray Serve
- MLflow
- Hugging Face
- Transformers
- Py Torch
- PEFT
- Deep Speed
- NVIDIA NIM
- Triton Inference Server
- TensorRT-LLM
- vLLM
- TGI
- SGLang
- Python
- Type Script or Java Script
- REST APIs
- Microservices
- CI/CD
- Git Hub or Azure Dev Ops
- API Design
- Distributed Systems
- Cloud-Native Engineering
- Test Automation
- Vector Databases
- Knowledge Graphs
- Document Processing
- Metadata Management
- Data Pipelines
- Object Storage
- Enterprise Search
- Structured and Unstructured Data Integration
- Build enterprise-grade LLM-powered applications and intelligent agent capabilities.
- Design reusable LLM patterns, services, APIs, and accelerators.
- Develop model interaction patterns for:
Reasoning, Summarization, Classification, Extraction, Planning, Decision support. - Build reusable prompt, context, retrieval, memory, and evaluation components.
- Support AI-for-SDLC agents across:
Requirements, Design, Coding, Testing, Security Review, Deployment, Operations. - Convert AI use cases into scalable production solutions.
- Build core intelligence services for enterprise agents.
- Develop reusable capabilities for:
Planning, Task decomposition, Reasoning, Tool usage, Agent collaboration. - Enable agent-to-agent interaction and multi-agent orchestration.
- Integrate LLMs with:
Agent runtimes, Tool registries, Workflow engines, MCP-based gateways. - Support human-in-the-loop, approval, escalation, and feedback workflows.
- Improve agent quality, accuracy, safety, and task completion.
- Design reusable prompt engineering standards, templates, and libraries.
- Create:
System prompts, Task prompts, Role prompts, Guardrail prompts, Evaluation prompts. - Develop context engineering strategies for better grounding, relevance, and personalization.
- Optimize:
Token usage, Context windows, Memory injection, Retrieval inputs. - Establish prompt versioning, testing, and governance practices.
- Design and implement enterprise RAG architectures.
- Build retrieval pipelines using:
Enterprise documents, Knowledge repositories, Structured data, Metadata. - Optimize:
Chunking, Embeddings, Indexing, Ranking, Reranking, Retrieval strategies. - Improve grounding, citation quality, precision, recall, and factual accuracy.
- Build reusable retrieval services for agents and business domains.
- Partner with data and knowledge management teams to onboard trusted data sources.
- Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs.
- Support domain-specific model development using approved datasets.
- Build supervised fine-tuning and model adaptation pipelines.
- Apply:
LoRA, QLoRA, Distillation, Quantization, Model compression. - Evaluate commercial, open-source, and internally hosted models.
- Select models based on:
Accuracy, Latency, Cost, Data residency, Security, Operational requirements.
- Build and support private AI capabilities for LLM/SLM hosting.
- Deploy models across:
On-premises, Hybrid, Private cloud environments. - Support GPU-enabled model hosting.
- Optimize latency, throughput, concurrency, resiliency, and GPU utilization.
- Build secure inference endpoints for internal applications and agents.
- Support air-gapped and restricted AI environments.
- Partner with infrastructure and platform teams on private AI hosting.
- Implement scalable model serving using modern inference frameworks.
- Build high-availability inference architectures.
- Optimize:
Token throughput, Response latency, Cost efficiency, Inference performance. - Implement:
Model routing, Load balancing, Caching, Fallback strategies. - Support batch and real-time inference.
- Develop reusable deployment templates for different model families.
- Build operational…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×