Job Description & How to Apply Below
vCluster Labs is hiring a Senior Inference Engineer to own the inference layer from ground up. You will deploy models to production on GPU infrastructure, stand up serving stacks with vLLM, SGLang, and TensorRT-LLM, and optimize latency and costs as traffic grows.
You will code in Python or Golang and shape the roadmap with our CTO. You will work closely with Product to decide what we build next and drive the engineering direction of inference across the company.
#J-18808-LjbffrNote that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×