Source description
About the role
Inference Engineer – LLM & Speech AI
Bengaluru, India
AI Infrastructure
In office
Full-time
We are looking for an experienced Inference Engineer to build and optimize high-performance inference systems for Large Language Models (LLMs), Speech AI systems, and multimodal AI workloads.
You will work on deploying production-grade AI systems with a strong focus on
-
low latency,
-
high throughput,
-
GPU efficiency,
-
scalable serving infrastructure,
-
distributed inference,
-
and cost optimization.
-
This role sits at the intersection of:
-
systems engineering,
-
deep learning infrastructure,
-
distributed computing,
-
and production AI deployment.
You will collaborate closely with
-
ML researchers,
-
platform engineers,
-
speech AI teams,
-
and product engineering teams.
-
Ready to apply?
-
Powered by
-
First name *
-
Last name *
-
Email *
-
LinkedIn URL *
-
Phone number *
-
Location *
-
Resume *
-
Click to upload or drag and drop here
-
Cover letter
-
Click to upload or drag and drop here
-
By applying you agree to Gem's terms and privacy policy.
-
Save your info to apply to other roles faster & help employers reach you.
-
Apply and saveApply without saving
-
Req ID: Soket 1
More at Soket AI
