Inference Performance Engineer

San Francisco
Remotemid
Confirmed open1 hour ago
First seentoday
Remoet re-checks every listing roughly once a day.
Apply on AdaptionYou will apply on Adaption's own site.Star Adaption to hear about new roles →

Mid-level inference performance engineer role focusing on vLLM and CUDA, offering a fully remote global work environment.

Summary, tech stack, seniority and salary here are extracted or inferred by Remoet, not the employer's own words.

Tech stack

PythonC++RustvLLMSGLangTensorRT-LLMCUDANCCL

Benefits

Flexible workTeam offsitesAnnual travel stipendWeekly meal allowanceComprehensive medical benefitsGenerous paid time off
Confirmed open1 hour ago
First seentoday
Remoet re-checks every listing roughly once a day.
Apply on AdaptionYou will apply on Adaption's own site.Star Adaption to hear about new roles →