Inference Performance Engineer

San Francisco
Hybridmid
Confirmed open4 days ago
First seen2 months ago
Remoet re-checks every listing roughly once a day.
Apply on AdaptionYou will apply on Adaption's own site.Star Adaption to hear about new roles →

Mid-level Inference Performance Engineer role using vLLM, SGLang, and TensorRT-LLM in a hybrid San Francisco office.

Summary, tech stack, seniority and salary here are extracted or inferred by Remoet, not the employer's own words.

Tech stack

PythonC++RustvLLMSGLangTensorRT-LLMCUDANCCL

Benefits

Flexible workAnnual travel stipendWeekly meal allowanceComprehensive medical benefitsGenerous paid time off
Confirmed open4 days ago
First seen2 months ago
Remoet re-checks every listing roughly once a day.
Apply on AdaptionYou will apply on Adaption's own site.Star Adaption to hear about new roles →