Home › Companies › Perplexity › Engineering Manager (AI Inference)

Engineering Manager (AI Inference)

Perplexity · San Francisco · Active · Ashby

Job facts

Field	Value
Company	Perplexity
Title	Engineering Manager (AI Inference)
Normalized title	-
Department / team	AI / AI
Location	San Francisco, CA, United States
Work model	-
Employment type	Full Time
Salary	-
Status	active
ATS provider	Ashby
Posted / first seen	— / 2026-05-29
Changed / last seen	2026-05-29 / 2026-06-04

Related slices

Page	What it contains	Open
Company jobs	Active postings from Perplexity.	Open
Company breakdowns	Role, location, ATS, and work model facets for this company.	Open
ATS provider jobs	Active postings observed through Ashby.	Open
Provider filtered search	The same provider as a filtered job collection.	Open
City jobs	Active postings in San Francisco.	Open
Department jobs	Active postings in AI.	Open
Lifecycle events	Open, update, close, and reopen events for this posting.	Open
Original posting	Canonical source or apply URL captured from the ATS.	Open

Linked records

Company	Perplexity
Source	9e1a7911-2863-49e5-b7be-114bf50b7e20
ATS provider	Ashby

Description

About the Role We are looking for an Inference Engineering Manager to lead our AI Inference team. This is a unique opportunity to build and scale the infrastructure that powers Perplexity's products and APIs, serving millions of users with state-of-the-art AI capabilities. You will own the technical direction and execution of our inference systems while building and leading a world-class team of inference engineers. Our current stack includes Python, PyTorch, Rust, C++, and Kubernetes. You will help architect and scale the large-scale deployment of machine learning models behind Perplexity's Comet, Sonar, Search, Deep Research products. Why Perplexity? Build SOTA systems that are the fastest in the industry with cutting-edge technology High-impact work on a smaller team with significant ownership and autonomy Opportunity to build 0-to-1 infrastructure from scratch rather than maintaining legacy systems Work on the full spectrum: reducing cost, scaling traffic, and pushing the boundaries of inference Direct influence on technical roadmap and team culture at a rapidly growing company Responsibilities Lead and grow a high-performing team of AI inference engineers Develop APIs for AI inference used by both internal and external customers Architect and scale our inference infrastructure for reliability and efficiency Benchmark and eliminate bottlenecks throughout our inference stack Drive large sparse/MoE model inference at rack scale, including sharding strategies for massive models Push the frontier with building inference systems to support sparse attention, disaggregated pre-fill/decoding serving, etc. Improve the reliability and observability of our systems and lead incident response Own technical decisions around batching, throughput, latency, and GPU utilization Partner with ML research teams on model optimization and deployment Recruit, mentor, and develop engineering talent Establish team processes, engineering standards, and operational excellence Qualifications 5+ years of engineering experience with 2+ years in a technical leadership or management role Deep experience with ML systems and inference frameworks (PyTorch, TensorFlow, ONNX, TensorRT, vLLM) Strong understanding of LLM architecture: Multi-Head Attention, Multi/Grouped-Query Attention, and common layers Experience with inference optimizations: batching, quantization, kernel fusion, FlashAttention Familiarity with GPU characteristics, roofline models, and performance analysis Experience deploying reliable, distributed, real-time systems at scale Track record of building and leading high-performing engineering teams Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism Strong technical communication and cross-functional collaboration skills Nice to Have Experience with CUDA, Triton, or custom kernel development Background in training infrastructure and RL workloads Experience with Kubernetes and container orchestration at scale Published work or contributions to inference optimization research

Full job record

Job ID	79306e6df2025f66522cac552b68df6e012ac298
Org ID	22236078-2ac1-4479-bbc4-5ae282c73695
Source ID	9e1a7911-2863-49e5-b7be-114bf50b7e20
Board ID	9e1a7911-2863-49e5-b7be-114bf50b7e20
Provider	ashby
Provider Job Key	2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9
Title	Engineering Manager (AI Inference)
Normalized Title	—
Status	active
Active	yes
Location Text	San Francisco
Department	AI
Team	AI
Employment Type	full_time
Workplace Type	—
Remote Policy	—
Country	United States
Region	CA
City	San Francisco
Salary Raw	—
Salary Min	—
Salary Max	—
Salary Currency	—
Salary Period	—
Source URL	https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9
Apply URL	https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9/application
First Seen At	2026-05-29 06:19:18Z
Last Seen At	2026-06-04 13:33:40Z
Last Checked At	2026-06-04 13:33:40Z
Last Changed At	2026-05-29 06:19:18Z
Inactive At	—
Source Posted At	—
Source Updated At	—
Raw Payload Uri	s3://bluework-jobs-prod-raw-590183727216/raw/provider=ashby/board=perplexity/date=2026-06-04/2026-06-04T13-33-00-183Z-c110938004660706490690598eb41b6d9d51ed6139f9cc802a90dc553598c7f5.json

Event Fields

{
  "content_hash": "33033f5385226c73a1505b8154d702b9945b9a207aa0c7e6491fd1e49ceed534",
  "source_hash": "fea6f71208a67681621b604c6ff44a5b203d27e2e572286c83dc9ef6423d4a04",
  "last_changed_at": "2026-05-29T06:19:18.716Z",
  "active_status": "active"
}

Parsed Structured

{
  "language": "en",
  "location": {
    "raw": "San Francisco",
    "city": "San Francisco",
    "region": "CA",
    "country": "United States",
    "is_remote": false,
    "confidence": 0.75
  },
  "salary_max": null,
  "salary_min": null,
  "inferred_at": "2026-06-04T13:33:40.662Z",
  "launch_scope": {
    "reason": "english_us_canada",
    "included": true,
    "language": "en",
    "location": {
      "raw": "San Francisco",
      "city": "San Francisco",
      "region": "CA",
      "country": "United States",
      "is_remote": false,
      "confidence": 0.75
    },
    "countries": [
      "United States"
    ]
  },
  "remote_policy": null,
  "salary_period": null,
  "workplace_type": null,
  "salary_currency": null
}

Extensions

{}

Native Structured

{
  "id": "2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9",
  "team": "AI",
  "title": "Engineering Manager (AI Inference)",
  "jobUrl": "https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9",
  "address": null,
  "applyUrl": "https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9/application",
  "isListed": true,
  "isRemote": false,
  "location": "San Francisco",
  "updatedAt": null,
  "apiVersion": "ashby-non-user-graphql-v1",
  "department": "AI",
  "publishedAt": null,
  "workplaceType": null,
  "employmentType": "FullTime",
  "secondaryLocations": []
}

Get this page with API

Rendered from the bluedoor Job Postings API. Reproduce it:

GET https://api.bluedoor.sh/job-postings/v1/jobs/79306e6df2025f66522cac552b68df6e012ac298?include=descriptionJSON

GET https://api.bluedoor.sh/job-postings/v1/orgs/22236078-2ac1-4479-bbc4-5ae282c73695JSON

GET https://api.bluedoor.sh/job-postings/v1/sources/9e1a7911-2863-49e5-b7be-114bf50b7e20JSON

GET https://api.bluedoor.sh/job-postings/v1/jobs/79306e6df2025f66522cac552b68df6e012ac298/eventsJSON

Docs · Get an API key