Home › Companies › Perplexity › Engineering Manager (AI Inference)
Engineering Manager (AI Inference)
Perplexity · San Francisco · Active · Ashby
Job facts
| Field | Value |
|---|---|
| Company | Perplexity |
| Title | Engineering Manager (AI Inference) |
| Normalized title | - |
| Department / team | AI / AI |
| Location | San Francisco, CA, United States |
| Work model | - |
| Employment type | Full Time |
| Salary | - |
| Status | active |
| ATS provider | Ashby |
| Posted / first seen | — / 2026-05-29 |
| Changed / last seen | 2026-05-29 / 2026-06-04 |
Related slices
| Page | What it contains | Open |
|---|---|---|
| Company jobs | Active postings from Perplexity. | Open |
| Company breakdowns | Role, location, ATS, and work model facets for this company. | Open |
| ATS provider jobs | Active postings observed through Ashby. | Open |
| Provider filtered search | The same provider as a filtered job collection. | Open |
| City jobs | Active postings in San Francisco. | Open |
| Department jobs | Active postings in AI. | Open |
| Lifecycle events | Open, update, close, and reopen events for this posting. | Open |
| Original posting | Canonical source or apply URL captured from the ATS. | Open |
Linked records
| Company | Perplexity |
| Source | 9e1a7911-2863-49e5-b7be-114bf50b7e20 |
| ATS provider | Ashby |
Description
About the Role We are looking for an Inference Engineering Manager to lead our AI Inference team. This is a unique opportunity to build and scale the infrastructure that powers Perplexity's products and APIs, serving millions of users with state-of-the-art AI capabilities.
You will own the technical direction and execution of our inference systems while building and leading a world-class team of inference engineers. Our current stack includes Python, PyTorch, Rust, C++, and Kubernetes. You will help architect and scale the large-scale deployment of machine learning models behind Perplexity's Comet, Sonar, Search, Deep Research products.
Why Perplexity? Build SOTA systems that are the fastest in the industry with cutting-edge technology
High-impact work on a smaller team with significant ownership and autonomy
Opportunity to build 0-to-1 infrastructure from scratch rather than maintaining legacy systems
Work on the full spectrum: reducing cost, scaling traffic, and pushing the boundaries of inference
Direct influence on technical roadmap and team culture at a rapidly growing company
Responsibilities Lead and grow a high-performing team of AI inference engineers
Develop APIs for AI inference used by both internal and external customers
Architect and scale our inference infrastructure for reliability and efficiency
Benchmark and eliminate bottlenecks throughout our inference stack
Drive large sparse/MoE model inference at rack scale, including sharding strategies for massive models
Push the frontier with building inference systems to support sparse attention, disaggregated pre-fill/decoding serving, etc.
Improve the reliability and observability of our systems and lead incident response
Own technical decisions around batching, throughput, latency, and GPU utilization
Partner with ML research teams on model optimization and deployment
Recruit, mentor, and develop engineering talent
Establish team processes, engineering standards, and operational excellence
Qualifications 5+ years of engineering experience with 2+ years in a technical leadership or management role
Deep experience with ML systems and inference frameworks (PyTorch, TensorFlow, ONNX, TensorRT, vLLM)
Strong understanding of LLM architecture: Multi-Head Attention, Multi/Grouped-Query Attention, and common layers
Experience with inference optimizations: batching, quantization, kernel fusion, FlashAttention
Familiarity with GPU characteristics, roofline models, and performance analysis
Experience deploying reliable, distributed, real-time systems at scale
Track record of building and leading high-performing engineering teams
Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism
Strong technical communication and cross-functional collaboration skills
Nice to Have Experience with CUDA, Triton, or custom kernel development
Background in training infrastructure and RL workloads
Experience with Kubernetes and container orchestration at scale
Published work or contributions to inference optimization research
Full job record
| Job ID | 79306e6df2025f66522cac552b68df6e012ac298 |
| Org ID | 22236078-2ac1-4479-bbc4-5ae282c73695 |
| Source ID | 9e1a7911-2863-49e5-b7be-114bf50b7e20 |
| Board ID | 9e1a7911-2863-49e5-b7be-114bf50b7e20 |
| Provider | ashby |
| Provider Job Key | 2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9 |
| Title | Engineering Manager (AI Inference) |
| Normalized Title | — |
| Status | active |
| Active | yes |
| Location Text | San Francisco |
| Department | AI |
| Team | AI |
| Employment Type | full_time |
| Workplace Type | — |
| Remote Policy | — |
| Country | United States |
| Region | CA |
| City | San Francisco |
| Salary Raw | — |
| Salary Min | — |
| Salary Max | — |
| Salary Currency | — |
| Salary Period | — |
| Source URL | https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9 |
| Apply URL | https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9/application |
| First Seen At | 2026-05-29 06:19:18Z |
| Last Seen At | 2026-06-04 13:33:40Z |
| Last Checked At | 2026-06-04 13:33:40Z |
| Last Changed At | 2026-05-29 06:19:18Z |
| Inactive At | — |
| Source Posted At | — |
| Source Updated At | — |
| Raw Payload Uri | s3://bluework-jobs-prod-raw-590183727216/raw/provider=ashby/board=perplexity/date=2026-06-04/2026-06-04T13-33-00-183Z-c110938004660706490690598eb41b6d9d51ed6139f9cc802a90dc553598c7f5.json |
Event Fields
{
"content_hash": "33033f5385226c73a1505b8154d702b9945b9a207aa0c7e6491fd1e49ceed534",
"source_hash": "fea6f71208a67681621b604c6ff44a5b203d27e2e572286c83dc9ef6423d4a04",
"last_changed_at": "2026-05-29T06:19:18.716Z",
"active_status": "active"
}Parsed Structured
{
"language": "en",
"location": {
"raw": "San Francisco",
"city": "San Francisco",
"region": "CA",
"country": "United States",
"is_remote": false,
"confidence": 0.75
},
"salary_max": null,
"salary_min": null,
"inferred_at": "2026-06-04T13:33:40.662Z",
"launch_scope": {
"reason": "english_us_canada",
"included": true,
"language": "en",
"location": {
"raw": "San Francisco",
"city": "San Francisco",
"region": "CA",
"country": "United States",
"is_remote": false,
"confidence": 0.75
},
"countries": [
"United States"
]
},
"remote_policy": null,
"salary_period": null,
"workplace_type": null,
"salary_currency": null
}Extensions
{}Native Structured
{
"id": "2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9",
"team": "AI",
"title": "Engineering Manager (AI Inference)",
"jobUrl": "https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9",
"address": null,
"applyUrl": "https://jobs.ashbyhq.com/perplexity/2a87ccbf-82ef-4fc7-b1ed-4dd18b11baf9/application",
"isListed": true,
"isRemote": false,
"location": "San Francisco",
"updatedAt": null,
"apiVersion": "ashby-non-user-graphql-v1",
"department": "AI",
"publishedAt": null,
"workplaceType": null,
"employmentType": "FullTime",
"secondaryLocations": []
}Get this page with API
Rendered from the bluedoor Job Postings API. Reproduce it:
GET https://api.bluedoor.sh/job-postings/v1/jobs/79306e6df2025f66522cac552b68df6e012ac298?include=descriptionJSONGET https://api.bluedoor.sh/job-postings/v1/orgs/22236078-2ac1-4479-bbc4-5ae282c73695JSONGET https://api.bluedoor.sh/job-postings/v1/sources/9e1a7911-2863-49e5-b7be-114bf50b7e20JSONGET https://api.bluedoor.sh/job-postings/v1/jobs/79306e6df2025f66522cac552b68df6e012ac298/eventsJSON