---
title: "Member of Technical Staff (Data Scientist, Evals)"
company: "Perplexity"
company_url: "https://www.remjobs.works/companies/perplexity"
url: "https://www.remjobs.works/job/perplexity-member-of-technical-staff-data-scientist-evals-e8e38619-adab-4b64-b21e-6ef0e76cf3b4"
apply_url: "https://jobs.ashbyhq.com/perplexity/4615ca06-bea7-47e3-9e57-f5cee52b75e6"
workplace: hybrid
location: "San Francisco"
employment_type: full-time
seniority: staff
role: data
region: united-states
skills: ["aws", "databricks", "llm", "python", "sql"]
date_posted: 2026-06-29T13:45:56.976Z
first_seen_by_remjobs: 2026-08-26T15:14:51.821Z
---

# Member of Technical Staff (Data Scientist, Evals)

**Perplexity** — San Francisco

Apply: https://jobs.ashbyhq.com/perplexity/4615ca06-bea7-47e3-9e57-f5cee52b75e6

## About Perplexity

Perplexity LLC is an Executive Coaching firm for SME

## About the role

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases. In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users.

##### Responsibilities

- Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness

- Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality

- Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices

- Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements

- Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality

##### Qualifications

- PhD or MS in a technical field or equivalent experience

- 4+ years of experience in data science or machine learning

- Strong proficiency in Python and SQL (expected to write production-grade code)

- Experience building within a modern cloud data stack, specifically AWS and Databricks

- Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster

##### Preferred Qualifications

- 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups

- Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale

- A strong research background, with experience applying research methods to real-world ML problems

- Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets

---

Source: Perplexity's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/perplexity-member-of-technical-staff-data-scientist-evals-e8e38619-adab-4b64-b21e-6ef0e76cf3b4
