---
title: "Applied AI Engineer, Evals"
company: "Sable"
company_url: "https://www.remjobs.works/companies/sable"
url: "https://www.remjobs.works/job/sable-applied-ai-engineer-evals-74d97a75-bfde-439b-b298-30a38395f366"
apply_url: "https://jobs.ashbyhq.com/sable/36addef7-4840-4a5e-bbcb-69502776801c"
workplace: onsite
location: "San Francisco"
employment_type: full-time
seniority: mid
role: ai-machine-learning
region: united-states
skills: ["llm"]
date_posted: 2026-09-15T21:23:45.563Z
first_seen_by_remjobs: 2026-09-15T21:31:40.470Z
---

# Applied AI Engineer, Evals

**Sable** · San Francisco

Apply: https://jobs.ashbyhq.com/sable/36addef7-4840-4a5e-bbcb-69502776801c

## About Sable

Sable is a digital banking platform and the only full-suite banking solution for internationals in the U.S., providing U.S. bank accounts, debit and credit cards. Newcomers to the U.S. can open a Sable account within five minutes, and start building their U.S. credit history from day 1, all without the need of an SSN or prior U.S. credit history. The Sable app is available on iOS and Android.

## About the role

#### About Sable

Sable built Aidan, the first AI employee who can lead customer calls using realtime voice, vision, and browser use. Aidan runs a live, two-way conversation inside a real product environment, clicking through the product like a human, watching the user's screen, and adapting the journey on the fly. Every conversation feeds a self-improving context graph we call the Brain, so Aidan gets smarter with each call.

#### The role

You turn "was that a good call" into clear, defensible signals, and then use those signals to make Aidan better. You define what good looks like across voice, vision, and browser-use, build the graders and environments that measure it, and carry what you learn back into the agent, the context, the data, and the models. Your numbers decide what ships and what's next.

#### What you'll do

- Design evaluation criteria for voice, vision, and browser actions and validate that LLM judges apply them confidently before we rely on them

- Build and maintain our internal benchmarks: accuracy, latency, and interaction quality.

- Build the harness that scores the whole runtime on customer-specific configurations and journeys

- Score live calls and identify gaps in current capabilities

- Use what the evaluations find to propose runtime, prompt, and model changes, and measure them

#### Who you are

- Engineer with applied ML experience: you have built evaluations, judges, or benchmarks for LLM or multimodal systems and know the difference between a metric that moves and a metric that means something

- Comfortable in a production codebase; you ship the harness, not a notebook

- Skeptical by default, with a habit of checking the wire before trusting a summary

- Bonus: experience with voice, realtime, or computer-use agents; post-training or RL environment work

---

Source: Sable's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/sable-applied-ai-engineer-evals-74d97a75-bfde-439b-b298-30a38395f366
