Applied AI Engineer, Evals
- Onsite
- All ai & machine learning jobs
- FullTime
- Engineering
About the role
About Sable
Sable built Aidan, the first AI employee who can lead customer calls using realtime voice, vision, and browser use. Aidan runs a live, two-way conversation inside a real product environment, clicking through the product like a human, watching the user's screen, and adapting the journey on the fly. Every conversation feeds a self-improving context graph we call the Brain, so Aidan gets smarter with each call.
The role
You turn "was that a good call" into clear, defensible signals, and then use those signals to make Aidan better. You define what good looks like across voice, vision, and browser-use, build the graders and environments that measure it, and carry what you learn back into the agent, the context, the data, and the models. Your numbers decide what ships and what's next.
What you'll do
Design evaluation criteria for voice, vision, and browser actions and validate that LLM judges apply them confidently before we rely on them
Build and maintain our internal benchmarks: accuracy, latency, and interaction quality.
Build the harness that scores the whole runtime on customer-specific configurations and journeys
Score live calls and identify gaps in current capabilities
Use what the evaluations find to propose runtime, prompt, and model changes, and measure them
Who you are
Engineer with applied ML experience: you have built evaluations, judges, or benchmarks for LLM or multimodal systems and know the difference between a metric that moves and a metric that means something
Comfortable in a production codebase; you ship the harness, not a notebook
Skeptical by default, with a habit of checking the wire before trusting a summary
Bonus: experience with voice, realtime, or computer-use agents; post-training or RL environment work
Description as published by Sable.