---
title: "Member of Technical Staff, Alignment"
company: "Abundant"
company_url: "https://www.remjobs.works/companies/abundant"
url: "https://www.remjobs.works/job/abundant-member-of-technical-staff-alignment-59c3e1cf-db9f-457f-83f4-ff612611ed65"
apply_url: "https://jobs.ashbyhq.com/abundant/0d11f5c1-036f-4ef9-872f-f9407fb79d57"
workplace: onsite
location: "San Francisco"
employment_type: full-time
seniority: staff
role: software-engineering
region: united-states
skills: ["llm"]
date_posted: 2026-09-16T12:03:55.658Z
first_seen_by_remjobs: 2026-09-16T12:19:20.453Z
---

# Member of Technical Staff, Alignment

**Abundant** · San Francisco

Apply: https://jobs.ashbyhq.com/abundant/0d11f5c1-036f-4ef9-872f-f9407fb79d57

## About Abundant

Hello! 👋

We are a team of former ML engineers, founders, roboticists and ops leads who obsess about data and its impact on safe, reliable AI.

We specialize in creating environments and datasets for RL by leveraging our experience in simulation and model training.

By the numbers:
  • Powering 3 of the top 6 global AI labs and multiple Fortune 500 enterprises
  • Billions of training tokens generated each month, 2x month over month
  • Exclusive, global network of over 500 domain experts

We believe humans are inherently creative, and thrive by pushing the frontier.

We are working towards an abundant future--one where everyone has access to infinite intelligence, services and goods.

Based in San Francisco, CA. We enjoy good food and good company.

-- more info below --

Abundant is building the NVIDIA of training data. AI models rely on two fundamental ingredients: compute and data. NVIDIA, the leader in compute, has a peak market cap of $5T and generated $130B in revenue last year as the need for scaling compute has exploded. We believe the need to scale data is just beginning, as we move beyond SFT and human supervision to RL and Learning from Experience.

Our founding team consists of second-time founders, ML engineers and data leads from Waymo, Google, Meta and AWS. Our team has previously collaborated with DeepMind to classify hate speech in YouTube videos, trained SOTA models for self-driving, and scaled data pipelines with thousands of human annotators. Our pioneering work in human computation, synthetic data, imitation learning and RL give us a solid advantage in delivering results to our customers.

Why now? Training data is more important and more scarce than ever before. Scaling laws dictate that linear improvement in model performance demands an exponential increase in training data. But there is only one World Wide Web and most of it has already been trained on. The next advances will require new, diverse, and high-quality datasets, making training data more important and scarce than ever before.

What happens if we succeed? Abundant will be the core enabler for AGI and beyond. Most of the challenges in model training are already solved. What’s missing is the data necessary to move from general knowledge to domain expertise; from chatbots to agents; and from digital intelligence to physical AI. Ask any AI researcher or roboticist: the core bottleneck to progress is the availability of data, i.e. “abundant data”.

Abundant works with the most advanced AI labs and startups, as well as F500 enterprises.

## About the role

### ABOUT ABUNDANT

Abundant is an applied research lab focused on scaling reinforcement learning for safe and reliable agentic capabilities. We are an extremely talent-dense team of researchers, roboticists, founders, and operators whose work includes [two-tower retrieval](https://dl.acm.org/doi/10.1145/2959100.2959190), [BERT](https://arxiv.org/abs/1810.04805), [web-scale graph neural networks](https://arxiv.org/abs/1806.01973), and the [Waymo Driver](https://waymo.com/blog/2020/10/waymo-is-opening-its-fully-driverless).

### THE ROLE

Frontier labs test their models in environments like ours, which means we are the first to see novel model behavior. You will make sure that alignment is baked into everything the do: the code we deploy, the runbooks we write, and the papers we publish.

### WHAT YOU’LL DO

- Run experiments on how our environments shape behavior: reward hacking, cheating the task, unsafe use of tools, deception over long runs. Then change how we build them.

- Break our own graders. Find the ways to score well without doing the work, close them, and keep the attacks as tests everyone runs.

- Build the tools that watch agents: monitors over trajectories, red team harnesses, and safety benchmarks that stay hard as models improve. On real runs, not toy ones.

- Keep our environments sealed. Network isolation, escapes, credentials, and how much damage an agent can do when it turns on us. Last summer’s failures came from here.

- Choose a research question about oversight, control, or long-horizon autonomy, answer it, and publish. This is part of the job, not something you do at night.

- Decide what we will not build. Tell partners plainly what they get from us is and is not safe for.

### WHO YOU ARE

- You are an engineer first. You build your own test pipelines, learn strange codebases quickly, and find the bug in the log.

- You know LLM safety in depth: attacks, red teaming, monitoring and control evaluations, reward hacking.

- You can turn a vague worry about a model into an experiment, run it, and say what the result does and does not prove.

- You read transcripts closely. You catch the small thing that makes the whole result wrong.

- You have put research into production systems, especially post-training, distillation, or evaluation that something depends on.

- You write clearly enough that your results change what other people do.

- You can decide fast with incomplete evidence and defend the call with data.

### NICE TO HAVE

- Published work on control, dangerous capability evaluations, oversight, or interpretability.

- RLHF or RLAIF experience, and a view on how training choices show up in behavior.

- You built or ran a public benchmark or agent task suite.

- You have run many agents at once: sandboxes, containers, logging.

- You have advised on AI safety or governance policy.

### COMPENSATION

**Base Salary**

$250,000 - $450,000

**Cash Bonus**

Sizable performance bonus tied to project and company milestones

**Equity**

Generous early-stage grant

**Benefits**

Health, dental, vision + flexible PTO

---

Source: Abundant's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/abundant-member-of-technical-staff-alignment-59c3e1cf-db9f-457f-83f4-ff612611ed65
