---
title: "Research Engineer, Benchmarks"
company: "Clera"
company_url: "https://www.remjobs.works/companies/clera"
url: "https://www.remjobs.works/job/clera-research-engineer-benchmarks-9d81cce6-3cdf-41aa-af0b-d834d39294f9"
apply_url: "https://jobs.ashbyhq.com/clera/36916df6-1948-4ce1-b336-2b9b1027fda6"
workplace: onsite
location: "Singapore"
employment_type: full-time
seniority: mid
role: ai-machine-learning
region: asia-pacific
skills: ["docker", "linux", "llm", "python"]
date_posted: 2026-09-11T17:38:25.161Z
first_seen_by_remjobs: 2026-09-15T16:29:34.147Z
---

# Research Engineer, Benchmarks

**Clera** · Singapore

Apply: https://jobs.ashbyhq.com/clera/36916df6-1948-4ce1-b336-2b9b1027fda6

## About Clera

Stop applying to startups. Start getting introduced. Clera connects you directly with hiring managers at the companies you want to work for.

## About the role

##### About the Role

This role sits at the heart of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to ensuring benchmarks are rigorous, credible, and practically meaningful.

##### What You'll Do

- Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.

- Partner with subject-matter experts to define realistic workflows and translate them into evaluation criteria.

- Build reliable infrastructure to run models and agents against benchmark tasks at scale using Python, Docker, and Linux environments.

- Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes.

- Validate that benchmark performance correlates with real-world evaluations and customer needs.

- Write clear technical documentation and benchmark reports for research and engineering audiences.

##### What We're Looking For

- 2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.

- Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.

- Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models.

- Experience developing metrics, statistical analyses, or validation studies to assess benchmark quality and real-world correlation.

- Experience collaborating with domain experts to translate workflows into structured evaluation tasks.

- Strong technical writing skills, with published papers or technical posts on AI benchmarking, model evaluation, or failure modes being a plus.

- Ability to reason from first principles about task design, scoring, and edge cases.

- Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.

- Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a bonus.

##### Compensation & Benefits

Salary range: **$150,000 to $250,000 USD annually.** Visa sponsorship is available.

##### Location

On-site in **Singapore**. This is a full-time, in-person role.

---

Source: Clera's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/clera-research-engineer-benchmarks-9d81cce6-3cdf-41aa-af0b-d834d39294f9
