---
title: "Research Engineer - Pre-training"
company: "Pluralis Research"
company_url: "https://www.remjobs.works/companies/pluralis-research"
url: "https://www.remjobs.works/job/pluralis-research-research-engineer-pre-training-bef4e6d8-55a1-49f9-b2f1-a0ee8e76cd62"
apply_url: "https://jobs.ashbyhq.com/pluralis-research/da969db9-7c3a-494e-ac43-fdb7b4320070"
workplace: remote
location: "San Francisco"
remote_scope: "San Francisco"
employment_type: full-time
seniority: mid
role: ai-machine-learning
region: united-states
skills: ["llm", "python", "pytorch"]
date_posted: 2026-08-31T13:59:41.075Z
first_seen_by_remjobs: 2026-09-17T11:39:12.860Z
---

# Research Engineer - Pre-training

**Pluralis Research** · San Francisco

Apply: https://jobs.ashbyhq.com/pluralis-research/da969db9-7c3a-494e-ac43-fdb7b4320070

## About the role

Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights ([tech report](https://arxiv.org/abs/2607.13332)). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read [A Third Path: Protocol Learning](https://pluralis.ai/blog/a-third-path-protocol-learning/).

This setting breaks nearly every assumption of datacenter training: communication-efficient training across different parallelism axes, fault tolerance as nodes join and drop mid-run, heterogeneous compute and networks, and robustness to malicious participants. Our published methods include [Subspace Networks](https://arxiv.org/abs/2506.01260), [Factored Gossip DiLoCo](https://arxiv.org/abs/2606.22768), [AsyncMesh](https://arxiv.org/abs/2601.22442), and [Sentinel](https://arxiv.org/abs/2603.03592).

As a Research Engineer you'll build the training system that takes Protocol Learning from the 8B run to frontier scale: large models on heterogeneous hardware, in physically different regions, connected by ordinary internet.

#### Key Responsibilities

- Distributed pretraining: Implement and optimize model-parallel training. Data, pipeline, and tensor parallelism for large models on heterogeneous GPUs under low-bandwidth, high-latency links.

- Performance optimization: Implement techniques that reduce communication overhead while maintaining model convergence in challenging network environments.

- Elasticity and fault tolerance: Make runs survive node churn. Robust checkpointing, state synchronization, and recovery as participants join and leave.

- Run instrumentation: Build the monitoring that shows throughput, bottlenecks, and model quality across hundreds of devices.

#### What We're Looking For

- Hands-on distributed training (required): You've trained models across many devices in PyTorch with FSDP, DeepSpeed, Megatron, or your own implementation. You understand data, tensor, and pipeline parallelism.

- Strong engineering: Production-quality Python. Concurrency, failure handling, profiling before optimizing.

- Evidence of execution: Shipped systems, research code, open-source work, or serious personal projects.

- Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI.

#### Nice to Have

- Hands-on experience training or serving large language models such as Nemotron, Qwen or OLMo.

- Experience with P2P networking and NAT traversal.

- Experience with post-training and RL.

- Experience with inference and serving systems.

- Experience at proprietary, open-weight and open-source AI labs

#### Compensation & Benefits

- Equity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary.

- Remote-First Culture: Flexible work environment with team members distributed globally.

- Visa Sponsorship: Optional full visa sponsorship and relocation support to either Australia or the US.

- Open Problems: Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones.

#### FYI's

- We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones.

- Applicants must have professional-level English proficiency (written and spoken).

- Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help.*We are backed by *[Union Square Ventures](https://www.usv.com/)* and other tier-1 investors, and we are a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We believe AI, and the world, end up on a better path if we succeed in implementing the protocol for intelligence. If this resonates, please apply.*

---

Source: Pluralis Research's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/pluralis-research-research-engineer-pre-training-bef4e6d8-55a1-49f9-b2f1-a0ee8e76cd62
