---
title: "Machine Learning Engineer - ML Training Platform"
company: "Pluralis Research"
company_url: "https://www.remjobs.works/companies/pluralis-research"
url: "https://www.remjobs.works/job/pluralis-research-machine-learning-engineer-ml-training-platform-1646b903-ad27-4414-abfb-1af36647de23"
apply_url: "https://jobs.ashbyhq.com/pluralis-research/7b107585-5a3b-428d-ad5d-1c85be9dca4e"
workplace: remote
location: "San Francisco"
remote_scope: "San Francisco"
employment_type: full-time
seniority: mid
role: ai-machine-learning
region: united-states
skills: ["aws", "azure", "docker", "gcp", "kubernetes", "python", "terraform"]
date_posted: 2026-08-31T22:13:24.346Z
first_seen_by_remjobs: 2026-09-17T11:39:12.860Z
---

# Machine Learning Engineer - ML Training Platform

**Pluralis Research** · San Francisco

Apply: https://jobs.ashbyhq.com/pluralis-research/7b107585-5a3b-428d-ad5d-1c85be9dca4e

## About the role

Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights ([tech report](https://arxiv.org/abs/2607.13332)). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read [A Third Path: Protocol Learning](https://pluralis.ai/blog/a-third-path-protocol-learning/).

Our training and inference doesn't happen in a datacenter. It happens on consumer nodes and cloud instances that are not co-located, connected by ordinary internet, joining and leaving mid-run. Your primary role is to architect, build, and scale the platform that keeps continuous experimentation and large-scale training running on top of that: infrastructure orchestration, distributed compute, and the services that tie them together.

#### Key Responsibilities

- Multi-cloud infrastructure: Design the resource management systems that provision and orchestrate compute across AWS, GCP, and Azure with infrastructure-as-code (Pulumi/Terraform). Handle dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes.

- Distributed training and inference systems: Architect fault-tolerant infrastructure for distributed ML. GPU clusters, NVIDIA runtime, S3 checkpointing, large-dataset management and streaming, health monitoring, and resilient retry strategies.

- Real-world networking: Build the systems that simulate and handle real network conditions such as bandwidth shaping, latency injection, packet loss. Managing node churn and keeping data flowing across workers with heterogeneous connectivity.

#### What We're Looking For

- Infrastructure and platform engineering (required): Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments, Docker/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at scale.

- Distributed systems and ML infrastructure: You understand distributed training workflows: checkpointing, data sharding, model versioning, long-running job orchestration.

- Decentralized networking: P2P, NAT traversal, traffic shaping, real bandwidth constraints.

- Systems programming and reliability: Strong Python engineering (asyncio, concurrency, retry logic, cloud SDKs, CLI tooling) with hands-on observability and SRE practice; Prometheus/Grafana, performance profiling, incident response.

- Environment fit: You've done this in a startup with heavy service orchestration, or at big-tech scale, and you can show which systems you owned.

- Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI.

#### Nice to Have

- Experience with foundation model pre-training, post-training, or RL.

- Experience at proprietary, open-weight and open-source AI labs

#### Compensation & Benefits

- Equity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary.

- Remote-First Culture: Flexible work environment with team members distributed globally.

- Visa Sponsorship: Optional full visa sponsorship and relocation support to either Australia or the US.

- Open Problems: Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones.

#### FYI's

- We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones.

- Applicants must have professional-level English proficiency (written and spoken).

- Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help.*We are backed by *[Union Square Ventures](https://www.usv.com/)* and other tier-1 investors, and we are a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We believe AI, and the world, end up on a better path if we succeed in implementing the protocol for intelligence. If this resonates, please apply.*

---

Source: Pluralis Research's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/pluralis-research-machine-learning-engineer-ml-training-platform-1646b903-ad27-4414-abfb-1af36647de23
