---
title: "Site Reliability Engineer"
company: "Latent"
company_url: "https://www.remjobs.works/companies/latent"
url: "https://www.remjobs.works/job/latent-site-reliability-engineer-48429dae-f45e-4131-82de-bed7a68a8350"
apply_url: "https://jobs.ashbyhq.com/latent/bbd1a8e3-b943-4c8c-b239-e87951504f71"
workplace: onsite
location: "San Francisco"
employment_type: full-time
seniority: mid
role: devops-infrastructure
region: united-states
skills: ["kubernetes", "postgres", "python", "redis", "terraform", "typescript"]
date_posted: 2025-12-05T23:21:20.724Z
first_seen_by_remjobs: 2026-08-27T09:28:13.835Z
---

# Site Reliability Engineer

**Latent** — San Francisco

Apply: https://jobs.ashbyhq.com/latent/bbd1a8e3-b943-4c8c-b239-e87951504f71

## About Latent

At Latent, we're building medical language models to tackle the trillions of operational overhead weighing down healthcare. Our flagship product streamlines authorizations for life-saving drugs by analyzing EHR records and surfacing the most relevant data.

We've built our product and signed enterprise contracts with some of the largest health systems in the US.

## About the role

### SRE

**Location:** San Francisco, CA (5 Days In-Office)

You are the infrastructure expert who enables our rapid product development and guarantees **99.9%+ stability and performance** of our clinical AI platform for major health systems. Your focus on operational excellence is directly tied to a patient's access to life-saving treatment.

#### What We Look for in a Great Engineer

You have the intensity and technical mastery to own mission-critical infrastructure. You hold yourself and others to high standards and thrive in a high-energy, in-office culture where everyone is in it to win it.

- Tool Proficiency: You are highly proficient with your tools—you speak command line fluently and have mastered keyboard shortcuts.

- Ownership: You thrive on owning complex systems and have a proven track record of scaling mission-critical deployments.

- Automation Drive: You love automating things, always finding new ways to increase your own leverage, and defining standards for operational excellence.

- Problem Solver: You won't wait for someone else to solve a problem that you're in a position to solve; you are willing to jump into whatever needs to get done.

#### What You'll Work On (Responsibilities)

As our SRE, you will own the entire production environment and improve the development experience:

- Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.

- Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.

- CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.

- DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.

- Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.

#### Technical Qualifications & Environment

- IaC & Orchestration: Deep, demonstrable experience with Kubernetes, Helm, and Terraform.

- Scaling Systems: Proven ability to architect and maintain complex, distributed systems with high-availability requirements.

- Deployment Experience: Hands-on experience optimizing deployment pipelines for both application code (TypeScript) and machine learning models (Python/ML). Also PostgreSQL, Redis, Kakfa.

- Core Team Member: Excitement about working five days per week in our San Francisco office.

---

Source: Latent's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/latent-site-reliability-engineer-48429dae-f45e-4131-82de-bed7a68a8350
