Research Engineer / Research Scientist, Confidential Computing
- $90,000–$150,000 per year
- Remote
- All ai & machine learning jobs
- Flexible
- FullTime
- Verification
About the role
The team
Some of the most consequential decisions of the AI era will depend on answering one question: can we verify what is happening inside AI datacenters? Answering it could unlock international cooperation, help countries protect their sovereignty, and enable trustworthy adoption of AI in high-stakes industries. Building tools that can earn trust across borders is an urgent technical and political challenge.
SASH’s Verification team builds and tests tools to verify agreements about AI. We prototype new verification mechanisms, lead international research collaborations, and work with policymakers around the world to show what these tools can do and inform how they’re used.
You’d join a small team combining technical expertise with international AI policy experience. Our technical team is led by Pascal Berrang, an Associate Professor at University of Birmingham, and brings experience from the Singapore Government. Our policy team brings experience from Oxford and the Centre for the Governance of AI, while our partners include experts from the Future of Life Institute and the University of Oxford.
The problem
Verifying what happens inside a datacenter running large AI models is the bottleneck on almost every serious agreement about AI — between states, between a regulator and a lab, between a company and the customers it wants to reassure. The obstacle is always the same shape: the party with something to prove cannot hand over its model weights, and the party doing the checking cannot simply take its word.
One solution to this problem puts its trust in hardware: a trusted execution environment isolates a computation and attests to what ran inside it, so a claim about what happened becomes a claim you can check against a chip. Confidential computing exists on current accelerators and runs near production speed.
That unlocks several problems. An audit tells you something about a model, but nothing connects that result to the weights actually being served — attesting the deployed model by weight hash closes the gap between "this model was evaluated" and "you are talking to the model that was evaluated." The most valuable evaluations are built by people who cannot publish their test set, against models whose weights cannot leave the developer; inside an enclave, neither side has to give way. Logs can be attested where they are produced, making usage auditable after the fact without being readable.
Your work
You would own the question of what a chip's word is actually worth, and then build the systems that are worth building on top of it.
Concretely, that means establishing what attestation on current accelerators does and does not prove, designing a registry that lets an outsider check an attested claim without anyone's permission, and making attested pipelines that hold together when someone hostile pulls on them. "It works" and "an adversary cannot make it lie" are different bars, and only the second one counts here. You'd work alongside cryptographers and hardware-security colleagues on the team and with external partners, and you'd be expected to publish since the whole approach depends on other people being able to check it.
Representative projects
Establish what a GPU confidential-computing attestation guarantees on current hardware, and publish where it breaks.
Design and build the registry: the attested-result format, the submission path for third-party auditors, and a verification client someone outside the trust chain can actually run.
Build attested logging inside an enclave, and work out what a privacy-preserving audit over those logs can and cannot establish.
Run an evaluation end to end under mutual secrecy, with an attestation recording whose code and whose weights were involved.
Chain attestations across a pipeline so a claim about a final output stays checkable back to its inputs, then find where that chain breaks under production load.
Get a confidential-computing stack running on accelerators from more than one supply chain.
Combine hardware attestation with the team's cryptographic work, so a verifier does not have to trust the chip alone.
Red-team our own attestation chain.
About you
We are hiring this role at a range of levels. We care more about ability and trajectory than years of experience.
You may be a good fit if you:
Have built something real with confidential computing or hardware enclaves (SGX, TDX, SEV-SNP, Arm CCA, GPU confidential computing) or have strong systems security fundamentals and want to learn them on the job.
Understand attestation chains well enough to know where they carry weight and where they are decoration.
Are hard to reassure: you read a security claim and look first for what it does not say.
Can ship an integration across organisational boundaries, which is as much about people as about code.
Thrive in ambiguous, early-stage environments where defining the problem is part of the job.
Strong candidates may also have:
ML inference serving experience and the operational scars of production deployment.
Background in audit, certification, or PKI: the registry is a certificate authority with unusual requirements.
Privacy-preserving computation experience beyond enclaves: secure computation, differential privacy, federated setups.
Familiarity with the AI evaluations landscape and what safety institutes actually need.
Experience explaining a hardware guarantee to someone who will never read a datasheet.
Description as published by Singapore AI Safety Hub.