Technical Intern
- Onsite
- FullTime
- Technical
About the role
About Us:
d_model is a fundamental AI research lab partnering with frontier labs to turn their models into capable interpretability and alignment researchers. Alongside our partnerships, we aim to use the agents we build for independent research.
Our team brings experience from places including OpenAI, Google, Anthropic, EleutherAI, and MATS.
About the Role:
As a Technical Intern, you'll work directly under a supervising Member of Technical Staff. You'll help build the reinforcement learning environments and evaluations we use to study how AI agents approach alignment problems. This role is a strong fit for someone early in their research career who wants hands-on experience in AI safety, interpretability, and reinforcement learning.
Primary Responsibilities:
Run experiments under guidance and document observed patterns in model behavior and failure modes.
Assist in exploring AI interpretability techniques within our reinforcement learning environments.
Support the development of reinforcement learning environments and evals that test how AI agents approach alignment problems, including implementing scoped components, writing test cases, and helping validate results.
Help test graders for robustness to specification gaming, including identifying edge cases and proposing improvements.
Participate in team ideation meetings and research discussions, sharing findings from your work and learning from ongoing research.
Collaborate with mentors and team members across research and engineering functions.
You May Be a Good Fit If You:
Have strong proficiency in Python and ML frameworks (PyTorch or JAX)
Are able to iterate quickly and collaboratively
Are familiar with the alignment and/or interpretability literature
Can generate research ideas in the field and implement others' ideas
Strong Candidates May Also Have:
Prior research experience, academic or independent
A completed MATS, SPAR, or similar AI safety research program
Personal projects or write-ups in interpretability or alignment
Published or presented research in AI safety, especially Mechanistic Interpretability
Experience post-training language models
What We Offer:
Flexible PTO policy
Catered breakfast and lunch
Monthly workspace and wellness allowances
Commuter benefits
Free SFMOMA admission and group subscriptions to Works in Progress, The New Yorker, Eye, and Aesthetica
Logistics:
Location Policy: Currently, we expect all staff to be in office 5 days a week in our San Francisco location.
Visa Sponsorship: We sponsor visas. Not every role and candidate combination works out, but when we make an offer we commit to a genuine effort, supported by an immigration counsel we work with directly.
Compensation: $15,000 per month
We encourage you to apply even if you don't believe you meet every single qualification.