Merlin LabsBoston
Senior Manager, AI Foundation Model
- Hybrid
- All ai & machine learning jobs
- Full Time
- Engineering
About the role
About Merlin:
Merlin (NASDAQ: MRLN) is a publicly traded aerospace and defense company building a non-human pilot to deliver full-stack autonomy for any aircraft from takeoff to touchdown. The Merlin Pilot autonomy system powers a growing range of aircraft and mission profiles and has been proven through hundreds of autonomous flights from Merlin's global flight test facilities, including Kerikeri, New Zealand; Quonset Point, Rhode Island; and soon, Bedford, Massachusetts. Headquartered in Boston, Merlin is expanding its organization to accelerate the development and deployment of its autonomy platform, helping customers solve some of aviation's most pressing challenges, from pilot shortages to improving flight safety. Backed by some of the world's leading investors prior to its public listing, Merlin continues to advance the certification and commercialization of autonomous flight across commercial and defense aviation.
About You: You have built learned decision-making systems that left the lab and ran on real hardware with real consequences. You are fluent in modern model architecture and post-training, but you are not a benchmark chaser — you have been in the room when a learned system was asked to justify itself to people who sign off on safety, and you know the difference between a model that performs well and a model whose behavior you can characterize. You want to work on a problem where “it works most of the time” is not a result.
About You: You have built learned decision-making systems that left the lab and ran on real hardware with real consequences. You are fluent in modern model architecture and post-training, but you are not a benchmark chaser — you have been in the room when a learned system was asked to justify itself to people who sign off on safety, and you know the difference between a model that performs well and a model whose behavior you can characterize. You want to work on a problem where “it works most of the time” is not a result.
Responsibilities:
- Technical strategy: own Merlin's foundation and world-model work — architecture selection, build-vs-adapt decisions, post-training approach, and the capability roadmap that supports it.
- Team leadership: lead and mentor a small team of world-model and post-training engineers; set the technical bar and the review culture for model work across AI Core.
- System interface: design the model interface to the rest of the autonomy stack — structured, schema-constrained plan outputs that a deterministic verifier can accept or reject, never free-form actuator authority.
- Evaluation: define what “good” means before training begins — build the evaluation harness, capability taxonomy, and regression suite that gate every model release, in partnership with the Data/Sim/Release pillar.
- Safety-relevant outputs: establish uncertainty quantification and out-of-distribution detection as first-class model outputs, not afterthoughts — downstream safety monitoring depends on them.
- Benchmarking: deliver an honest, reproducible comparison between learned planning and Merlin's current rule-based behavior planning across representative mission profiles, including the cases where the learned approach loses.
- Certification partnership: work with Systems Engineering, Certification, and the Chief Architect to keep model design inside what is defensible to a regulator, and to shape what “defensible” will mean for learned components.
- Research judgment: track the external research frontier and make disciplined calls about what Merlin adopts, builds, or ignores.
Qualifications:
- Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or a related subject.
- 8+ years building AI systems, with 3+ years leading technical teams or owning a major model program.
- Proven team management experience shipping high-tech, AI-powered models into production — hiring and developing AI engineers, setting technical direction and priorities, and owning delivery from research through deployment.
- Demonstrated ownership of a learned system that shipped into a physical, real-time product — robotics, autonomous vehicles, aerospace, or industrial autonomy.
- Depth in at least two of: world models and learned dynamics; sequence models applied to planning or control; post-training (SFT, preference optimization, RL fine-tuning); structured or constrained generation.
- Rigorous evaluation practice: you have built eval harnesses that caught regressions before customers did, and you can explain why a model's aggregate metric improved while a specific behavior got worse.
- Strong PyTorch; comfortable reading and reasoning about the C++ real-time systems your models feed.
- You write clearly. Architecture decisions here get read by systems engineers, safety engineers, and regulators — not only by other AI engineers.
Nice to Have:
- Experience with learned components in a certified or regulated product (DO-178C, ISO 26262, IEC 62304).
- Background in classical planning, behavior trees, MCTS, or hierarchical task networks — you'll be replacing and interoperating with exactly these.
- Familiarity with aviation domain structure: flight phases, ARINC 424 procedures, ATC phraseology.
- Publications or open-source contributions in embodied AI, world models, or robot learning.
Description as published by Merlin Labs.