Manager, Machine Learning Engineering
- $170,000–$210,000 per year
- Remote
- All ai & machine learning jobs
- US
- Full-time
- Technology
About the role
The Role
We’re looking for a Manager, Machine Learning Engineering to lead Tala’s ML Platform team. This person will manage a team of Machine Learning Engineers responsible for building the platforms, frameworks, and infrastructure that enable our Data Science teams to securely train, deploy, monitor, and operate machine learning models at scale.
This is a player-coach management role. You’ll be responsible for developing and growing the team while also providing enough technical leadership to guide architecture, engineering practices, reliability, and production systems. The role has a particular focus on real-time machine learning inference and streaming data systems, as well as the platforms that support batch model development and deployment.
What You'll Do
- Manage and develop a team of 4–6 Machine Learning Engineers across mid-to-senior levels.
- Hire, source, interview, and close strong MLE talent.
- Establish clear expectations, provide regular feedback, and create development plans for direct reports.
- Coach engineers toward growth and promotion while addressing performance gaps directly and thoughtfully.
- Create opportunities for engineers to take on challenging projects and grow their technical leadership.
- Set quarterly goals and ensure the team consistently delivers against them.
- Own prioritization across product roadmap work, run-the-business activities, and operational excellence.
- Balance team capacity across new development, maintenance, technical debt, and production support.
- Improve team productivity by reducing context switching and delegating effectively.
- Partner with engineers and technical leads to estimate and scope complex work.
- Guide the development of platforms and frameworks that allow Data Scientists and Analysts to explore data, develop features, and train, test, deploy, and monitor ML models.
- Provide technical leadership across model infrastructure, real-time inference, streaming feature extraction, batch processing, and production ML systems.
- Drive strong engineering practices around testing, automation, observability, fault tolerance, infrastructure-as-code, and deployment.
- Own and improve SLOs, on-call health, capacity planning, reliability, and incident response.
- Review technical designs and help drive architectural standards and technical debt reduction.
- Work closely with Data Science, Data Engineering, Data Platform, Product, Credit, and Business Development teams.
- Translate business and technical needs into scalable ML platform solutions.
- Coordinate dependencies and delivery across multiple engineering and data teams.
- Help create structure and clarity in an environment where priorities and requirements can evolve.
Lead & Grow the Team
Own Engineering Delivery
Provide Technical Leadership
Partner Across the Organization
What You'll Need
- 2+ years of directly managing engineers, including hiring, performance management, coaching, and career development.
- Experience managing a team through at least one full performance cycle.
- Demonstrated ability to coach engineers toward promotion and address underperformance effectively.
- Experience owning team goals, prioritization, estimation, and delivery.
- Experience with production on-call, incident response, and capacity planning.
- Willingness to be actively involved in sourcing, interviewing, and closing engineering talent.
- 6+ years of backend software engineering experience in consumer-scale applications.
- At least 3 years of hands-on Python experience.
- Experience building and operating machine learning or causal inference systems in production.
- Earlier-career experience personally building and deploying ML models or ML infrastructure.
- Ability to participate in technical architecture and system-design discussions and provide technical direction without needing to be the primary coder.
- Strong understanding of software quality, security, reliability, testing, and production operations.
- Languages: Python, SQL
- Machine Learning: Jupyter, Pandas, Scikit-Learn, XGBoost, TensorFlow, PyTorch, Hugging Face
- Cloud & Infrastructure: AWS, GCP, Azure, Kubernetes, Docker
- Streaming: Kafka, Kinesis, Beam, Flink, Spark Streaming
- Batch Processing: Airflow, Metaflow
- Databases: MySQL, PostgreSQL, Cassandra, Snowflake, Druid, and/or similar technologies
- APIs: REST, GraphQL, gRPC, Protocol Buffers
- Production Engineering: DevOps, SLOs, monitoring/observability, on-call, capacity planning, root-cause analysis
- ML/Analytics: Machine learning, causal inference, scalable algorithms
Management Experience
Technical Experience
Technical Skills
We’re particularly interested in candidates with experience across:
Description as published by Tala.