AI Research Lead
- Hybrid
- All ai & machine learning jobs
- FullTime
- Engineering
About the role
About PointFive
PointFive is the AI Efficiency OS. From the cloud to the coding agent, we're the only platform that manages AI spend everywhere it happens.
Engineering and FinOps teams use PointFive to make their organizations more efficient and their cloud and AI more effective. We don't just show what you spend. We show what you're wasting, and we fix it autonomously. NuBank saw ROI in 10 days. Customers average 1,200%+ ROI and a 4.9 rating on G2.
Founded by the team behind IntSights (acquired by Rapid7), PointFive recently closed a $60M Series B led by Accel, with participation from Entrée Capital and Salesforce Ventures.
About the Role
PointFive is building infrastructure for the next generation of AI-powered engineering.
As AI agents become embedded into developer workflows, the underlying model layer is changing rapidly. Organizations are no longer choosing between a handful of hosted APIs. They are increasingly operating across frontier models, open-weight models, locally deployed models, specialized models, and dynamically routed combinations of them.
We’re looking for an AI researcher to take an integral part in PointFive’s research into how these models behave, how they should be evaluated, and how they can be deployed and used more efficiently across real engineering workloads.
This is a deeply technical research role with direct product impact.
You’ll study frontier and open-source models, develop evaluation frameworks, investigate inference and model optimization techniques, explore local deployment strategies, and help answer questions such as:
Which model should handle a particular task?
How much capability do we lose when we quantize it?
Can a smaller local model replace a frontier API for a specific workload?
How should models be routed across latency, cost, privacy, and quality constraints?
How do we measure whether one model is actually better than another for agentic software engineering?
Your work will directly shape PointFive’s model strategy and the intelligence behind our AI infrastructure products.
What You'll Do
Define and lead PointFive’s LLM research agenda across model evaluation, inference optimization, local models, routing, and model adaptation.
Continuously evaluate frontier and open-weight models across real-world software engineering and agentic workloads.
Design rigorous evaluation frameworks for model quality, reasoning, tool use, code generation, command execution, summarization, and agent behavior.
Build benchmarks that reflect actual developer workflows rather than generic academic tasks.
Research model efficiency techniques including quantization, distillation, speculative decoding, prefix caching, KV-cache optimization, batching, and context management.
Investigate when smaller or locally deployed models can replace expensive frontier models without materially degrading task quality.
Evaluate model architectures, parameter sizes, quantization formats, and runtime configurations across heterogeneous hardware.
Research and benchmark local inference stacks including llama.cpp, MLX, vLLM, SGLang, ONNX Runtime, Ollama, TensorRT-LLM, and emerging runtimes.
Study inference performance across Apple Silicon, NVIDIA GPUs, AMD GPUs, CPUs, Windows workstations, Linux machines, and other endpoint configurations.
Develop model-routing strategies that optimize across quality, latency, cost, privacy, context size, and hardware availability.
Explore intelligent cascades where smaller models handle common tasks and more capable models are invoked only when necessary.
Research model specialization through fine-tuning, LoRA, adapters, distillation, prompt optimization, and other model adaptation techniques.
Investigate model behavior under constrained environments, including offline execution, limited memory, limited compute, and local-only inference.
Evaluate agent-specific model behavior, including planning, tool selection, shell interaction, code editing, error recovery, and long-running task execution.
Analyze failure modes such as hallucination, tool misuse, context degradation, reasoning collapse, excessive token consumption, and unstable agent loops.
Design experiments that quantify the tradeoffs between model capability, inference cost, token consumption, latency, and resource utilization.
Build internal research infrastructure for reproducible model experiments, benchmarking, dataset management, and evaluation.
Track frontier model releases and emerging research, rapidly determining which developments are meaningful for PointFive’s products.
Collaborate closely with engineering and product teams to translate research results into production capabilities.
Build and lead a small, exceptional LLM research team over time.
What We're Looking For
Must-have
Deep understanding of modern large language models and transformer-based architectures.
Strong hands-on experience evaluating and experimenting with both frontier and open-weight models.
Strong understanding of inference behavior, including prefill, decoding, KV caches, context windows, batching, memory usage, and token generation performance.
Experience with model optimization techniques such as quantization, distillation, LoRA, fine-tuning, or model compression.
Strong experimental mindset — able to formulate hypotheses, design controlled experiments, and draw meaningful conclusions from noisy results.
Experience building evaluation frameworks for LLM quality and behavior.
Strong Python proficiency and familiarity with the modern ML ecosystem.
Ability to read, understand, and reproduce ideas from current ML research papers.
Comfortable working with ambiguous research problems where there may not yet be an established best practice.
Strong ability to bridge research and production — understanding not only whether something works, but whether it is practical to deploy.
Nice to have
Experience with open-weight models such as Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, or similar model families.
Experience with frontier model APIs including OpenAI, Anthropic, Google, and other leading providers.
Experience with inference frameworks such as vLLM, SGLang, llama.cpp, MLX, TensorRT-LLM, Ollama, or ONNX Runtime.
Deep understanding of quantization techniques including FP8, INT8, INT4, AWQ, GPTQ, GGUF, and related approaches.
Experience with GPU performance optimization, CUDA, Metal, ROCm, or DirectML.
Familiarity with distributed inference and multi-GPU serving.
Experience with model routing, mixture-of-model systems, cascades, or adaptive inference.
Experience with reinforcement learning, preference optimization, DPO, GRPO, or related post-training techniques.
Familiarity with agentic systems, coding agents, tool-use models, and computer-use models.
Experience building or evaluating coding benchmarks and software-engineering agents.
Experience with synthetic data generation, dataset curation, and automatic evaluation.
Research publications or meaningful contributions to open-source ML projects.
Experience leading a small applied research or ML research team.
Research Areas
Some of the problems we expect this team to work on include:
Model Routing
Choosing the optimal model dynamically based on task difficulty, latency requirements, cost, privacy constraints, and hardware availability.Local vs. Cloud Inference
Determining which workloads can reliably move from cloud models to models running directly on developer endpoints.Model Compression
Understanding how far models can be quantized, distilled, or otherwise optimized before meaningful capability is lost.Agentic Model Evaluation
Building evaluation methods for agents that operate over codebases, shells, developer tools, and long-running workflows.Inference Efficiency
Improving throughput, latency, memory consumption, and token efficiency across different model architectures and runtimes.Model Specialization
Investigating whether smaller specialized models can outperform general-purpose frontier models for narrow engineering tasks.Long-Context Behavior
Understanding how models behave as context grows, what information gets lost, and how context can be compressed or structured more intelligently.Hardware-Aware AI
Matching models and inference strategies to available GPUs, CPUs, NPUs, Apple Silicon, and other endpoint hardware.Cost vs. Intelligence Tradeoffs
Quantifying when additional model capability actually produces better outcomes — and when it simply produces more expensive tokens.
Our Tech Stack
Python, Go, PyTorch, Hugging Face, vLLM, SGLang, llama.cpp, MLX, CUDA, Metal, AWS, Cloudflare, Snowflake, and a rapidly evolving ecosystem of frontier and open-weight models.
Why PointFive
A rare research surface
We operate where LLM research meets real enterprise infrastructure. The questions we work on have immediate implications for how thousands of engineers use AI every day.
Access to real workloads
Instead of optimizing against abstract benchmarks, you'll be able to study how models perform across real software engineering and agentic workflows.
Research that ships
This isn't a research lab disconnected from product. Successful ideas can move rapidly from experiment to production.
Models are becoming infrastructure
Enterprises will increasingly operate fleets of models across cloud APIs, private infrastructure, and developer endpoints. Deciding how those models are selected, optimized, deployed, and governed is becoming a fundamental infrastructure problem.
Frontier moves fast
New models, architectures, inference techniques, and agent systems appear constantly. Your job is to understand which developments matter — and turn them into an advantage for PointFive.
Founders with a track record
Built and sold IntSights to Rapid7. Backed by top-tier investors with deep conviction in the category.
Early-stage leverage
You'll define PointFive's LLM research strategy, research methodology, and eventually the team itself.
Equal Opportunity Statement
PointFive is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We welcome candidates from all backgrounds, experiences, and perspectives to apply.
Description as published by PointFive.