---
title: "Product Analyst"
company: "Bolna AI"
company_url: "https://www.remjobs.works/companies/bolna-ai"
url: "https://www.remjobs.works/job/bolna-ai-product-analyst-96caa681-57e4-4943-9ab4-304cdfb3c7a3"
apply_url: "https://jobs.ashbyhq.com/bolna/a970416f-30f3-445e-8f22-e9d452c518cb"
workplace: onsite
location: "Bengaluru"
employment_type: full-time
seniority: mid
role: product
region: india
skills: ["azure", "llm", "python", "sql"]
date_posted: 2026-09-24T06:07:17.236Z
first_seen_by_remjobs: 2026-09-24T06:37:08.618Z
---

# Product Analyst

**Bolna AI** · Bengaluru

Apply: https://jobs.ashbyhq.com/bolna/a970416f-30f3-445e-8f22-e9d452c518cb

## About Bolna AI

Scale customer support, sales, recruitment with Bolna Voice Agents. Inbound and outbound calls in vernacular Indian languages like English, Hindi, Hinglish.

## About the role

### About Bolna

Bolna is a YC-backed voice AI orchestration platform built for the Indian market—powering multilingual, vernacular voice agents across Hindi, Hinglish, Tamil, and 10+ languages at sub-500ms latency across collections, recruitment, sales, and e-commerce use cases. We are an orchestration layer, not a model company: our moat is outcome-labelled vernacular data, rigorous evaluation infrastructure, and a growing taxonomy of how Indian enterprise voice AI fails in production.

#### Why This Role Exists

Product decisions at Bolna increasingly hinge on rigorous, code-mixed-aware data analysis—and not just one kind. On one side, there is model and evaluation rigor: LLM benchmarking for post-call intelligence, ASR/WER evaluation, inter-rater reliability on human-labelled calls, and routing and latency economics. On the other, there is product and growth insight: understanding where self-serve users drop off in their journey, what patterns emerge across lakhs of monthly calls, and which use cases and configurations are actually working.

Both currently sit with the Head of Product alongside strategy and roadmap ownership. We need a dedicated analyst to own the execution and recurring cadence across both-freeing product leadership to act on findings rather than produce them.

#### What You’ll Do

#### Model and Evaluation Analysis

- LLM and model benchmarking: Run structured comparisons across model providers such as Sarvam, DeepSeek, Gemini, and Claude variants for tasks including post-call extraction and LLM-as-judge scoring. Evaluate cost, accuracy, fill rate, and TTR, with particular attention to Hinglish and code-mixed content.

- Evaluation infrastructure: Build and maintain LLM-as-judge pipelines using tools such as DeepEval, design and track evaluation metrics, and run inter-rater reliability analysis such as Krippendorff’s alpha across human call reviewers.

- Golden dataset creation: Support the construction of golden datasets for ASR and transcript labelling, including flagging conventions such as code-switch scripting in Devanagari versus Roman script and transliteration normalization before scoring.

- ASR and voice benchmarking: Evaluate WER and related quality metrics across ASR providers and models for Indic languages, using public benchmarks and academic references where relevant.

- Infrastructure and latency analytics: Analyse routing, latency, and cost data, including Azure PTU utilization and percentile latency distributions, to inform infrastructure and routing decisions.

- Agent behaviour analytics: Support population-level analysis of graph-agent behaviour, including node-level aggregates, designed-versus-observed graph differences, stuck-in-loop detection, and similar failure-pattern metrics.

#### Product and Growth Insight Generation

- Self-serve journey analysis: Instrument and analyse the self-serve funnel from signup to activation, habit, and expansion; identify where users drop off and surface friction points for the product team.

- Cross-customer call insights: Mine aggregate call data across customers and use cases for patterns, including completion rates by use case, the best-performing model and configuration combinations, and emerging failure patterns across prompt templates.

- Ad hoc product analysis: Serve as a fast, reliable “pull me the data on X” resource for pod PMs-covering usage patterns, cohort behaviour, and feature adoption-and turn raw usage data into a clear, actionable read.

#### Across Both Areas

- Reporting and tooling: Build repeatable dashboards and scripts—not one-off notebooks—so these analyses run as an ongoing cadence. Present findings to product, ML, and infrastructure stakeholders in a form they can act on.

#### What We’re Looking For

#### Must-Have

- 0–2 years of experience as a new graduate or early-career professional in a data or product analyst, applied ML, or research-adjacent role. We are hiring for raw analytical strength and trainability, not a finished track record.

- Strong SQL and Python skills, including pandas, developed through coursework, internships, or prior work. The candidate should be comfortable writing and debugging their own queries and scripts without hand-holding.

- Solid statistical fundamentals, including distributions, basic hypothesis testing, and agreement or reliability concepts. Production experience with inter-rater reliability metrics is not required, but the candidate should be able to learn new statistical concepts quickly.

- Genuine comfort with ambiguous, messy real-world data, including the ability to notice when something looks wrong and flag it clearly even if the resolution is not theirs to make.

- Basic funnel and cohort analysis instincts, with comfort thinking in terms of drop-off stages and segments.

- Native or near-native fluency in Hindi-or another Indian language-and English, with comfort reading and labelling code-mixed or Hinglish text.

#### Strong Plus

- Exposure to LLM evaluation concepts such as prompt-based scoring and LLM-as-judge, speech or ASR evaluation such as WER and transcript QA, or evaluation frameworks such as DeepEval. Coursework and personal projects count.

- Exposure to product or growth analytics, including funnel analysis, retention curves, and cohort behaviour, through a prior role, internship, or personal project.

- Familiarity with cloud-inference economics, including token-based billing or provisioned throughput models.

- A portfolio of self-directed analysis—a project, competition, or write-up—that demonstrates an ability to look for the “so what,” not just the number.

#### First 90 Days – Success Looks Like

- Run the LLM benchmarking comparison for post-call extraction end-to-end under direction-executing the sweep and producing clean cost, accuracy, and fill-rate tables while escalating judgment calls rather than making them alone.

- Contribute meaningfully to the golden dataset build by clearly flagging inconsistencies in transliteration and script conventions, then applying the agreed decision consistently across the dataset.

- Run the inter-rater reliability pipeline on a recurring basis once it is set up, without needing to redesign it each time.

- Produce a clear first-pass view of the self-serve funnel-signup, activation, habit, and expansion-with at least one concrete drop-off point identified and flagged for action.

- Become the reliable first pass for “pull me the data on X” across at least two areas: model benchmarking, ASR evaluation, routing and latency, agent behaviour analytics, or self-serve and usage patterns.

#### What We Offer

- Innovative culture: Be part of a generational opportunity to define the trajectory of AI while working with a team pushing the boundaries of what is possible.

- Growth paths: Join a dynamic team with opportunities to drive impact beyond the immediate role and responsibilities.

- Learning and development: Bolna proactively supports professional development, with the relevant processes being established.

- Competitive compensation and meaningful ESOPs.

- In-person team collaboration from the Bengaluru office.

---

Source: Bolna AI's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/bolna-ai-product-analyst-96caa681-57e4-4943-9ab4-304cdfb3c7a3
