Back to Research
Research2026-07-265 min read

Cost per Speaking Hour: A Transparent Comparison Model

VP

Vlad Podoliako

Founder & CEO, LinguaLive

Vlad Podoliako is the founder of LinguaLive, an AI-powered language learning platform focused on making useful speaking practice available on demand.

Follow on LinkedIn

Executive finding

Language programmes should compare cost per completed learner speaking hour, not price per seat, scheduled class hour, or app subscription. A class may last 60 minutes while each learner speaks for eight; an app licence may cover a month while the learner completes no voice practice. The denominator must be observed spoken participation, and the result must sit beside quality, safety, learning, and equity measures.

This brief introduces a calculation template. It is not an experimental study, price benchmark, or claim that automated practice is equivalent to instruction.

The core formula

Cost per completed learner speaking hour = total programme cost ÷ total observed learner speaking hours

For a cohort, calculate the denominator by summing each learner's verified voice minutes and dividing by 60. Do not multiply scheduled class hours by enrolment and call the result speaking time; that counts listening, teacher talk, silence, administration, absence, and unused capacity.

Numerator: total programme cost

Include the costs required to deliver the measured period:

  • teacher, facilitator, assessor, and administrator time;
  • employer taxes, benefits, and contractor overhead where applicable;
  • licences and usage charges;
  • onboarding, training, curriculum configuration, and content review;
  • integration, identity, analytics, accessibility, and support;
  • devices, rooms, connectivity, and scheduling where borne by the programme;
  • privacy, security, procurement, legal, and safeguarding review;
  • assessment and quality assurance;
  • learner support and remediation;
  • an allocated share of one-time implementation cost.

Report recurring delivery cost and allocated first-year cost separately. A mature rollout should not hide setup cost, and a pilot should not make every future year carry the full initial integration twice.

Denominator: completed learner speaking hours

Count audible or signed productive participation in the target language that fits the programme definition. Exclude:

  • absent or logged-in-but-inactive time;
  • listening-only time;
  • teacher talk;
  • silent preparation;
  • menus, loading, and technical failure;
  • microphone tests and accidental background audio;
  • recordings too incomplete to serve the task.

Privacy-preserving measurement can use on-device or aggregated speech-activity duration without retaining raw audio. Document the method and error rate; speech detection is not perfect and may undercount pauses that are part of legitimate turn planning.

Calculating different delivery modes

Teacher-led group

Use observed learner-speaking proportions rather than a fixed assumption:

completed speaking hours = sum of each attendee's target-language speaking minutes ÷ 60

If direct measurement is unavailable, sample representative sessions and report the speaking-share estimate with a range. Group size, pedagogy, level, and task design can change the result substantially.

One-to-one tutoring

Do not count the full session automatically. Subtract tutor talk, explanation, silent work, and non-target-language administration. These activities may be valuable; they simply are not learner speaking minutes.

AI voice practice

Use completed learner voice time, not “minutes in app.” Exclude system speech, processing delays, menus, and inactive sessions. Track the percentage of sessions that fail because of speech recognition, connectivity, safety stops, or model errors.

Language exchange

Include platform or facilitation cost and learner time if the institution values time economically. Split the session by target language. A 60-minute reciprocal call with 30 minutes in each language is not 60 learner-speaking minutes for either participant, because each also listens during their target-language half.

Worked illustration—not a market benchmark

Assume a 12-week programme with 100 learners. The invented figures below show the calculation only and must not be used as product pricing evidence.

Group programme

  • total delivery and allocated setup cost: £13,800;
  • 100 learners scheduled for 24 one-hour sessions;
  • attendance: 80%;
  • sampled learner-speaking share of attended time: 35%.

Calculation:

100 × 24 × 0.80 × 0.35 = 672 completed learner speaking hours

£13,800 ÷ 672 = £20.54 per completed learner speaking hour

Voice-practice programme

  • total licence, usage, support, review, and allocated setup cost: £6,400;
  • verified completed learner voice time: 54,000 minutes;
  • invalid/technical-failure voice time already excluded.

Calculation:

54,000 ÷ 60 = 900 completed learner speaking hours

£6,400 ÷ 900 = £7.11 per completed learner speaking hour

This does not prove equal value. The teacher programme may deliver diagnosis, social interaction, explanation, safeguarding, and accountable assessment that the voice tool does not. The next section prevents cost from swallowing quality.

Pair cost with an outcome and quality dashboard

Report cost per speaking hour beside:

Dimension Example measure
Reach Eligible learners who begin and complete
Speaking dose Median completed voice minutes; distribution, not only mean
Task performance Success on an unpractised parallel speaking task
Intelligibility Blind human-listener understanding
Interaction Response to genuine follow-ups and repair
Retention Delayed outcome after access ends
Equity Completion and scoring gaps across groups/devices
Reliability Failed-session and unusable-feedback rate
Safety Reported incidents, privacy exceptions, escalations
Experience Learner and teacher usefulness ratings

A lower cost with no improvement may be waste. A higher cost may be justified by outcomes or services the cheaper mode does not provide.

Sensitivity analysis

Publish at least three cases:

  • low use / conservative speaking share;
  • observed base case;
  • high but credible use / speaking share.

Also change completion, staff time, implementation allocation, and support cost. If the preferred conclusion disappears under a small assumption change, the business case is fragile.

Procurement questions

  1. Can the provider export verified learner voice minutes without raw audio?
  2. What counts as an active or completed voice minute?
  3. Are system speech and waiting time excluded?
  4. How are failed and abandoned sessions reported?
  5. What teacher, QA, privacy, and support work shifts to the institution?
  6. Are usage overages, minimum commitments, and renewal increases modelled?
  7. Which learning and equity outcomes accompany the cost metric?
  8. What happens to the economics at realistic—not contractual maximum—use?

Editorial disclosure

LinguaLive sells AI speaking practice and therefore benefits if speaking-hour economics are valued. Before publication, finance and education researchers should validate the formula, cost categories, measurement protocol, and worked arithmetic. Any public comparison must replace illustrative values with dated, auditable local evidence and preserve the non-equivalence caveat.

Related Topics

cost per speaking hour comparison modelcost per speaking houroriginal methodology brieflanguage learning evidence

Share this article

Ready to Start Learning?

Try LinguaLive's AI-powered conversation practice free. 10 minutes a day can transform your fluency.

Start Free - 10 Min Daily

Your first sentence is one tap away.

Hear the tutor, take the mic for three minutes, then keep 10 free minutes a day with an account.

Replay the demo

No credit card required · 10 min/day free · Cancel anytime