Cost per Speaking Hour: A Transparent Comparison Model
Vlad Podoliako
Founder & CEO, LinguaLive
Vlad Podoliako is the founder of LinguaLive, an AI-powered language learning platform focused on making useful speaking practice available on demand.
Follow on LinkedInExecutive finding
Language programmes should compare cost per completed learner speaking hour, not price per seat, scheduled class hour, or app subscription. A class may last 60 minutes while each learner speaks for eight; an app licence may cover a month while the learner completes no voice practice. The denominator must be observed spoken participation, and the result must sit beside quality, safety, learning, and equity measures.
This brief introduces a calculation template. It is not an experimental study, price benchmark, or claim that automated practice is equivalent to instruction.
The core formula
Cost per completed learner speaking hour = total programme cost ÷ total observed learner speaking hours
For a cohort, calculate the denominator by summing each learner's verified voice minutes and dividing by 60. Do not multiply scheduled class hours by enrolment and call the result speaking time; that counts listening, teacher talk, silence, administration, absence, and unused capacity.
Numerator: total programme cost
Include the costs required to deliver the measured period:
- teacher, facilitator, assessor, and administrator time;
- employer taxes, benefits, and contractor overhead where applicable;
- licences and usage charges;
- onboarding, training, curriculum configuration, and content review;
- integration, identity, analytics, accessibility, and support;
- devices, rooms, connectivity, and scheduling where borne by the programme;
- privacy, security, procurement, legal, and safeguarding review;
- assessment and quality assurance;
- learner support and remediation;
- an allocated share of one-time implementation cost.
Report recurring delivery cost and allocated first-year cost separately. A mature rollout should not hide setup cost, and a pilot should not make every future year carry the full initial integration twice.
Denominator: completed learner speaking hours
Count audible or signed productive participation in the target language that fits the programme definition. Exclude:
- absent or logged-in-but-inactive time;
- listening-only time;
- teacher talk;
- silent preparation;
- menus, loading, and technical failure;
- microphone tests and accidental background audio;
- recordings too incomplete to serve the task.
Privacy-preserving measurement can use on-device or aggregated speech-activity duration without retaining raw audio. Document the method and error rate; speech detection is not perfect and may undercount pauses that are part of legitimate turn planning.
Calculating different delivery modes
Teacher-led group
Use observed learner-speaking proportions rather than a fixed assumption:
completed speaking hours = sum of each attendee's target-language speaking minutes ÷ 60
If direct measurement is unavailable, sample representative sessions and report the speaking-share estimate with a range. Group size, pedagogy, level, and task design can change the result substantially.
One-to-one tutoring
Do not count the full session automatically. Subtract tutor talk, explanation, silent work, and non-target-language administration. These activities may be valuable; they simply are not learner speaking minutes.
AI voice practice
Use completed learner voice time, not “minutes in app.” Exclude system speech, processing delays, menus, and inactive sessions. Track the percentage of sessions that fail because of speech recognition, connectivity, safety stops, or model errors.
Language exchange
Include platform or facilitation cost and learner time if the institution values time economically. Split the session by target language. A 60-minute reciprocal call with 30 minutes in each language is not 60 learner-speaking minutes for either participant, because each also listens during their target-language half.
Worked illustration—not a market benchmark
Assume a 12-week programme with 100 learners. The invented figures below show the calculation only and must not be used as product pricing evidence.
Group programme
- total delivery and allocated setup cost: £13,800;
- 100 learners scheduled for 24 one-hour sessions;
- attendance: 80%;
- sampled learner-speaking share of attended time: 35%.
Calculation:
100 × 24 × 0.80 × 0.35 = 672 completed learner speaking hours
£13,800 ÷ 672 = £20.54 per completed learner speaking hour
Voice-practice programme
- total licence, usage, support, review, and allocated setup cost: £6,400;
- verified completed learner voice time: 54,000 minutes;
- invalid/technical-failure voice time already excluded.
Calculation:
54,000 ÷ 60 = 900 completed learner speaking hours
£6,400 ÷ 900 = £7.11 per completed learner speaking hour
This does not prove equal value. The teacher programme may deliver diagnosis, social interaction, explanation, safeguarding, and accountable assessment that the voice tool does not. The next section prevents cost from swallowing quality.
Pair cost with an outcome and quality dashboard
Report cost per speaking hour beside:
| Dimension | Example measure |
|---|---|
| Reach | Eligible learners who begin and complete |
| Speaking dose | Median completed voice minutes; distribution, not only mean |
| Task performance | Success on an unpractised parallel speaking task |
| Intelligibility | Blind human-listener understanding |
| Interaction | Response to genuine follow-ups and repair |
| Retention | Delayed outcome after access ends |
| Equity | Completion and scoring gaps across groups/devices |
| Reliability | Failed-session and unusable-feedback rate |
| Safety | Reported incidents, privacy exceptions, escalations |
| Experience | Learner and teacher usefulness ratings |
A lower cost with no improvement may be waste. A higher cost may be justified by outcomes or services the cheaper mode does not provide.
Sensitivity analysis
Publish at least three cases:
- low use / conservative speaking share;
- observed base case;
- high but credible use / speaking share.
Also change completion, staff time, implementation allocation, and support cost. If the preferred conclusion disappears under a small assumption change, the business case is fragile.
Procurement questions
- Can the provider export verified learner voice minutes without raw audio?
- What counts as an active or completed voice minute?
- Are system speech and waiting time excluded?
- How are failed and abandoned sessions reported?
- What teacher, QA, privacy, and support work shifts to the institution?
- Are usage overages, minimum commitments, and renewal increases modelled?
- Which learning and equity outcomes accompany the cost metric?
- What happens to the economics at realistic—not contractual maximum—use?
Editorial disclosure
LinguaLive sells AI speaking practice and therefore benefits if speaking-hour economics are valued. Before publication, finance and education researchers should validate the formula, cost categories, measurement protocol, and worked arithmetic. Any public comparison must replace illustrative values with dated, auditable local evidence and preserve the non-equivalence caveat.
Related Topics
Share this article
Ready to Start Learning?
Try LinguaLive's AI-powered conversation practice free. 10 minutes a day can transform your fluency.
Start Free - 10 Min DailyMore Articles
AI-Assisted Language Learning Evidence Map, 2023–2026
Recent studies suggest that AI chatbots and speech-feedback tools can increase practice volume and improve selected speaking outcomes in structured short-term…
Foreign-Language Speaking Anxiety: Evidence Map and Practice Implications
Foreign-language speaking anxiety is a context-sensitive barrier that can affect participation, attention, willingness to communicate, and self-evaluation. It…