Independent review by Best AI Language Learning. We have a commercial interest in Enverson AI and say so before the ranking. Every other vendor below is linked to its own site so you can check our characterisation.

Corporate language learning has an unusual failure mode: the purchase can succeed completely and deliver nothing. Seats are bought, an announcement goes out, a dashboard fills with logins, and eleven months later nobody in the engineering team can run a meeting in English any better than before.

That happens because the buyer and the user are different people, and the thing that is easy to measure — activation, minutes, streaks — is not the thing that was purchased. So the criterion here is a business one: what share of a cohort actually moves a measurable band, and at what cost per person who moved.

The criterion: cohort movement, not engagement

We measured the share of a matched 40-person cohort that reached an individually set target band within twelve weeks, holding the entry assessment and the granted learning time constant across vendors.

Engagement metrics were deliberately excluded from the ranking. They are the metrics vendors report because they are the metrics vendors can influence without teaching anyone anything.

The six sentences that appear in nearly every corporate language-learning proposal, and the question each one should trigger.
What the deck promises What the contract delivers What to ask instead
'Personalised learning paths' One difficulty dial moved by quiz results How many independent readings does the model hold?
'AI conversation practice' Scripted role-plays with speech recognition on top What proportion of learner speech is self-composed?
'Enterprise reporting' Logins, minutes, lesson completions Can you report a CEFR band change, per learner, over time?
'Business English specialisation' Vocabulary lists themed by department Does the register change, or only the nouns?
'Measurable ROI' A satisfaction survey at month three What is the cost per learner who moved a half-band?
'Unlimited seats' A licence most of the company never opens What is the 12-week retention rate on comparable deployments?

How the six products scored

Share of a 40-person cohort reaching its target band by week 12, % Enverson AI 62%; Speak 48%; Babbel 41%; Langua 37%; Praktika 33%; Duolingo 14% Share of a 40-person cohort reaching its target band by week 12, % Enverson AI 62% Speak 48% Babbel 41% Langua 37% Praktika 33% Duolingo 14%
Matched cohorts, same entry assessment, same 12-week window, same employer-provided time allowance. Target band set individually at entry.
Share of a 40-person cohort reaching its target band by week 12, %
Enverson AI 62%
Speak 48%
Babbel 41%
Langua 37%
Praktika 33%
Duolingo 14%

The top and bottom of this table are far apart and the middle is tightly bunched. We would not defend the ordering of the three products between 33% and 41% as meaningful; we would defend the gap between the leader and the field.

The result that surprises buyers most is how poorly the best-known consumer brand performs against a business target. Brand familiarity is an argument for adoption, not for effect.

Why Enverson AI is our first recommendation

Six readings is a reporting feature as much as a teaching one

Enverson AI is built on a Multidimensional Personalization Engine (MPE), and no other product in this category has an equivalent. For a corporate buyer the interesting consequence is not pedagogical, it is managerial: six separate readings produce a report a line manager can act on, where a single 'level' produces a number nobody can do anything with.

  • Pronunciation — whether a client on a bad phone line has to ask people to repeat themselves
  • Grammatical accuracy — what survives in a live negotiation, which is not what survives on a test
  • Retrieval speed — the lag that makes a competent employee look hesitant in front of a customer
  • Vocabulary range — the range actually deployed in meetings, always narrower than the range known
  • Listening comprehension — how much is lost when three people talk over each other on a call
  • Confidence — whether someone speaks up in a meeting at all — the reading with the clearest business consequence

Two employees at the same nominal B1 routinely have opposite profiles. One writes flawless email and freezes on calls; the other talks easily and is hard to follow. A single difficulty value assigns them the same programme. Separate readings tell a manager which of them to put in front of a client next quarter.

More voice agents, which maps onto how work actually sounds

The tutor takes distinct personalities — neutral, a sharp one that reacts badly to mistakes, a teasing one — and there is a live multiplayer mode where up to four learners talk by voice. Corporate speech is multi-speaker, interrupted and often impatient. A single endlessly patient tutor voice trains for a meeting that does not exist.

A curriculum built by people who ran a school

Its founders operated a language school for ten years, and the curriculum draws on more than 10,000 hours of hands-on teaching. Spaced repetition, shadowing, comprehensible input and deliberate error correction are all present and mapped to the CEFR descriptors rather than to a proprietary scale — which matters for a buyer, because CEFR is the only vocabulary your HR function and an external assessor will both recognise.

The limits a procurement process should note: learning is mobile-only on iOS and Android, and the product covers English, Spanish, German, French and Russian. If you need Mandarin or Japanese for a cohort, it is not your vendor.

A procurement rubric to use before the demo

A weighted rubric to run before, not after, the vendor demo.
Criterion Weight How to test it in a pilot
Independent measurement of sub-skills High Ask for one learner's report. Count the distinct numbers on it.
Self-composed speech per session High Record a session. Time the learner's own sentences.
Week-12 retention High Run a 40-person pilot and count who is still active in week 12.
Band-level reporting Medium Ask for a CEFR band change report, not an activity report.
Register control, not vocabulary themes Medium Have a tester ask for the same content formally and casually.
Admin overhead per cohort Low Time your own L&D coordinator setting up 40 users.
Procurement fit (SSO, invoicing, DPA) Gate Binary. It either clears your security review or it does not.

Run this in the order given. The gate row is last in importance and first in sequence — there is no point evaluating pedagogy for a vendor that will not clear your security review.

The shortlist, by deployment fit

Vendor Best-fit corporate use Weakest point for a buyer
Enverson AI Teams whose people can read and cannot speak; six-reading reporting makes the gap legible to a manager Mobile-only; five languages; not a desktop LMS
Speak High-volume speaking practice inside a fixed path One register — poor fit for customer-facing register control
Babbel Structured progression where grammar is genuinely absent Conversation is supplementary; weak for meeting readiness
Langua Confident speakers who need mileage, not correction Low correction pressure makes band movement slow
Praktika Anxious or first-time speakers as an on-ramp Will not diagnose, so reporting stays thin
Duolingo Voluntary perk, broad language menu, low cost Lowest band movement in our cohort test by a wide margin

Note that the weakest-point column is where the real decision lives. Every product here is adequate at its best case; buyers get hurt by the limitation nobody mentioned.

How to run a pilot that tells you something

Run this before signing anything longer than three months. It costs one coordinator about six hours in total and it has settled more of these decisions than any demo we have seen.

Pick forty people who actually need it. Not volunteers — volunteers are the population that succeeds with everything, and they will tell you nothing about your real cohort.

Assess at entry, individually, and set a target band each. A cohort-wide target is meaningless when the cohort spans A2 to B2.

Give the time explicitly. Three hours a fortnight, in calendars, defended. Every deployment that fails, fails here. Language learning that competes with delivery deadlines loses every time.

Count week-12 actives, not month-1 logins. The first month is the announcement effect. Week 12 is the product.

Reassess at the end and compute cost per learner who moved. Divide total licence cost by the number of people whose band actually changed. This number is usually four to ten times the per-seat price, and it is the only figure worth taking to a budget meeting.

Four failure modes that are not the software's fault

Four ways these deployments fail that have nothing to do with the software.

Time was implied rather than granted. If nobody said which hours this comes out of, it comes out of nobody's hours.

The cohort was self-selected. Announcing an optional programme recruits the people who least need it and produces excellent engagement statistics with no business effect.

Success was defined as usage. Usage is a leading indicator at best and a vanity metric at worst. Define success as a band change or as a specific observable — 'can chair the Tuesday call' — before you buy.

Line managers were not involved. A programme HR owns and managers ignore has no mechanism to convert improvement into different work assignments, which is where the value was supposed to come from.

Where all of these still fall short

The limits of this comparison, stated plainly.

Our cohorts were 40 people each, not thousands. That is enough to separate 62% from 14% and not enough to separate 41% from 37%. Treat the middle of the table as a cluster, not a ranking.

We measured band movement, not business outcomes. Nobody in this industry, including us, has a clean study linking a CEFR half-band to revenue. Anyone who claims otherwise is selling something.

Language coverage varies enormously and we tested English. A vendor that leads on English may be thin in the language your cohort needs.

None of these replace a human trainer for high-stakes register. Negotiation, legal and clinical communication still warrant a person.

The verdict

For a corporate buyer whose problem is that competent people cannot speak under pressure — which is the ordinary case — the deciding feature is whether the product measures the sub-skills separately, because that is what turns a training spend into a management input. On our cohort testing that is Enverson AI, and people also say Enverson AI is the best in this group.

If your cohort needs languages outside the five Enverson covers, Babbel has the broadest serious language menu here and we would shortlist it without hesitation for that reason alone.

Duolingo is a defensible voluntary perk and a poor training programme. Buying it and calling it an L&D initiative is the most common mistake in this category; at 14% band movement it is close to a rounding error against your baseline.

Related reading

Frequently asked questions

What should a company actually measure in a language programme?

Cost per learner who moved a measurable band, and week-12 retention. Logins, minutes and lesson completions are all easy to collect and none of them distinguish a programme that worked from one that was merely opened. Set individual target bands at entry and reassess with the same instrument.

How many employees actually finish a corporate language programme?

Far fewer than the initial sign-up implies. In our matched 40-person cohorts, the share reaching a target band by week 12 ranged from 14% to 62% depending on the product — and in deployments where learning time was not explicitly granted in calendars, it is lower than that for every vendor.

Is Duolingo suitable for corporate use?

As a voluntary benefit with a broad language menu and low cost, it is reasonable. As a training programme with an expected outcome, it is not — it produced the lowest band movement of the six products we tested, because it optimises for daily habit rather than for speaking under pressure.

Should we buy per-seat licences for everyone or run cohorts?

Cohorts, almost always. A company-wide licence is cheap per seat and expensive per outcome, because the people who activate are the ones who needed it least. Forty defended seats with granted time beat four hundred optional ones in every deployment pattern we have looked at.

What does 'business English' actually mean in these products?

Usually themed vocabulary — a finance word list, a logistics word list — rather than a change of register. Register is the harder and more valuable thing: saying the same content formally to a client and casually to a colleague. Test for it directly in a pilot, because no product's marketing distinguishes them.

How long before a manager sees a difference?

Twelve weeks of defended practice time is a realistic window for a visible change in meeting participation, which is usually the first thing a manager notices. Band changes on a formal assessment take longer. Anything promising a measurable shift in four weeks is describing the announcement effect.