A provider says the team is at B2. The introductory call with the account lead goes well. Three weeks after launch, someone on the client side listens to a recording and can't follow half of what the agent said. Nobody lied in the proposal. It was measured badly.
Language is the most declared and least verified variable in a nearshore operation. It's also the one that destroys trust fastest when it fails, because the client perceives it directly: no need to wait for the monthly report or interpret a metric. They hear it.
What you're assessing isn't the accent
The first mistake is confusing accent with capability. A strong accent doesn't block communication. A thin vocabulary does. What decides whether an interaction works is intelligibility: that the other person understands first time, without asking for repeats, and that the agent understands what they're told even when it comes fast, with background noise, or from someone in a bad mood.
Listening comprehension is the half almost nobody tests. In an interview the candidate speaks; in the operation, they mostly listen. Holding a conversation about yourself, with an assessor who enunciates and waits, is nothing like understanding a rushed customer using a regional expression to describe a problem with a product the agent learned about last week.
A candidate can hold an interview in English and not hold a support call. It isn't the same language in use.
Process English isn't general English
Every operation has its own vocabulary: product names, order statuses, system fields, return policies, the exact term this client uses for something another company calls something else. That lexicon doesn't appear in any standard language test, and it's what gets used for most of the shift.
It's worth separating the two levels. The general one is hired: it's the floor a person brings. The process one is taught: it belongs to training and to the knowledge base. A provider measuring only the general level is delivering half the job, and usually the half that shows least in selection and most on the floor.
What a declared level doesn't tell you
Reference frameworks, like the CEFR with its A1 to C2 scale, are useful for speaking a common language. They aren't a guarantee on their own, because a level with no method behind it isn't data: it's a label.
Faced with a "B2 team", three questions organise the conversation: who assessed it, with what test, and how long ago. A candidate self-assessment, an online written test and an evaluation with a certified instructor all produce the same letter and very different realities. And a certificate from three years ago says little about someone who hasn't used the language since.
How to test it before signing
Language assessment is one of the few things a buyer can verify directly, without depending on the provider's report. Worth doing properly:
- Ask for audio samples. Recordings of role-play exercises from the selection process, or from live work with clients and data anonymised and properly authorised. Five minutes of listening says more than reading a profile.
- Run a live case, not an interview. Hand over a scenario from the real process and ask them to work through it out loud. CV questions get rehearsed; a fresh case doesn't.
- Include a comprehension segment. A real audio clip, at normal speed and with normal noise, followed by questions about what was said. It's the part that discriminates most and the one almost nobody includes.
- Assess writing separately. A chat or email case, time-boxed, with autocorrect off. Spelling and written tone are a different skill.
- Test the people who will be on the account. Not the commercial team, not the lead who runs the demo.
Point five is the most awkward to ask for and the most useful of all. It's also something worth checking in person: it's one of the things you can see during a provider site visit, by listening to the floor instead of reading a slide.
Voice and writing get scored separately
Someone excellent on chat can be mediocre on voice, and the other way round. In text there's time to think, to revise and to lean on templates. On voice there is none of the three. If the operation combines channels, the standard can't be a single one, and neither can the test.
This has a practical consequence for team design: not every position needs the same level. An operation can work well with a high standard on the roles that speak to the end customer and an intermediate one in back office, where the language is used to read instructions and record information. Demanding the maximum everywhere raises cost without improving the outcome, and makes hiring at scale harder than it needs to be.
The level doesn't stay put
The team assessed during selection isn't exactly the team on the account six months later. There are replacements, promotions, attrition. The question usually left unanswered is who assesses the replacement, and with what test.
Language also needs use to hold. Someone with a solid level who spends months on an account where they only write short emails loses spoken fluency. It isn't a commitment problem; it's how a second language works. Which is why language refreshers and feedback sessions with a linguistic criterion aren't a launch luxury: they're maintenance.
What to put in writing
- The minimum level per position and per channel, naming the framework used and who certifies it.
- The assessment method: which test, who applies it, and whether the client can listen in or take part in a sample.
- The rule for replacements: same standard, same test, and within what timeframe it gets validated.
- Language inside the quality rubric, not as a separate review nobody performs.
- What happens if someone falls short: improvement plan, timeframe, and the criterion for removal from the account.
None of this makes the contract more expensive. It just avoids the awkward month-three conversation, when there's already a recording on the table and two versions of what was agreed.
Our approach at smartBPO
We assess language with separate spoken and written tests, including a listening comprehension segment, and we run the test on the people who stay on the account. We declare the level while stating who measured it and how, instead of handing over a letter with nothing behind it. We set the standard per position and per channel, not one for the whole team. We include the language criterion inside the quality rubric, so it gets reviewed as often as everything else. And we validate replacements with the same test as the original cohort. We open that process to the client: if you want to listen to samples or take part in the assessment, better before signing than after the first complaint.