Overview
Every remote candidate's profile says "fluent English." Most first calls go fine. And a surprising number of hires still fail on English, three months in, when the first difficult customer email or the first ambiguous instruction arrives and the person cannot handle it. The gap is not that candidates lie. It is that "fluent" is not a measurement, a friendly call is not a test, and the English a job actually needs is rarely written down before the search starts.
Testing English proficiency properly means three things. Deciding which level the role requires, in a scale that means something. Testing written and spoken English separately, because they are different skills and candidates are routinely strong in one and weak in the other. And testing under conditions that resemble the work rather than a quiz.
This guide covers each, with the Philippines as the worked example, since it is where most English-language remote hiring lands.
Start with the level the job needs
The scale that hiring teams around the world use is the CEFR, the Common European Framework of Reference, which places a person on six bands from A1 to C2. Two of them matter for remote hiring.
B2 is the "independent user" band. A B2 speaker follows an argument without help, writes clearly and in some detail on familiar subjects, understands complex work instructions and can present an idea in a meeting. Hiring guides converge on B2 as the minimum for any role with regular client contact, and the reason given is practical: below B2, communication breakdowns in professional settings become frequent enough to cost money.
C1 is "proficient." The step from B2 to C1 is not more grammar; it is precision and flexibility under pressure. A C1 writer expresses a nuanced point without visibly hunting for words, adjusts register for a lawyer and a customer in the same hour, and keeps the quality up when the topic gets hard or the deadline gets short. Roles where the person writes on the company's behalf, handles escalations, or drafts anything a client will read unedited are C1 roles, and hiring at B2 for them is where the three-month failures come from.
The exercise, before a single candidate is contacted, is to name the band per task: reading tickets (B1 is often enough), replying to customers in writing (B2 at least), phone support (B2 spoken, with accent and listening weighed separately), drafting proposals or managing a client relationship (C1). A role is hired at its highest-band task, not its average.
Why written and spoken English are tested apart
The two skills diverge more than most managers expect. A candidate who speaks warmly and confidently on a call can produce written English full of errors that a customer will notice on the first line. A candidate who writes precise, well-structured English can be halting on a call, especially under the stress of an interview with a stranger in another country.
Roles differ in which one they need. A chat and email support seat is a written-English job; the phone may be irrelevant. An appointment-setting or sales role is a spoken-English job; the emails are templates. Testing one and inferring the other is the most common vetting mistake, and it is a mistake in both directions: strong writers rejected on a nervous call, weak writers hired on a charming one.
Written English is best tested with a real writing task, not a grammar quiz. A grammar quiz measures whether someone can pick the right option among four; it says nothing about whether they can draft an apology to an angry customer in a tone the company would sign. Give the candidate a realistic situation and ask for the email, the summary or the reply, with the same time pressure the job carries. Score it against a standard: is the meaning unambiguous, is the register right for the reader, would this go out unedited. That is a B2/C1 judgment and it can be made consistently once the standard is written down.
Spoken English is best tested in a live conversation by someone trained to rate it, on more than one axis. Clarity, listening comprehension, the ability to handle an unexpected turn, and whether the accent is intelligible to the customers this company actually serves. Note the last point: intelligibility to a specific audience is the job requirement, not a neutral accent, and a rater who knows the customer base makes a better call than a generic proficiency score.
The Philippines, as the worked example
The Philippines is the largest English-speaking labor market for remote work, and the reason is structural rather than incidental. English is an official language, a language of instruction from primary school, and the language of law, business and government. In the 2025 EF English Proficiency Index the country ranks 28th of 123 countries with a score of 569, in the "high proficiency" band, second in Asia behind Malaysia and well above the global average of 488.
That average hides a wide distribution, which is the whole point of testing. A national ranking says a Filipino applicant is very likely to be somewhere between B1 and C1. It does not say where. Two candidates with identical résumés from the same university can sit a band apart in written English, and a manager who hires on the national reputation instead of the individual test gets the distribution rather than the person.
This is the logic behind how Flex screens its own applicants. Every candidate who reaches a client's shortlist has sat an assessment that Flexie, the company's AI examiner, builds around the applicant's own field and runs live, so no two people see the same exam; written English is graded there against a fixed standard rather than against the other applicants. Spoken English is rated separately, in a recorded interview run by a member of the Flex team, on its own scale alongside communication, problem solving and field expertise. A client hiring through Flex meets only people who have cleared both, which is a different thing from meeting people whose profiles say "fluent."
Conditions that resemble the work
A last principle, which applies to any test in any country: proficiency measured in comfortable conditions overstates proficiency in real ones. The three conditions that matter are time, ambiguity and stakes.
Time: a writing task with an hour and a dictionary measures a different skill than a reply drafted in eight minutes between two other tickets. Test at the job's pace. Ambiguity: real instructions are incomplete; a test that asks the candidate to act on a deliberately underspecified brief, and watches whether they ask the right question or guess, reveals more than a clean prompt does. Stakes: an interview in which the candidate knows they are being rated on English produces careful, rehearsed English. A conversation about the work itself, rated on English afterwards, produces the English the customer will get.
A test that is too easy passes people who fail in the job. A test that is too hard, or too artificial, rejects people who would have been excellent, and in a market as deep as the Philippines that is a real cost, because the excellent ones have other offers.
Conclusion
Testing English proficiency in remote hires is a matter of deciding the band per task before the search, testing written and spoken English as the separate skills they are, and testing them under something like the job's own conditions. The CEFR gives the vocabulary: B2 for independent professional work, C1 wherever the person writes or speaks on the company's behalf without a second pair of eyes.
The Philippines makes the case for individual testing better than any other market. The national proficiency is high and the spread is wide, which means the reputation is a reason to search there and never a reason to skip the test.