- Screening job applications is named in the EU AI Act's high risk list, so switching on a CV filter moves you into a regulated category rather than just a faster process.
- New York City has required an annual independent bias audit and a published result since 2023, with candidate notice at least ten business days before the tool is used.
- The scrutiny follows the decision, not the software. A tool that ranks a shortlist for a human sits in a different place from one that rejects people on its own.
- Bias does not need a protected field to reappear. Postcode, university and a gap in a career history carry it back in after you delete the obvious columns.
- You can run the core audit yourself on a spreadsheet of last year's hiring. The arithmetic is one division per group and it takes an afternoon.
- The high risk obligations apply from 2 December 2027, which is two budget cycles rather than a distant horizon.
Before anything else: automated CV screening is a compliance topic wearing a productivity costume. Whatever the vendor deck says about hours saved, the moment a system filters or scores job applicants you are operating in a category that European law names explicitly and that New York City already audits. That is not a reason to avoid it. It is a reason to buy it with different questions.
Here is the specific text. Annex III of the EU AI Act puts recruitment and selection in the high risk category, and it names the activities: placing targeted job adverts, analysing and filtering applications, evaluating candidates. A second entry covers promotion, termination, task allocation by personal traits and performance monitoring. Most of that is the standard feature list of every applicant tracking system sold today.
What does the high risk label actually oblige?
Different things depending on whether you built the tool or bought it, and for almost every business reading this the answer is bought. That makes you a deployer, and the deployer duty list is about use rather than construction: operate the system according to its instructions, assign human oversight to somebody competent with authority to intervene, keep the logs, watch for problems, and tell candidates where decisions about them are being made this way.
The timing has moved and it is worth getting right. The Commission's page on the regulatory framework for AI now gives 2 December 2027 for high risk systems in sensitive areas, following the AI Omnibus that entered into force on 27 July 2026. We mapped each duty to the document that discharges it in the piece on the AI Act obligation map, and the screening case is the clearest example of why that mapping matters: nothing on the list is hard, and all of it is invisible until somebody asks.
Task by task, and where the risk concentrates
Not every use of a model in hiring carries the same weight. The variable is how much of the decision the system makes without a person, which is why the table below is organised by task rather than by product.
| Screening task | What can go wrong | The control |
|---|---|---|
| Keyword and knockout filters | A rule nobody revisited removes qualified people invisibly, and the rejected set is never inspected | Sample the rejected pile monthly. If nobody has ever read a rejected application, the filter is unsupervised. |
| Ranking and scoring | A score becomes a decision because the reviewer only reads the top of the list | Present the score without a cut off, and require a stated reason for rejection that is not the score |
| Interview scheduling | Low risk in itself, and the place vendors point when asked about AI in hiring | Little needed. Do not let it stand in for the answer about the other rows. |
| Automated rejection with no human review | This is the row that attracts enforcement, and it is the one most likely to be switched on by default | Turn it off, or require a named person to confirm every rejection above a volume you can actually review |
The last row deserves the emphasis. A system that suggests is a tool. A system that rejects is a decision maker, and every regulatory framework in this area treats those differently. Check the default setting rather than the documentation, because the two often disagree.
Why does removing the obvious fields not fix bias?
Because the model does not need the field. It needs a correlate, and hiring data is full of them.
Train a scorer on who you hired before and it learns your past preferences, including the ones you did not intend and would not defend. Delete name and gender and the pattern survives in postcode, which tracks where people can afford to live. It survives in university, which tracks who could afford to go where. It survives in a career gap, which tracks who took time out for caring. It survives in the phrasing of a CV, which tracks first language and class. Each one is a legitimate looking feature. Together they reconstruct the thing you removed.
This is why fairness in screening cannot be established by inspecting the inputs. You have to measure the outputs, which is the entire point of the audit below, and it is why an accuracy figure from a vendor tells you nothing about your own pipeline. Their number describes their test set. Yours describes your applicants.
The strongest signal that a screening model has learned something you did not intend is that it performs unusually well. A model that predicts your historical hiring decisions with high accuracy has learned your historical hiring decisions, including whatever was wrong with them. Accuracy against past outcomes is a warning, not a result.
The audit you can run yourself
New York City wrote the template, and it is simple enough to run on a spreadsheet whether or not the law applies to you. Under the city's rules, an employer using an automated employment decision tool must have an independent bias audit performed and published, calculating the selection rate for each category and the impact ratio against the group with the highest rate, and must give candidates notice at least ten business days before the tool is used, including the qualifications and characteristics it will assess and how to request an alternative process.
The threshold comes from older ground. The federal uniform guidelines on employee selection state that a selection rate for any race, sex or ethnic group which is less than four fifths, or eighty percent, of the rate for the group with the highest rate will generally be regarded as evidence of adverse impact, with the caveats that smaller differences can still matter where they are significant in statistical and practical terms, and that larger differences may not where the numbers are small and not statistically significant.
Here is the arithmetic, worked once on illustrative numbers so the shape is clear. Suppose last year 200 applicants in the group with the highest rate reached interview at a rate of 60 out of 200, which is 30%. In a second group, 10 of 50 applicants reached interview, which is 20%. The impact ratio is 20 divided by 30, or 0.67. That sits below the four fifths threshold, and it is a flag to investigate rather than a verdict, because the second group's 50 applicants may be too few to be statistically meaningful.
| Step | What you do | What you need |
|---|---|---|
| 1. Define the stage | Pick one gate, usually application to interview, and measure only that | Your applicant tracking export |
| 2. Count by group | Applicants and advancers for each category you hold data on | Whatever demographic data you lawfully collect, which may be none |
| 3. Divide | Selection rate is advancers divided by applicants, for each group | A spreadsheet |
| 4. Compare | Impact ratio is each group's rate divided by the highest group's rate | The same spreadsheet |
| 5. Interpret | Below 0.8 is a flag. Check whether the counts are large enough to mean anything. | Judgement, and a note of what you decided |
Two honest caveats. Many businesses do not hold demographic data on applicants and should not start collecting it casually, since that is its own legal question in most jurisdictions. And a formal audit under the NYC rules must be performed by an independent auditor, so the spreadsheet version above is a management control rather than a compliance artefact. It is still the fastest way to discover that something is wrong.
What do you owe the candidate?
Three things, and they are cheap. Tell them a system is involved, before they apply rather than after they are rejected. Tell them what it assesses, in plain terms, so the disclosure is informative rather than defensive. And give them a route to a human, whether that is an alternative process or a request for review.
The New York rules put numbers on the first two: notice at least ten business days ahead, covering the job qualifications and characteristics the tool will use, along with instructions for requesting an alternative selection process or accommodation. Even where no such rule binds you, that is a reasonable standard to hold yourself to, because it is the version you would want to be able to describe if a rejected candidate complained publicly.
The third one is where most policies quietly fail. An appeal route that lands in an unmonitored inbox is not a route. If you offer human review, somebody has to own the queue.
What the tool is actually optimising
Worth saying plainly, because it explains most of the disappointment. A screening model is not trying to find good employees. It is trying to predict which applications a recruiter would have advanced, because that is the only labelled data anybody has. Nobody holds a dataset of applications matched to whether the person turned out to be good at the job three years later, and even where performance reviews exist they encode the same preferences the screening was supposed to remove.
So the target is agreement with past decisions, and that has two consequences worth planning for. Your model inherits your historical shortlisting habits, including the ones you would change if you could see them. And it cannot get better than the process that generated its training labels, no matter how much data you feed it.
The practical version of that for a small business is to stop asking whether the tool is accurate and start asking whether it is consistent. Consistency is real value. Every applicant assessed against the same criteria, on a Friday afternoon as well as a Monday morning, is a genuine improvement over a human reading in a hurry. Consistency is also auditable, which accuracy against a hidden benchmark is not.
Where the records have to live
Two years from now somebody will ask why a specific candidate was screened out. What you want to be able to produce is short: the version of the tool in use on that date, the criteria it applied, the score or outcome it returned, whether a person reviewed it, and who that person was. That is five fields on a row, and almost no applicant tracking system stores them by default.
Check yours before you rely on it. Ask the vendor whether decision logs are retained for the whole period you might need them, whether they survive a configuration change, and whether you can export them without asking support. If the answer to any of those is no, your oversight exists only in the present tense, which is not where the question gets asked.
The retention period is a judgement call rather than a formula, and the sensible anchor is the time limit for bringing a claim in your jurisdiction plus a margin. Write the number down and apply it deliberately, because the default in most systems is either forever or until somebody tidies up, and both are the wrong answer for applicant data.
What to ask a vendor before you switch it on
Five questions, and the useful ones are about defaults and evidence rather than about capability.
Ask whether the system can reject a candidate with no human step, and whether that behaviour is on by default. Ask what it was trained on, specifically whether it learned from a customer's historical hiring outcomes. Ask what bias testing exists, who performed it, and whether the result is published or merely asserted. Ask what you will be able to show a candidate who asks why they were screened out. And ask what happens to the applicant data, which is among the most sensitive material a small business ever handles.
A vendor who answers those in writing has thought about the problem. One who answers with a compliance certification has not, because a certificate describes their product and your obligation attaches to your deployment. That distinction runs through the whole regulation, and we set out how the duty lists differ in the piece on the AI Act's risk tiers and timeline.
Is this worth doing at all?
For high volume roles, often yes. If you receive four hundred applications for a warehouse shift, no human reads four hundred applications properly, and an honest comparison is not against a careful process but against a tired one that stops at the fiftieth CV. A system that ranks consistently can be fairer than a person skimming under time pressure, provided you can show what it did.
There is a second case for it that vendors rarely make and that holds up better than the time argument. A documented, consistent screen is easier to defend than an undocumented inconsistent one. If you are ever asked to explain how a role was filled, a stored rule and a decision log are a better answer than a manager recalling a Tuesday in March. That benefit only exists if you keep the records, which is why the logging section above sits where it does rather than at the end as an afterthought.
For low volume specialist roles, usually not. Twelve applications is a reading task, and the overhead of oversight, notice and audit exceeds anything you save. The trap is buying a tool for the first case and leaving it switched on for the second.
What decides it either way is whether you can inspect the outcome. If you can pull a rejected pile and read it, run the audit and explain a decision, the system is a tool you supervise. If you cannot, you have outsourced hiring judgement to something that cannot be questioned, and every framework in this area is built to stop exactly that. The governance version of the same argument, applied across every AI system in a business rather than to hiring alone, is in our piece on the first ninety days of AI governance, and the written rules that make it stick belong in something like a published AI policy rather than in an email thread.
The one thing to do this week
If you take a single action from all of the above, make it this one, because it costs nothing and it is the only test that uses your own applicants rather than somebody else's benchmark.
Export the rejected applications from your last open role and read twenty of them. Not the hires, the rejections. That single act tells you more about your screening than any vendor benchmark, and if you find somebody in there you would have interviewed, you have learned the most valuable thing this article can offer you before you have configured anything at all.