- New York City's rule on automated hiring tools applies to employers of any size, with no exemption for a shop hiring its first assistant.
- That rule wants an annual bias audit, the results published on your site, and 10 business days of notice to candidates before the tool is used.
- The federal fairness test compares selection rates between groups, and it is arithmetically useless when you are choosing one person from twelve.
- Because the statistical check is unavailable at your scale, the defensible safeguard is procedural: fixed criteria written before you read anything.
- You remain liable for a tool a vendor built, and a vendor's assurance about fairness does not transfer the responsibility.
- In the EU, placing targeted job advertisements is itself listed among the high risk recruitment uses, not only the sifting.
The first hire is the moment a business stops being one person with a laptop. It is also, for most owners, the first time they have ever read an application, and the temptation to hand the reading to software is entirely understandable when sixty people answer an advert for one part time role.
What almost nobody expects is that ai hiring rules were not written with a headcount floor. The main American rule on this subject applies to an employer of any size. The European framework puts recruitment in its most heavily regulated category. Neither asks how many people work for you.
Does any of this apply to a shop with two employees?
Yes, and that is the single most surprising fact in this subject. New York City's Local Law 144 is the clearest case: the city's own guidance on automated employment decision tools describes a law that applies to employers of any size operating in New York City, with no exemption based on company size.
The obligations attach to the tool rather than to the payroll. Where an automated employment decision tool is used, the city requires that it has been subject to a bias audit within one year of its use, that information about that audit is publicly available, and that notices have been provided to employees or job candidates. Enforcement by the Department of Consumer and Worker Protection began on 5 July 2023, and the notice period candidates are owed is 10 business days before the tool is used in their evaluation.
Read what that means operationally for a small employer. Publishing a bias audit summary on your own website is not something a two person business has infrastructure for, and commissioning an independent annual audit of a tool you pay forty dollars a month for is plainly disproportionate. The realistic response is not compliance at that scale. It is to avoid using a tool that meets the definition in the first place, which is a design choice available to you and much cheaper than the alternative.
What counts as an automated employment decision tool?
Something that substantially assists or replaces discretionary decision making in hiring or promotion. The city's own frequently asked questions use exactly that framing, and the word doing the work is substantially.
A tool that produces a ranked list of candidates substantially assists the decision, because the order is the decision in practice. A tool that scores applicants against a profile does the same. A tool that transcribes an interview does not. A model that reformats sixty CVs into a consistent layout so you can read them faster does not, because you are still doing the reading and the judging.
That distinction is the whole practical answer for a small business. Use the technology to make the material readable and leave the ordering to yourself, and you stay outside the definition while getting most of the time saving. Ask it who the best five are, and you have bought an obligation.
Why the standard fairness test does not work at your size
Because it is a ratio between selection rates, and a selection rate computed from twelve applicants carries no information. This is the part of the subject that nobody writes honestly about, and it changes what a small employer should actually do.
The federal test is the four fifths rule. As Littler's summary of the EEOC technical assistance issued on 18 May 2023 puts it, where a selection rate for any race, sex, or religious or ethnic group is less than 80 percent of the rate of the group with the highest selection rate, that generally indicates disparate impact. The worked example in the guidance is clean: 48 of 80 men advance, a rate of 60 percent, against 12 of 40 women, a rate of 30 percent. Half of 60 is well under the four fifths threshold, so adverse impact is indicated.
Now try it on your hire. Twelve applicants, one job. Seven of the applicants are men and five are women, and you interview three men and one woman. The male selection rate is 43 percent and the female rate is 20 percent, giving a ratio of 47 percent, which fails the test dramatically. Move one candidate and the numbers invert. The test is measuring nothing at this scale, which is precisely why the guidance warns that the rule is a benchmark rather than a verdict and may be inappropriate where the data lacks statistical significance.
This cuts both ways and it is worth being clear about it. A small employer cannot prove fairness statistically, and equally cannot be shown to be unfair statistically from a single hire. What remains available to anyone examining your decision is the process: what criteria you used, whether they were fixed in advance, whether everyone was asked the same things, and whether you can explain the rejections. That is the record you should be building.
The procedural substitute, which is also just good hiring
Write the criteria before you look at a single application. This is the whole method, it costs twenty minutes, and it is the only defence that works at a scale where numbers cannot speak for you.
Four to six criteria, each one observable, each one genuinely connected to the job. For a shop assistant: availability on the days you actually need, experience handling cash or a till, comfort with the physical demands, a spoken language your customers use. Score each applicant against those and nothing else. The discipline is not bureaucratic. It stops you from noticing, halfway through a pile of CVs, that you have started selecting for people who remind you of yourself.
| Hiring step | Safe use of AI | What crosses the line |
|---|---|---|
| Writing the advert | Drafting and tightening the copy | Targeting the advert by inferred characteristics |
| Receiving applications | Reformatting into a consistent layout | Ranking or scoring candidates |
| Sifting | Extracting stated facts you asked for | Inferring traits not stated in the application |
| Interviewing | Transcribing with notice and consent | Scoring tone, expression or word choice |
| Deciding | Summarising your own notes | Producing the shortlist for you |
| Rejecting | Drafting a reply you read and send | Sending automated rejections unread |
The right hand column has a pattern. Every item in it is a moment where the software forms a view about a person rather than handling text about a person. That is the boundary, and it is easier to hold than any rule about specific products.
The specific ways a model gets a CV wrong
Not by being prejudiced in the way people imagine, but by being confidently wrong about facts it inferred. Three failure modes account for most of what goes wrong when software reads applications, and all three are recognisable once you know to look.
The first is inference from absence. An application that does not mention a qualification is treated as an application from someone who lacks it, which is frequently false for people who wrote a short CV or who described the same thing in different words. Career changers, people returning after a break and anyone whose experience sits outside the vocabulary of your industry lose disproportionately to this. The fix is to ask for the facts you need in structured fields on the application form rather than hoping to extract them from prose.
The second is proxy signals. Nothing in a model needs to reason about a protected characteristic to correlate with one. A postcode, the name of a school, the years a role covered, a gap in dates, the phrasing of an email address: each carries information about background that has nothing to do with whether somebody can do the job. You cannot inspect a commercial tool for this, and you can avoid handing it the fields where such signals live. Strip addresses and dates of birth before anything automated touches an application.
The third is fluent summarisation that loses the thing that mattered. A four line summary of a two page CV is an editorial act, and the detail it drops is the unusual one, which is often the reason a candidate is interesting. If you read only summaries you will systematically hire the most conventional applicants, which is a real cost rather than a compliance problem.
Is an AI interview a good idea for a small employer?
No, and this is the clearest recommendation in the piece. A one way video interview scored by software is the highest risk thing on the menu and it buys a small business almost nothing.
The risk is concentrated because such a tool evaluates a person directly, which puts it squarely inside every definition discussed above, and because the signals it uses, speech patterns, pace, expression, are the ones most closely tied to disability, to accent and to first language. Several jurisdictions have legislated specifically about video interview analysis, and the direction of travel everywhere is toward more constraint rather than less.
The benefit is small because you are hiring one person. Asynchronous video screening exists to compress a funnel of two thousand applicants into two hundred. At twelve applicants you can simply speak to the six who look plausible, which takes an afternoon, produces far better information, and leaves the candidate with an impression of a business that treated them like a person. For a local employer competing for staff against larger companies, that impression is an asset rather than an overhead.
Who is responsible when the tool gets it wrong?
You are. This is settled and it is the point most likely to be misunderstood by a business that reasonably assumes buying a product transfers the problem to whoever built it.
Mayer Brown's reading of the same EEOC guidance on algorithmic decision making tools records the position directly: a third party's assurances or representations about compliance will not necessarily shield an employer from liability. The guidance goes further and suggests employers ask vendors what metrics they have used to assess whether their tools produce adverse impact, and conduct their own analysis on an ongoing basis rather than relying on what they were told.
For a small employer that recommendation is aspirational rather than achievable, which returns you to the same conclusion. If you cannot audit the tool and cannot escape responsibility for it, the sensible position is not to delegate the judgment to it. Buy the time saving, keep the decision.
What does Europe add?
Breadth. The EU classifies recruitment among its high risk uses, and the category is drawn wider than most employers expect, reaching the advertising stage rather than starting at the sift.
Annex III of the EU AI Act covers systems intended for the recruitment or selection of natural persons, and it names placing targeted job advertisements, analysing and filtering job applications, and evaluating candidates. A second limb covers decisions affecting terms of a working relationship, promotion, termination, task allocation based on individual traits, and monitoring or evaluating performance.
The targeted advertising inclusion is the one worth noticing, because it catches a practice a shop would never think of as a hiring tool. Boosting a job post to an audience an advertising platform assembled for you is a form of selection happening before anyone applies, and it is the stage where the applicant pool is shaped most decisively. The narrower point about monitoring people once they work for you is a separate subject we covered in the piece on what staff monitoring is not allowed to do.
Where AI earns its place in a small hire
In everything that is writing or logistics rather than judgment, which is a surprising share of the work.
The job advert is the best case. Most small business adverts are simultaneously too vague about the role and too specific about the person, and a model is good at separating those. Describe the work, the hours, the pay and the conditions precisely, and describe the person loosely, which is both better practice and a narrower target for a complaint.
Scheduling is the second. Arranging eight interviews around a shop rota is pure logistics and the tool has no view about anybody.
Interview notes are the third, with a condition. A transcript frees you to listen instead of writing, which measurably improves the interview. It is also a recording of an identifiable person, so the candidate has to be told before it starts and given a real option to decline, for exactly the reasons set out in the piece on consent when a notetaker joins a call. Store the transcripts where only you can read them, and delete them on a schedule rather than leaving them in a shared workspace forever. Where that data lives matters as much as what is in it, which is the thinking behind our approach to holding data.
And rejections. Writing sixty polite declines is the job owners avoid, which is why so many applicants hear nothing, which is a small cruelty with a real reputational cost in a local labour market. Drafting them is a legitimate use. Sending them without reading is not, and the difference takes a minute.
What to keep, and for how long
Keep the criteria you wrote, the scores you gave, and a one line reason for each rejection. A year is a reasonable default in most places, and your local rules may set a longer minimum.
The reason to keep them is not litigation, which is unlikely. It is that these records are the only version of your decision that exists outside your memory, and memory reconstructs hiring decisions in a flattering and unreliable way. If you hire twice more over the next three years, the criteria you wrote the first time are the starting point for the second, and comparing them tells you something true about how your idea of the role changed.
Keep them somewhere that is not the same shared drive your part time staff use. Applicant data includes information people gave you in confidence about their circumstances, and the most common failure in small business hiring is not discrimination, it is a CV sitting in a folder everyone can open. The general question of who may see what inside a small team is one we set out in the piece on writing an AI use policy.
That sentence is the reason this article recommends the boring version. A small employer has no bias audit, no statistics that mean anything, and no way to inspect a vendor's model. What it does have is the ability to write down what it was looking for before it started looking, ask everyone the same questions, and keep the answers. Those three things are worth more than any tool on the market, and they were good practice long before software was involved.