A Bias Audit May Not Cover How You Use an AI Hiring Tool
A vendor’s audit PDF is evidence about a particular test, not a compliance passport. Employers need to match its data, jobs and decision point to the deployment.
August 9, 2026 · 7 min read

The document to scrutinize is the bias-audit summary attached to an AI hiring product, especially when procurement receives it before anyone has mapped how the employer will use the product.
Consider a customer-support requisition with two automated screens. A resume model ranks applicants and sends the highest-scoring group to a recorded interview, where another model scores responses for communication traits. A recruiter then reviews those scores before deciding who reaches a live interview.
The vendor supplies an audit covering the resume model’s recommendations across several employers. That report may be relevant evidence. It does not, by itself, establish that the interview scorer was tested, that customer-support applicants appeared in the audit data, or that the employer gave the notices required where the job is located.
This distinction matters under New York City’s Local Law 144, the most concrete US example of an enforced bias-audit rule for automated hiring. Enforcement began in July 2023. The law says an employer or employment agency may not use a covered automated employment decision tool unless the tool “has been the subject of a bias audit conducted no more than one year prior to the use of the tool” and a summary of the results is publicly available.
The requirement attaches to use. It does not certify a product for every customer and workflow.
Start with the decision the system makes
New York City defines an automated employment decision tool through both its technology and its role. It must use machine learning, statistical modeling, data analytics, artificial intelligence, or a related computational process to issue a simplified output, such as a score or recommendation, that “substantially assist[s] or replace[s] discretionary decision making” in hiring or promotion.
The city’s rules make that operational. A tool substantially assists when an employer relies only on its output, gives that output more weight than any other criterion, or uses it to overrule conclusions based on other factors. A product containing AI is not covered merely because it supports recruiting. The way the employer places its output in the decision chain matters.
Return to the customer-support workflow. If the audited resume rank determines who receives an interview, the audit needs to correspond to that screening decision. If managers instead use the score to overturn live-interview recommendations, the same output occupies a different selection stage and carries different weight. An audit based on early-stage advancement may not describe disparities at the final-offer stage, where the candidate pool and decision rule have changed.
That is the first comparison to make: identify the exact output, who sees it, what happens immediately afterward, and whether a person can reverse it. “Recruiting support” is too broad for this purpose.
Job category is part of the evidence
A pooled audit can conceal deployment differences. New York City’s implementing rules address this by requiring calculations to be made separately for each employment category when a tool is used across different categories, while the published summary must explain the data source and disclose exclusions.
That requirement reflects a technical problem rather than a paperwork preference. A resume scorer may rank candidates differently when one job rewards a license and another rewards previous call-center experience. Applicant populations also differ. An impact ratio, which compares a group’s selection or scoring rate with that of the highest-rate group, can look different after the employer changes the requisition, threshold, knockout questions, or meaning assigned to the score.
The customer-support employer should therefore look beyond the vendor’s product name. The audit summary should reveal whether the data included the relevant employment category, whether results were aggregated across customers, and whether small or missing groups were excluded. When a report says only that a platform was tested on “applicants,” it leaves unanswered whether the tested population resembles the one passing through this requisition.
New York City’s rules permit historical data from one employer or multiple employers. Test data may be used when sufficient historical data is unavailable, but the summary must identify and explain the source. Synthetic or otherwise constructed test data can help before deployment, yet it cannot reproduce every relationship among local labor supply, application channels, job requirements, and employer-specific thresholds.
A mismatch does not automatically prove unlawful discrimination. It limits what the audit can support.
An impact test is not job validation
Bias audits under the city rules calculate selection or scoring rates and impact ratios across specified sex and race or ethnicity categories, including intersectional categories. These calculations can expose a disparity in who advances or how groups score. They do not explain why the disparity arose, establish that a model measures a job requirement accurately, or cover every class protected by employment law.
That last boundary is easy to miss. A report focused on sex and race or ethnicity does not test disability access, age discrimination, or whether speech and facial-analysis features disadvantage particular applicants. Federal civil-rights obligations continue to apply even when a local audit has been completed.
The US Equal Employment Opportunity Commission has warned that an employer may remain responsible when a vendor’s software causes discriminatory outcomes. Under federal employment law, “validation” has a more specific meaning than showing an acceptable impact ratio: employers may need evidence that a selection procedure is job-related and consistent with business necessity, particularly when it produces adverse impact.
For the customer-support role, that means asking what a “communication” score measures. If the recorded-interview model evaluates word choice, vocal patterns, response timing, or visual behavior, a resume-model audit says nothing about those inputs. Even an audit of the interview scorer would not, on its own, show that the score predicts successful support work or that applicants can use the system accessibly.
The fallback is human review, but only if it changes the decision rather than rubber-stamping a ranking. Employers should document what reviewers receive, which underlying application materials remain visible, and when a reviewer can advance someone despite the model’s output. Otherwise the nominal human step may leave the automated threshold intact.
Notice is a separate control
New York City’s audit requirement and candidate-notice requirement sit in the same law, but satisfying one does not satisfy the other. For covered uses, the employer or employment agency must notify a candidate who resides in the city that an automated tool will be used and identify the “job qualifications and characteristics” it will assess. The notice must arrive at least 10 business days before use.
The law also requires a route for requesting an alternative selection process or accommodation. Its data-disclosure provisions address information collected by the tool, the source of that data, and the employer’s retention policy, subject to the law’s terms and exceptions.
A vendor’s public audit page cannot deliver applicant-specific timing. Nor does a sentence in a general privacy policy necessarily name the characteristics assessed for a particular requisition. In the example workflow, candidates may need notice before the resume model screens them, not when they reach the recorded interview days later.
Location complicates the control. Hiring teams often use one configuration across offices and remote roles, while notice duties can turn on where a job or candidate falls within a jurisdiction’s coverage. Other state and local laws may impose consent, disclosure, accommodation, or recordkeeping duties without requiring the same New York City bias-audit format.
These are enacted controls only where the relevant law is in force and covers the deployment. Pending algorithmic-accountability bills, voluntary frameworks, and vendor commitments can guide system design, but they are not interchangeable with an enforced audit or notice requirement. Teams should verify current effective dates and agency guidance rather than converting a proposal into a compliance checkbox.
Build a deployment-to-audit crosswalk
Before approving the customer-support workflow, place the audit summary beside the requisition configuration and record six fields: the exact model or component tested; its configured output and threshold; the employment category represented; the selection stage measured; the dates and source of the audit data; and the jurisdictions in which the workflow will run.
Then capture the operational evidence. Save the candidate notice, its delivery timestamp, the characteristics disclosed, the accommodation route, and the version of the system active when each decision occurred. This adds administrative work and may require separate audits when job categories or configurations differ, but it creates a traceable record instead of asking one vendor PDF to support claims it never made.
Material changes deserve another comparison. Switching on the recorded-interview scorer, changing a cutoff, replacing historical data with a new applicant pool, or moving the score from advisory use to automatic rejection can break the match between report and deployment even when the product’s brand name stays the same.
The practical decision is narrower than “Has this AI been audited?” It is whether the available audit examined this component, for this employment category, at this stage, under a use close enough to the employer’s configuration that the report remains relevant evidence. This analysis describes compliance mechanics, not legal advice.
Questions people ask
Does a vendor’s bias audit cover every employer using its tool?
No. A vendor audit may use pooled historical data or test data and may cover only particular components, job categories, or decision stages. An employer should compare the audited configuration and population with its own deployment before relying on the report.
Does passing a bias audit prove an AI hiring tool is lawful?
No. An impact-ratio calculation can identify group differences, but it does not establish job-relatedness, accessibility, or compliance with every anti-discrimination rule. Employers remain responsible for how the tool operates in their selection procedure, including decisions made by vendors.
Can human review keep an AI hiring system outside an audit rule?
Not necessarily. Under New York City’s rules, a tool can substantially assist a decision when its output receives more weight than other criteria or overrides a human conclusion. The authority given to the score matters more than the presence of a recruiter somewhere in the workflow.
Is publishing the audit enough to satisfy candidate notice rules?
No. New York City separately requires advance notice that a covered tool will be used and disclosure of the job qualifications and characteristics it will assess. The employer must manage timing and delivery for the relevant candidates; a vendor’s public audit page does not do that work.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



