A NYC Hiring Bias Audit Needs More Than a Published PDF
New York City requires annual bias audits for covered hiring tools, but a results page cannot prove the test used the right data, workflow, or software version.
August 9, 2026 · 8 min read

Picture one concrete workflow: a résumé-ranking system scores New York City applicants, and recruiters advance those above a configured threshold. The employer’s website links to a bias-audit PDF showing selection rates and impact ratios, along with the date of the audit.
That PDF may satisfy part of New York City’s publication requirement. It cannot, by itself, establish that the auditor tested the same tool, threshold, applicant population, and hiring stage that recruiters use today.
The evidence that matters sits behind the PDF: the frozen applicant export, the definition of “selected,” the records excluded from analysis, the auditor’s independence documentation, copies of candidate notices, and a change log connecting the audited system to the production system. Without those records, an employer has a result but little support for how it was produced.
The rule requires an audit, not a passing grade
New York City’s Local Law 144 restricts employers and employment agencies from using a covered automated employment decision tool, or AEDT, unless the tool received a bias audit no more than one year before use and a summary of the latest results was made publicly available. Enforcement began in 2023.
An AEDT is a computational process that issues a score, classification, or recommendation and substantially assists or replaces discretionary decision-making in hiring or promotion. The definition matters. A tool does not enter or leave scope merely because its vendor calls it artificial intelligence, and software that only schedules interviews or stores applications may not perform the decision function described by the law.
The audit calculates selection or scoring rates for demographic groups and compares them through impact ratios. A selection rate is the share of people in a group who move forward; an impact ratio compares that rate with the rate for the group receiving the highest selection rate. The rules call for calculations across sex, race or ethnicity, and intersectional categories where the available data supports them.
Local Law 144 does not set an impact-ratio score that makes a tool lawful, nor does publication certify compliance with federal, state, or city discrimination law. The enforced obligations concern a covered tool’s audit, public summary, and notices. An adverse result may warrant investigation or remediation, but the city’s rule does not convert the audit into a pass-fail license.
That distinction should shape the evidence request. Asking only for a vendor’s “compliance certificate” collapses several separate questions: whether the employer’s workflow is covered, whether the right population was tested, whether the calculations can be reproduced, and whether the audited tool is the one in use.
Start with the frozen applicant export
Return to the résumé ranker. Before calculating anything, the auditor needs a defined population and outcome: applicants evaluated during a stated period, by a stated configuration, for stated jobs, with “selected” tied to a specific point such as advancement to recruiter review or interview.
The employer should request and retain a frozen copy of that input dataset, or a controlled extract from which it can be reconstructed. The accompanying data dictionary should identify each field, its source, the demographic categories used, the hiring-stage outcome, and the method used to join demographic records with system decisions. Access can be restricted because the file may contain sensitive employment and demographic data; reproducibility does not require publishing personal records.
Selection-rate calculations can change sharply when the denominator changes. An export might include people who withdrew before scoring, duplicate applications, test accounts, applicants routed around the tool, or records created after a position closed. Conversely, a failed join between the applicant-tracking system and a voluntary demographic survey can remove real candidates from the audit without changing the tool’s output log.
For each exclusion, the evidence packet should show a rule, a count, and a reason. The city’s rules allow certain small categories to be excluded from impact calculations and require the public summary to identify exclusions and explain them; records with unknown demographic information also need visible treatment rather than quiet deletion. The useful artifact is an exclusion table that reconciles the original export to the final analysis population.
The employer should also ask whether the audit used its own historical data, pooled data from multiple employers, or test data. Test data means records constructed to evaluate the tool rather than outcomes from actual candidates. Each source answers a different question: historical data reflects a deployed workflow but can reproduce old recruiting patterns, while test data can cover missing groups yet may not represent the employer’s jobs, thresholds, or applicant behavior.
Recalculate the published rates
A defensible audit workbook should let another qualified reviewer move from group counts to the published summary without guessing. For a selection system, that means the number assessed and the number selected in each reported category, the resulting selection rates, the comparison group used for each impact ratio, and the calculations for intersectional groups.
Scoring tools need equal care. The employer should identify which score or classification the audit tested, particularly when the product emits several outputs or recruiters see a composite ranking assembled from them. Testing a communication score does not cover a final recommendation that also incorporates an assessment result and a knockout rule.
Small samples deserve context, but they should not disappear behind a general statement that the data was insufficient. The auditor’s workpapers should preserve the category counts, the applied exclusion rule, and any sensitivity analysis showing whether a different time window or job grouping materially changes the result. Combining unrelated jobs can increase the sample while hiding that the ranker behaves differently for warehouse applicants and software engineers.
This is where the public PDF reaches its limit. A summary can report rates and exclusions, as the rules require, without exposing enough information to verify the source query, detect duplicate records, or confirm that “selected” means the same thing in the audit and the live recruiting workflow.
Independence needs records, not a label
The city’s rules define an independent auditor through objective and impartial judgment. The auditor cannot have been involved in using, developing, or distributing the AEDT, cannot have an employment relationship with the employer or vendor during the audit, and cannot hold a direct financial interest or a material indirect financial interest in them.
An employer should therefore retain the engagement letter, scope of work, conflict statement, payment arrangement, and a record of who supplied data, wrote the analysis code, reviewed exceptions, and approved the report. Paying an auditor does not by itself defeat independence, but a document that merely calls a vendor partner “independent” does not establish the rule’s criteria either.
The boundary is especially important when a vendor coordinates the audit. The evidence should show that the auditor controlled the analytical judgment and received enough underlying data to test the calculations, rather than signing a report assembled by the company whose product was under review.
Notices need delivery receipts
The audit is only one control. Before using a covered AEDT, an employer or employment agency must give a New York City resident candidate or employee notice at least 10 business days in advance, identify that an AEDT will be used, and state the job qualifications and characteristics it will assess.
For the résumé ranker, the notice record should connect the words candidates saw to the actual workflow. A generic statement that technology may assist recruiting is weak evidence if the system ranks experience, credentials, or other stated characteristics before recruiter review.
Employers can document the notice path with the approved template, publication or delivery timestamps, job-page snapshots, email logs, and the rule used to determine which candidates received it. They should also retain the posted or supplied instructions for requesting an alternative selection process or accommodation. The law requires a route to make that request; it does not itself require the employer to grant an alternative process.
Separate disclosure rules cover the type and source of data collected by the AEDT and the employer’s data-retention policy. A link in the audit summary is not a substitute for evidence that the applicable information was available through the required channel and that written requests were handled within the rule’s timeframe.
A tool change can strand last year’s audit
Finally, match the audit to production. The employer should record the product release, model or ruleset identifier, scoring configuration, threshold, enabled features, job families, and deployment dates covered by the audit. A vendor name and product name are too coarse when the same service supports different models or customer settings.
Suppose the résumé ranker was audited when recruiters reviewed the top band, then the employer lowered the threshold, enabled a new knockout filter, or moved the score earlier in the workflow. The published rates describe the old setup. Even if the vendor considers the product name unchanged, the employer needs a documented assessment of whether the change affects selection and whether a new audit is needed before continued use.
Keep that assessment with release notes, configuration exports, procurement records, and approval tickets. The practical control is a deployment gate: the hiring team cannot activate a consequential model, threshold, or workflow change until the owner compares it with the audited scope and records the decision. That costs staff time and may delay a release, but it is cheaper than trying to reconstruct an obsolete configuration after a complaint.
Questions people ask
Does a published
NYC bias-audit report prove the hiring tool complies?
No. It can demonstrate that required results were made public, but it does not alone prove that the correct workflow, population, outcome, or software configuration was audited. It also does not establish compliance with other employment-discrimination laws.
What data should an employer keep behind the selection rates?
Keep a controlled applicant-level extract, its data dictionary, group and outcome counts, join logic, exclusion table, calculation workbook, and the date range and jobs tested. Sensitive records need access controls, but a reviewer should still be able to reproduce each published rate.
Can a vendor’s audit cover every employer using the same tool?
A vendor-coordinated audit may support reliance in some circumstances, but the employer still needs to compare its deployment with the audited version, configuration, use, and data source. A pooled or test dataset may not represent the employer’s threshold, job mix, or definition of selection.
When should a hiring-tool change trigger another review?
Review any change to the model, rules, threshold, inputs, enabled features, or point in the hiring workflow. The key record is a written comparison between the audited setup and production, completed before the changed system affects candidates.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



