What exactly does NYC Local Law 144 require, and who is covered?
LL 144 (enforced by the Department of Consumer and Worker Protection since July 5, 2023) covers automated employment decision tools, computational processes derived from machine learning, statistical modeling, data analytics, or AI that issue simplified outputs (scores, classifications, rankings) used to substantially assist or replace discretionary decisions in hiring or promotion, when used for candidates or employees in New York City. Three obligations. Bias audit: an impartial evaluation by an independent auditor (not the employer or the vendor with a stake in the tool) completed within one year before each use, calculating at minimum: for selection tools, selection rates and impact ratios by sex, race/ethnicity, and intersectional categories; for scoring tools, median scores and the rate of scores above the median (scoring rate) with the same demographic breakdowns; using historical data of the tool's use (or test data where historical data is insufficient, with disclosure). Publication: a public summary of the most recent audit's results, the distribution date of the tool, and the data sources, posted on the employment section of the employer's website before use. Notice: NYC candidates get at least 10 business days' notice that an AEDT will be used, the job qualifications and characteristics it assesses, and information about requesting an alternative selection process or accommodation. Penalties run $500 for a first violation and up to $1,500 for each subsequent one, per tool per day, and each day of use without a compliant audit or notice is a separate violation, small numbers that compound quickly across requisitions.
What statistics belong in a bias audit, and what is the four-fifths rule's actual role?
The workhorse metrics: selection rate per demographic group (candidates advanced or selected divided by candidates assessed), impact ratio (each group's rate divided by the most-favored group's rate), and for scoring tools, median-score comparisons and scoring-rate ratios, all computed by sex, race/ethnicity, and intersectionally (Black women, Hispanic men), because intersectional disparities routinely hide inside clean marginal numbers. The four-fifths rule, from the EEOC's Uniform Guidelines on Employee Selection Procedures (1978), treats a selection rate under 80% of the most-favored group's as evidence of adverse impact; it is an enforcement heuristic, not a statutory line, and LL 144 deliberately requires publishing impact ratios without setting a pass threshold. Serious audits go past the heuristic: statistical significance testing (two-proportion z-tests, Fisher's exact for small cells) because an 0.75 ratio on 40 candidates and on 40,000 mean different things; confidence intervals on the ratios; and sample-size disclosure per cell, with small-cell suppression handled transparently. Beyond selection outcomes, mature methodology examines score distributions (calibration by group: do equal scores mean equal outcomes across groups), feature diagnostics (proxies for protected classes, ZIP code, school, employment gaps), and error asymmetry (false-negative rates by group for classification tools). The methodological choice with the biggest effect on results is population definition: applicant pool versus assessed pool versus pipeline stage, and honest audits state it, hold it constant across cycles, and resist the temptation to shop populations until the ratios look right.
Beyond NYC, what laws require or reward bias testing?
A widening set. Colorado's AI Act (effective June 30, 2026): deployers of high-risk AI in consequential decisions owe impact assessments analyzing algorithmic-discrimination risk before deployment, annually, and after substantial modifications, bias testing is the analytical core, and the named frameworks (NIST AI RMF, ISO 42001) both embed fairness evaluation. The EU AI Act: high-risk providers owe data-governance examination of possible biases in training, validation, and test data (Article 10), accuracy metrics across the system's context of use, and post-market monitoring; employment and credit systems are squarely Annex III high-risk on the August 2026 timeline. Illinois: the AI Video Interview Act requires notice and consent for AI-analyzed interviews with demographic reporting for some users, and 2026 amendments to the Illinois Human Rights Act make discriminatory AI use in employment decisions actionable. Federal law needs no AI-specific statute: Title VII disparate-impact doctrine, the ADEA, the ADA (screen-out of candidates with disabilities), the FCRA where third-party scores feed decisions, and ECOA/Regulation B in credit all apply to algorithmic decisions exactly as to human ones, and the EEOC has litigated AI-driven screening (its 2023 settlement with iTutorGroup over automated age screening was the first). Insurance regulators (Colorado's SB 21-169 quantitative testing regime for insurers, NY DFS circulars) push the same direction in their sector. The strategic read: jurisdiction-specific audit mandates vary, but disparate-impact liability is universal, so the testing program is the constant and the reporting artifacts are the per-jurisdiction variables.
Who can perform an audit, and what makes one independent and credible?
LL 144 defines independence functionally: the auditor exercises objective and impartial judgment, was not involved in using, developing, or distributing the tool, has no employment relationship with the employer or vendor beyond the audit, and no direct or material indirect financial interest in either. In practice the market spans specialized algorithmic-audit firms, accounting and consulting practices building AI-assurance lines, and academic groups; vendors' self-published fairness reports, whatever their quality, do not satisfy the independence requirement for the deployer's audit. Credibility markers beyond the legal minimum: a written methodology fixed before data arrives (population definitions, category mappings, statistical tests, small-cell handling), so results cannot quietly reshape the method; access to record-level data rather than vendor-summarized tables (auditors who only re-total the vendor's spreadsheet are re-publishing, not auditing); demographic-data handling that is defensible (self-identified data preferred, imputation methods like BISG disclosed with their error characteristics if used); intersectional reporting even where marginal numbers look clean; documented limitations (coverage, missing-data rates, inference caveats) in the published summary; and no contingency between the fee and the findings. For the employer, the procurement questions that separate real auditors: show a redacted prior methodology, explain your small-cell policy, describe a case where you reported adverse findings, and confirm you will not accept the engagement without record-level access. A cheap audit that would never find anything is a liability generator with a certificate on top.
The audit found disparities. What now, and how does remediation interact with legal risk?
First, orient: a disparity is not automatically illegal, and publishing it is not an LL 144 violation, disparate-impact law asks whether the practice causing it is job-related and consistent with business necessity, and whether a less discriminatory alternative exists that serves the same interest. That legal frame dictates the response sequence. Diagnose: trace the disparity to its mechanism, training-data composition, proxy features, threshold placement, interaction with a particular pipeline stage, because remediation targets mechanisms, not ratios. Validate: if you will defend the tool, the validation evidence (criterion validity: does the score predict actual job performance) must exist and be current; a tool with adverse impact and no validation study is indefensible, and this is where many vendor tools quietly fail. Remediate: options ordered by robustness, retrain with corrected data composition; remove or transform proxy features; adjust thresholds or scoring bands (within the bounds of the Ricci line on ex-post race-conscious changes, counsel is mandatory here); constrain the tool to a narrower role (advisory rather than screening); or replace it, and re-audit after any change, because remediation is itself a modification. Document deliberately and under privilege where appropriate: the audit is discoverable and, once published under LL 144, public; the remediation record is what converts 'they knew' from an accusation into evidence of reasonable care, the same duty-of-care posture Colorado's statute rewards. Structurally: pre-commit to a remediation budget and decision path before commissioning any audit, an organization that tests without capacity to act on findings has purchased a plaintiff's exhibit; one that tests, fixes, re-tests, and monitors has built the defensible program every regulator's framework describes.