A predictive model or formula learns from past data, so it inherits whatever that data contains. Past decisions may reflect unequal treatment, uneven access to services or differences in who was watched more closely. A variable that looks neutral, such as a zip code, the number of prior contacts with public systems or a gap in employment, can act as a proxy for race, disability, poverty or national origin. Missing data can be concentrated in some groups, for example when language preference or disability information is recorded inconsistently. And data may simply be wrong. In the Idaho Medicaid case K.W. v. Armstrong, involving a formula that set budgets for adults with developmental disabilities, the ACLU reported that expert review found errors in the historical data so serious that about two-thirds of records had to be discarded, along with regional disparities the state could not explain.
Testing for bias starts with clear questions. Is the tool accurate overall for the decision it supports? Is it similarly accurate across groups defined by race, ethnicity, language, disability, age, sex and region, where data allows? Do false positives (flagging someone who should not be flagged) and false negatives (missing someone who needs attention) fall more heavily on some groups? Does the tool's output lead to different results for groups in practice, once staff act on it? These questions echo the disparate impact analysis used in civil rights law: is there a disproportionate adverse effect, is the tool necessary to meet a legitimate goal, and is there a less discriminatory alternative that would work as well? Legal conclusions belong to counsel and the Equal Opportunity and Access Division, but program leaders should insist that the testing happens.
The National Institute of Standards and Technology (NIST) has published a voluntary risk management framework for these systems. It organizes the work into four functions: govern (set policies, roles and accountability), map (understand the context, purpose and people affected), measure (analyze and track risks, including harmful bias), and manage (prioritize and act on risks). It describes trustworthy systems as valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed. NIST has indicated the framework is being revised, so check the current version.
Testing is not a one-time event. Populations, policies and data systems change. Plan tests before launch, on a regular schedule after launch and whenever the underlying data or rules change, and publish a summary of results in plain language.