Most conversations about AI bias stop at “the training data wasn't representative.” That is one real route bias takes into a system — but a widely cited framework names two others that have nothing to do with the dataset at all, and a system can pass every demographic-balance check and still not be fair.
Key takeaways
The clearest available breakdown of how bias actually enters an AI system comes from the U.S. National Institute of Standards and Technology (NIST) — a United States federal agency, cited here for its framework rather than as a Canadian authority, alongside the Canadian regulatory source below. NIST's AI Risk Management Framework states that it has “identified three major categories of AI bias to be considered and managed: systemic, computational and statistical, and human-cognitive. Each of these can occur in the absence of prejudice, partiality, or discriminatory intent” (NIST AI RMF, AI Risks and Trustworthiness — a U.S. framework). That last sentence matters on its own: none of the three categories requires anyone to have intended a biased outcome.
(All three definitions: NIST AI RMF, AI Risks and Trustworthiness.) The practical consequence is that a bias audit focused only on the training dataset addresses one of three named categories. A system can be carefully rebalanced on the data side and still carry systemic bias from how it was commissioned and measured, or human-cognitive bias in how the people relying on its output interpret an ambiguous result.
NIST draws a distinction worth stating plainly, because it cuts against an intuitive assumption: “systems in which harmful biases are mitigated are not necessarily fair. For example, systems in which predictions are somewhat balanced across demographic groups may still be inaccessible to individuals with disabilities or affected by the digital divide or may exacerbate existing disparities or systemic biases. Bias is broader than demographic balance and data representativeness” (NIST AI RMF, AI Risks and Trustworthiness). A model that produces demographically even outputs on a standard fairness test can still fail people it was never tested against — which is one reason accessibility obligations and bias mitigation are related but separate questions, not the same checkbox.
Canada's Privacy Commissioner puts a concrete obligation on developers that maps onto the computational-and-statistical category above. Under the Fairness discussion in its generative-AI principles, developers should evaluate “the training data sets to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases” (OPC, Principles for responsible, trustworthy and privacy-protective generative AI). The same guidance connects this directly to high-stakes use: unaddressed bias is flagged as most consequential “where they are used as part of an administrative decision-making process…or in highly impactful contexts such as health care, employment, education, policing, immigration, criminal justice, housing or access to finance.”
Canada's Cyber Centre offers a mechanism for how the computational category specifically tends to happen in practice with large language models: “most of the training datasets fed into LLMs come from the open Internet. As such, generated content has a fundamental bias in that only limited amounts of the world's total data are online and available for AI to use” (ITSAP.00.041). Whatever isn't well represented on the open internet in the first place has no realistic path into the training data at all — a structural gap, not a fixable oversight in how a specific dataset was assembled.
A business builds an AI tool to triage incoming customer complaints by urgency. If the historical complaint data used to build it came predominantly from customers who complain in a particular channel or writing style, the model may learn to associate urgency with that style rather than with the substance of the complaint — computational and statistical bias, from a non-representative sample. Separately, if the team defines “successfully triaged” as “matches what a human agent historically did,” and that historical pattern itself reflected uneven attention across customer groups, the model inherits systemic bias regardless of how balanced its training sample is. And if the staff reviewing flagged edge cases give more weight to the model's confident-sounding output than their own judgment on ambiguous cases, that is human-cognitive bias, occurring entirely after training is finished. Fixing only the first of the three leaves the other two untouched.
Not necessarily. NIST is explicit that demographically balanced predictions can still be inaccessible to people with disabilities, still be affected by the digital divide, or still exacerbate other disparities — bias is broader than dataset balance alone.
Yes — NIST states this explicitly for all three of its named categories: systemic, computational and statistical, and human-cognitive bias “can occur in the absence of prejudice, partiality, or discriminatory intent.” Intent is not a precondition for any of them.
No. NIST describes it as present across the entire AI lifecycle, including “design, implementation, operation, and maintenance” — meaning it affects how staff interpret and act on a system's output long after training is complete.
Whether a target's AI system was tested for all three categories of bias, not only dataset balance, is a real diligence question with a factual answer.