Treadstone Associates
Article · 8 min read

How AI systems get sorted by risk

“High-risk AI” sounds like a single, portable label. In practice it means three different things depending on which document assigned it, because the three sourced methods for sorting AI by risk do not work the same way at all.

Treadstone Associates · Updated 2026

Key takeaways

  • • Canada’s federal government sorts its own automated decision systems into four impact levels (I–IV), self-assessed against context, the decision’s reversibility, and data sensitivity — and it applies only to federal departments.
  • • The European Union’s AI Act does not self-assess at all — it lists specific banned practices and specific high-risk use categories directly in the law itself.
  • • The US-based NIST framework is not a risk tier system in the first place — it is a continuous four-function process (GOVERN, MAP, MEASURE, MANAGE) with no output level.
  • • Confusing these three produces real mistakes — assuming a Canadian federal “Level III” system would automatically be “high-risk” under the EU’s law, or assuming NIST issues a rating at all.

Three organizations, three different jobs, three different ways of answering how risky a given AI system actually is. None of them is wrong. They are not measuring the same thing, and treating one method’s output as if it were another’s is the fastest way to misjudge what a rating actually tells you.

Canada’s federal method: a self-assessed scale, scoped to government

The Treasury Board’s Directive on Automated Decision-Making sets out four Impact Assessment Levels, and the assessment is done by the federal department deploying the system, using its own algorithmic impact assessment tool. Level I describes a context with “low levels of risk,” where the decision will likely have “little to no, easily reversible, and brief impacts” and the data used “likely presents low levels of risk.” The scale climbs from there: Level II is “moderate” risk with “short-term impacts”; Level III is “high” risk with impacts that are “difficult to reverse and potentially ongoing”; Level IV is the top of the scale, where impacts are “very high, irreversible and perpetual” (TBS, Directive on Automated Decision-Making, Appendix B). Three factors drive the level at every tier: the context the system operates in, how reversible and how lasting the decision’s impact would be, and how sensitive the underlying data is. The directive applies only to systems in production “used to make an administrative decision or a related assessment about a client” inside the federal government — it says nothing about a system a private business runs, and nothing about provincial or municipal government systems.

The EU’s method: no self-assessment, a fixed list in the law itself

The European Union’s AI Act works on an entirely different logic — it does not ask an organization to score its own system against a scale. It defines four categories directly in the regulation and lists specific practices inside each one. “Unacceptable risk” systems are banned outright, and the list is concrete: harmful AI-based manipulation, social scoring, untargeted scraping of the internet or CCTV material to build facial recognition databases, and AI systems generating non-consensual intimate imagery are named specifically, not left to a case-by-case judgment. “High-risk” use cases are also named directly — AI safety components in critical infrastructure, CV-sorting software for recruitment, credit-scoring systems, and AI used in the administration of justice all appear on the EU’s own list (European Commission, the AI Act — European Union). A “transparency risk” tier requires disclosure rather than restriction — a chatbot has to make clear it is a machine — and a residual “minimal or no risk” tier, which the Commission says covers “the vast majority of AI systems currently used in the EU,” carries no rules at all. The practical difference from Canada’s method is total: an EU provider does not run an assessment tool and land somewhere on a scale. It checks whether its system’s use case appears on a list, because the categories were fixed by the legislature, not calculated per-system. For what this means for a Canadian business specifically, see what the EU AI Act means for Canadians.

NIST’s method: not a tier system at all

The United States’ NIST AI Risk Management Framework is the one most likely to be mistaken for a rating scheme, and it is not one. It has no levels and produces no score. It structures a continuous process around four functions: GOVERN (“legal and regulatory requirements involving AI are understood, managed, and documented”), MAP (“context is established and understood”), MEASURE (evaluating what “will not — or cannot — be measured,” and documenting that gap), and MANAGE (deciding “whether the AI system achieves its intended purposes… and whether its development or deployment should proceed”) (NIST, AI RMF Core — United States). There is no Level I through IV, and no unacceptable/high/minimal categories — an organization following NIST’s framework is not producing a risk tier that could be compared against Canada’s or the EU’s at all. It is documenting a process. The Office of the Privacy Commissioner of Canada’s own generative-AI principles cite this same NIST page — specifically its treatment of validity and reliability — as the source behind its own instruction to “evaluate the validity and reliability of the generative AI tool for the intended purpose” (OPC, generative AI principles), which is the one place a Canadian regulator formally borrows from the American framework — not its structure, just its treatment of one characteristic.

Why mixing these up produces real mistakes

Picture a Canadian government contractor whose system is assessed at Level III under the federal directive because a wrong administrative decision would be hard to reverse. It would be a mistake to conclude the same system is therefore “high-risk” under the EU’s Act — that status turns on whether the use case appears on the EU’s own list (recruitment, credit scoring, critical infrastructure, and so on), not on a Canadian department’s self-scored reversibility judgment. It is an equal mistake to assume a vendor “followed NIST” and therefore know what risk tier its product sits at — NIST does not assign one. Each of the three answers a genuinely different question: how bad would getting this wrong be for the person affected (Canada’s method), does this use case appear on a fixed legislative list (the EU’s method), or has the organization run a sound process at all (NIST’s method).

The level is not only descriptive inside the directive — it changes a concrete operational fact about the system itself. Appendix C ties the two lowest levels to language permitting the system to decide on its own: “the system may make decisions and assessments without direct human involvement.” At Level III and IV, that flips: “the final decision must be made by a human,” and “decisions cannot be made without having clearly defined human involvement during the decision-making process.” (TBS, Directive on Automated Decision-Making, Appendix C) Nothing comparable to that binary sits inside the EU’s fixed categories or NIST’s four functions — it is a consequence unique to how Canada’s own scale actually works.

Related: Ottawa’s own rules for automated decisions and how public bodies in Canada disclose AI use.

Common questions

Is Canada’s Level IV the same as the EU’s “unacceptable risk”?

No. Level IV describes the worst end of a self-assessed federal-government scale — it can still be built and used, subject to the directive’s requirements. The EU’s “unacceptable risk” category is banned outright by name; the two are not comparable positions on one shared scale.

Does NIST assign a risk score to an AI system?

No. NIST’s AI Risk Management Framework structures a process (GOVERN, MAP, MEASURE, MANAGE); it does not output a level or score, so there is nothing to compare against Canada’s four levels or the EU’s four categories.

Does a private Canadian business have to use the federal government’s four levels?

No. The Treasury Board directive and its Appendix B levels apply to federal government automated decision systems only — a private business has no legal obligation to use that scale, though nothing stops it from adapting the method internally.

Trying to work out how risky your own AI use actually is?

A short conversation can walk through which framework, if any, actually applies to what you’re building.