Reinforcement learning is a way of training an AI system by consequence rather than by example: instead of studying a set of already-correct answers, the system takes actions inside an environment, receives a reward signal that scores those actions, and gradually adjusts its behaviour toward whatever earns it more reward over repeated attempts.
The United States’ National Institute of Standards and Technology — labelled here as a US, not Canadian, body — defines the mechanism precisely in its adversarial-machine-learning taxonomy: reinforcement learning is the paradigm “in which an agent interacts with an environment and learns an optimal policy to maximize its reward”.
No Canadian statute or regulator defines reinforcement learning, or treats it differently from any other way of building a model. The Treasury Board’s Directive on Automated Decision-Making — which binds federal government departments, not private Canadian businesses — applies to “any automated decision system in production used to make an administrative decision or a related assessment about a client”, and that test does not ask which training method produced the system.
A system trained by reinforcement learning that assesses a benefit claim is covered by that directive on exactly the same terms as one built from a simple rules engine — the directive grades the decision’s impact on the person affected, not the technique that produced it.
The clearest consumer-facing use of reinforcement learning today sits behind the chat assistants millions of people already use. NIST’s own taxonomy names reinforcement learning from human feedback as one documented mitigation approach: a model generates several candidate replies, human reviewers rank which reply they prefer, and that preference ranking becomes the reward signal the model is tuned against in further training rounds.
That is a genuinely different training loop from feeding the model a fixed set of correct answers: nobody writes out the ideal reply in advance. The model discovers, through repeated reward signals, which kinds of replies tend to be preferred — which is also why the same NIST document flags reward signals themselves as a target: an attacker who can influence the ranking step can nudge the model’s behaviour without ever touching its training data directly.
See also: supervised learning, unsupervised learning, machine learning.
Deciding whether a build genuinely needs a reward-driven loop, or whether a simpler supervised approach solves the same problem for less effort, is a scoping question — custom-ai-solutions covers how that choice actually gets made.