Facts
The Office of the Privacy Commissioner of Canada (OPC), Quebec’s Commission d’accès à l’information (CAI) and the Information and Privacy Commissioners of British Columbia and Alberta jointly investigated OpenAI OpCo, LLC. The investigation examined the collection, use and disclosure of personal information of individuals in Canada to develop and deploy ChatGPT, focusing on the GPT-3.5 and GPT-4 models that powered it when the inquiry began. More than 99 percent of pre-training data had been obtained by crawling publicly accessible sources, and the rest came from licensed datasets. Users’ conversations with ChatGPT were also used for fine-tuning. Sources included social media and discussion forums, which can contain sensitive information and information about children.
Question
The investigation addressed seven questions: whether the information was collected for purposes a reasonable person would consider appropriate and limited to what was necessary; whether valid consent was obtained; whether OpenAI met its openness obligations; whether it took reasonable steps to keep the information ChatGPT generates about individuals accurate; whether individuals could access and correct their information; whether there were proper retention and disposal procedures; and whether OpenAI was accountable for the information under its control. A central dispute was OpenAI’s position that it could rely on implied consent for information scraped from publicly accessible sources.
Decision
In findings of 6 May 2026 the regulators held that OpenAI’s collection of personal information from public websites and licensed sources for training had been overbroad and therefore inappropriate. Individuals would not reasonably have expected information posted about them online to be used for training, and valid consent had not been obtained. OpenAI also fell short on openness, accuracy, access and correction, retention and accountability. After OpenAI retired the older models, introduced a filter masking personal information and made further commitments, the OPC found the complaint well-founded and conditionally resolved. The British Columbia and Alberta commissioners concluded that consent for models built on scraped data had not been and could not be obtained. The CAI’s conclusions varied by issue.
Why it matters
The findings show regulators in one country reaching different results on consent for LLMs trained on scraped data. The OPC, applying what it called a pragmatic and flexible interpretation of PIPEDA, accepted that OpenAI’s commitments would resolve the matter. The two provincial commissioners, applying more specific statutes, held that consent for such models could not be obtained. The report records that OpenAI generally disagreed with the findings. The regulators will monitor its commitments through quarterly reports.
Related stages
OpenAI committed to show, within three months of the report, a notice in the signed-out ChatGPT web experience that chats may be reviewed and used for training, and to make its data exports more accessible within six months.