Grounds for Concern: Fair Use, Terms of Use and the Grist for AI
Much ink has been spilled about copyright and AI. Almost all of it pools in two places: the works used as model training data, and the works that models generate as potentially and allegedly infringing output 1David Tan, “The Best Things in Life are Not for Free: Copyright and Generative AI Learning” Singapore Law Gazette (Apr 2023).. There is a third pool that gets far less airtime: retrieval-augmented generation. It is the practice of handing a large language model context that grounds its answer 2Cheryl Seah, “Liability for AI-generated Content” Singapore Law Gazette (Mar 2024)., in the hope of mitigating the hallucinations that everyone dreads. But whether the grounding itself can be infringing is not oft-discussed.
First, a quick word on “hallucination”, which is a bit of a misnomer. Veracity was never baked into the primary goal of such models: generation. A model generates the way a bedtime story generates: fluently, plausibly, even entertainingly, with no particular duty to the real world. Asking it to stop hallucinating is like asking a storyteller to stop making things up. Grounding is how we drag generation back towards the truth. We tell the storyteller that they now have constraints, a tie to reality.
As AI agents reach out across the internet to ground what they produce, this article argues the following:
- Fair use will be invoked more, and not just at the margins;
- Terms of use will evolve to reckon with grounding, especially where access to information is the live issue;
- The exceptions for text-and-data mining and computational data analysis – narrow or broad, in whichever jurisdiction will be reshaped in the process.
Let me make it concrete with an example closest to home, and the common law system: legal databases.
Is grounding really so different from a meme?
We already accept fair use as the ordinary currency of social media. Reusing, remixing, riffing on someone else’s clip – a great deal of this is treated transformative and mostly permitted 3David Tan, “‘We Don’t Need Permission to Dance’: Key Features of the New Copyright Act 2021” Singapore Law Gazette (Dec 2021); Copyright Act 2021, ss 190-191.. So it is worth asking: is there any qualitative difference between an influencer riffing off someone else’s material and a generative model drawing on primary sources to ground a response, a riff of sorts?
The next question is whether grounded generation is preferable to the ungrounded kind? In a post-truth world, obviously yes. A model generating an answer is doing something uncomfortably close to a social-media influencer putting knowledge of a ”hidden gem” out into the world: it might point you to the real thing (the law), or just as easily invent it (is it really a gem?). The difference, arguably, is human authorship on one side and, at least arguably, none on the other 4David Tan, “AI and Copyright: Death of the Author?” Singapore Law Gazette (Nov 2022).. But that does nothing to remove the need to be grounded. If anything, it sharpens it.
Here is the rub: if fair use is available and defensible, generative AI leans on it at a scale no influencer could dream of. Which is exactly why terms of use matter.
Public, private and the question of lawful access
If we agree that both influencers and AI ought to be grounded in the law, the next question is where they go to get it. Primary, public sources are ideally easy to reach. However, as every practitioner knows, they are not always easy to reach, and might even be arcane in language. But let’s for a second assume primary sources are the easy destination. They usually sit in databases.
First, let’s set private databases aside: a different beast. Unlawful access to one happens through hacking, a leak or plain carelessness, none of which has much to do with grounding or fair use; exfiltrate from a private database and the wrongdoing is rarely in doubt.
The genuinely vexed case – the one this article is about – is the public database, usually governed not by statute or precedent, but largely by terms of use 5Thomson Reuters Enterprise Centre GmbH v Ross Intelligence Inc (D Del, 11 Feb 2025, on appeal to the 3rd Cir); Getty Images (US) Inc v Stability AI Ltd [2025] EWHC 2863 (Ch)..
A brief detour: what actually protects a database?
Surprisingly little. Copyright never protects facts or data – only an original selection or arrangement of them – and ”sweat of the brow” is almost certainly dead 6Feist Publications, Inc v Rural Telephone Service Co 499 US 340 (1991); Global Yellow Pages Ltd v Promedia Directories Pte Ltd [2017] 2 SLR 185; CCH Canadian Ltd v Law Society of Upper Canada 2004 SCC 13; IceTV Pty Ltd v Nine Network Australia Pty Ltd [2009] HCA 14; Telstra Corp Ltd v Phone Directories Co Pty Ltd [2010] FCAFC 149.. In other words, labour doesn’t necessarily result in copyright. A database earns only “thin” copyright, if any. The EU and UK bolt on a second tool, the sui generis database right, which generally protects investment rather than creativity 7Directive 96/9/EC, Art 7; Football Dataco Ltd v Yahoo! UK Ltd (C-604/10); British Horseracing Board Ltd v William Hill Organization Ltd (C-203/02). – but Singapore and the United States have none.
Investment is sometimes difficult to separate from creativity. There are usually two versions of any judgment: the bare, neutral-citation version the court hands down, and the value-added law-report version (headnotes, catchwords, editorial apparatus) over which a publisher has a far stronger copyright claim, and in which creation there is a fair amount of creativity involved. For grounding, one would intuit that the neutral version is the appropriate version to utilise, to avoid clear-cut infringement. From this perspective, ownership is almost ancillary to access: Singapore Statutes Online 8www.sso.agc.gov.sg expressly asserts copyright over Singapore legislation, yet the practitioner’s real problem is rarely “who owns this”. It is “can I reach it, and may I compute on it?”.
The other lock on the door: personal data
Copyright is not the only thing standing between a model and a judgment. Judgments are stuffed with personal data: names, financial affairs, medical histories, the occasional spent conviction. Post-publication redaction is often a concern. From this perspective, data privacy, beyond statute, is a real issue 9Personal Data Protection Act 2012, First Schedule, Part 1, para 1 (publicly available data).. This is, in truth, the harder objection. Some of the anxiety about letting machines loose on case law is not really about who owns the text; it is about anonymisation, and whether sensitive facts about real people should be vacuumed up, recombined and regurgitated by a model that has no sense of discretion. The custodians who worry about this tend to treat anonymisation as something to be fixed, protected, and maintained at source, rather than be allowed to cascade.
I do not propose to solve that here, but grounding has a similar rationale: a model made to cite an actual, properly published source is far easier to audit, and correct, than one left to confabulate from the ether. The personal data question is real, but it is a reason to design computational access carefully, not to wall it off entirely. One could even mandate, to use the technical terminology, diff checks 10https://en.wikipedia.org/wiki/Diff, which are tools that identify changes between versions, and the use of latest versions.
Mechanics: why unlawful access collapses the exception
This is where the statutory plumbing bites. Section 244 of the Copyright Act 2021 (“the computational data analysis exception) 11https://sso.agc.gov.sg/Act/CA2021?ProvIds=P15-#pr244- provides that a site should be mined in accordance with its terms of use. Breach it, and the computational data analysis exception falls apart 12In this author’s opinion, the policy choices in section 244(2)(d) must have been difficult ones, but the final choice was still an overly broad one.. The same fault line runs through other regimes 13Copyright, Designs and Patents Act 1988 (UK), s 29A; UK IPO, “Copyright and AI” consultation (Dec 2024).. Pull the terms of use of legal databases around the world and you will find them strikingly similar, if not identical 14The National Archives, Open Justice Licence v2.0, https://caselaw.nationalarchives.gov.uk/open-justice-licence; BAILII, Terms of Service, https://www.bailii.org/bailii/copyright.html; LawNet, https://www.lawnet.sg., due to their being rooted in a move towards open legal information in decades past 15Declaration on Free Access to Law (Montreal, 2002), https://www.ittig.cnr.it/LawViaTheInternet/index002f.html?option=com_content&view=article&id=58&Itemid=71; Graham Greenleaf, “Free Access to Legal Information, LIIs, and the Free Access to Law Movement” [2011] UNSWLRS 40..
The obvious retort is that there is already a lawful path: subscribe, or seek permission. Payment narrows access to those who can afford it; permission, in most bureaucracies, can take months, and runs counter to the promise of using technology to make legal information more reachable. It’s even worse when multiple permissions are required from the various organisations that host primary, public sources 16https://graemejohnston.substack.com/p/quasi-regulation-of-the-use-of-uk. Another usual justification for keeping the gate shut is currency (leaving aside safeguarding bandwidth for now, an issue also ameliorated by modern cloud computing), the worry that a scraper freezes a judgment at a moment which may grow stale with the passage of time (and also related to the anonymisation concerns). That is, with respect, now a solved problem: an API serving the current version, or a simple timestamp and version flag, fixes it far more cleanly than a blanket ban ever could. And since much of this material is meant to be open to the public, the case for revisiting these terms, now with the added force of legal information needing also to be available computationally to mitigate the threat of confabulations, is only stronger.
Change the terms, or broaden the exception?
So which lever do we pull: rewrite the terms of use, or widen the statutory exceptions? Legislation should not ride roughshod over the private bargains people strike or offer. But once the terms are updated, the exceptions become far more useful 17https://lawgazette.com.sg/feature/copyright-fair-use-in-the-face-of-technological-developments-staying-ahead-or-limping-behind-part-2/ – the real bottleneck will arguably be the private terms of use, not statute. In that sense, modernising terms of use is necessary, and statute is only facilitative.
If you want to see “evolving terms of use” in the concrete, it is worth considering how the EU has already wired the opt-out into its mining exception. The right-holder keeps a veto, but only if it is exercised in a machine-readable way:
The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online. 18https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng
That is the shape of the future: not a blunt yes-or-no to computation, but a structured reservation a model can actually read and obey: this was the goal of robots.txt. This functions like the opposite of the U-turn in Singapore: instead of only proceeding when told you may, you may unless told otherwise.
Despite the apparent cultural clash embodied by the U-turn, this squares with where our own regulators have landed. Both the IMDA Model AI Governance Framework 19https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf for Generative AI, and its more recent counterpart for Legal Responsibility for Agentic AI 20https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/agents-legal-responsibility.pdf stress mitigating hallucination risk with RAG, choosing the right actions available to agents and taking responsibility for actions taken by agents. If that is the goal, then access to primary sources that is commensurate with the scale of computational analysis now possible is essential to achieving those aims. None of which makes server load, currency, or anonymisation unimportant; they simply must be weighed against the value of grounded generation, the balance tipping hardest where the information is public and central to rule of law. In order to comply, one must first be able to able to access that information. All the better if it is easy to access.
A loop worth closing
One contradiction remains. Everything above defends the input side, the grounding. But it cuts the other way on output. The surest way to make a model quote a source verbatim is to hand it that source, and tell it to quote verbatim. Better grounding can mean more, not less, reproduction of the original expression. This is an unresolved tension in many governance frameworks, which ask us in one breath to reduce hallucination by grounding, and in the next to respect content provenance and the rights of creators. The New York Times suit against OpenAI, for example, cites examples of ChatGPT output that cites the Times verbatim 21https://hls.harvard.edu/today/does-chatgpt-violate-new-york-times-copyrights/. I will not pretend that it resolves neatly. The cleanest discipline, for now, is the oldest one: when the machine reproduces a primary source, make sure it is acknowledged, attributed, and in quotation marks.
That still leaves a genuinely open question: can you invoke fair use even where you are in breach of a database’s terms of use, reaching past the statutory exceptions to the general fair-use provision 22Copyright Act 2021, s 184 (fair use operates independently of the computational data analysis exception)., when such genuinely ‘public’ information is in question? And note the obvious limit even if it does: a clean fair-use answer to a copyright claim does nothing for a breach-of-contract claim, or for liability under the Computer Misuse Act 1993 23Computer Misuse Act 1993, s 3 (unauthorised access)..
Conclusion
Information generated in today’s quantities has to be anchored to something. Rules like the hearsay rule exist for a reason: report the report of a report and you are playing broken telephone; it ends in falsehood. Generative AI sometimes gets you to the same lack of veracity without much effort, or actual reporting. Going back to the source will always matter.
If we take seriously the idea of law as a tool of social coordination that enables human flourishing 24Finnis, John. Natural Law and Natural Rights. 2nd ed. Oxford: Oxford University Press, 2011. ISBN: 9780199599141., a question follows that the profession will have to answer: do the providers of primary, public information carry some duty or responsibility – moral, ethical or otherwise – to make it available properly, computationally, and easily, so that this new tool of positive transformation or havoc in equal measure can stand on ground that tips it towards the former?
Publication note: The author is the Director of Publications of the Singapore Law Gazette. To avoid any conflict of interest, the author took no part in the editorial decision to publish this article, which was independently reviewed and approved by the Publications Committee.
Endnotes
| ↑1 | David Tan, “The Best Things in Life are Not for Free: Copyright and Generative AI Learning” Singapore Law Gazette (Apr 2023). |
|---|---|
| ↑2 | Cheryl Seah, “Liability for AI-generated Content” Singapore Law Gazette (Mar 2024). |
| ↑3 | David Tan, “‘We Don’t Need Permission to Dance’: Key Features of the New Copyright Act 2021” Singapore Law Gazette (Dec 2021); Copyright Act 2021, ss 190-191. |
| ↑4 | David Tan, “AI and Copyright: Death of the Author?” Singapore Law Gazette (Nov 2022). |
| ↑5 | Thomson Reuters Enterprise Centre GmbH v Ross Intelligence Inc (D Del, 11 Feb 2025, on appeal to the 3rd Cir); Getty Images (US) Inc v Stability AI Ltd [2025] EWHC 2863 (Ch). |
| ↑6 | Feist Publications, Inc v Rural Telephone Service Co 499 US 340 (1991); Global Yellow Pages Ltd v Promedia Directories Pte Ltd [2017] 2 SLR 185; CCH Canadian Ltd v Law Society of Upper Canada 2004 SCC 13; IceTV Pty Ltd v Nine Network Australia Pty Ltd [2009] HCA 14; Telstra Corp Ltd v Phone Directories Co Pty Ltd [2010] FCAFC 149. |
| ↑7 | Directive 96/9/EC, Art 7; Football Dataco Ltd v Yahoo! UK Ltd (C-604/10); British Horseracing Board Ltd v William Hill Organization Ltd (C-203/02). |
| ↑8 | www.sso.agc.gov.sg |
| ↑9 | Personal Data Protection Act 2012, First Schedule, Part 1, para 1 (publicly available data). |
| ↑10 | https://en.wikipedia.org/wiki/Diff |
| ↑11 | https://sso.agc.gov.sg/Act/CA2021?ProvIds=P15-#pr244- |
| ↑12 | In this author’s opinion, the policy choices in section 244(2)(d) must have been difficult ones, but the final choice was still an overly broad one. |
| ↑13 | Copyright, Designs and Patents Act 1988 (UK), s 29A; UK IPO, “Copyright and AI” consultation (Dec 2024). |
| ↑14 | The National Archives, Open Justice Licence v2.0, https://caselaw.nationalarchives.gov.uk/open-justice-licence; BAILII, Terms of Service, https://www.bailii.org/bailii/copyright.html; LawNet, https://www.lawnet.sg. |
| ↑15 | Declaration on Free Access to Law (Montreal, 2002), https://www.ittig.cnr.it/LawViaTheInternet/index002f.html?option=com_content&view=article&id=58&Itemid=71; Graham Greenleaf, “Free Access to Legal Information, LIIs, and the Free Access to Law Movement” [2011] UNSWLRS 40. |
| ↑16 | https://graemejohnston.substack.com/p/quasi-regulation-of-the-use-of-uk |
| ↑17 | https://lawgazette.com.sg/feature/copyright-fair-use-in-the-face-of-technological-developments-staying-ahead-or-limping-behind-part-2/ |
| ↑18 | https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng |
| ↑19 | https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf |
| ↑20 | https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/agents-legal-responsibility.pdf |
| ↑21 | https://hls.harvard.edu/today/does-chatgpt-violate-new-york-times-copyrights/ |
| ↑22 | Copyright Act 2021, s 184 (fair use operates independently of the computational data analysis exception). |
| ↑23 | Computer Misuse Act 1993, s 3 (unauthorised access). |
| ↑24 | Finnis, John. Natural Law and Natural Rights. 2nd ed. Oxford: Oxford University Press, 2011. ISBN: 9780199599141. |

