Exorcising Automagic Legal AI
A Decision-making Framework for Legal AI Use
Readers may be aware about the Reddit and Linkedin firestorm surrounding a particular legal AI company that has the same name as a TV character from the show Suits1Phrased to avoid the implication that the startup was deliberately named after a very good looking but unethical lawyer..
I have emphasised before that legal AI should be available to all, not just lawyers. This is an important point. The legal world tends to talk of legal AI as if only lawyers had a stake in it. The truth is that we should really be examining its effect – potential and realised – on everyone.
TL;DR – this is a checklist to help you – whether you are a lawyer or not – make an informed decision about legal AI tools, and to help2This is deliberate. I do think many vendors need help and encouragement to be honest. vendors remain honest.
When a vendor claims their legal AI has been “trained on legal documents”, what does that actually mean? When they promise “domain-specific” capabilities, what are you really buying? What should you be looking out for?
Here’s a framework to cut through the noise.
This is intentionally comprehensive, rather than too brief and fragmented. Legal AI vendors thrive on partial understanding. When you only grasp pieces, they can claim anything. You will not need every section for every vendor, but having the full map prevents you from getting lost in rhetoric.
What We Talk About When We Talk about Legal AI
When we talk about the efficacy, performance, or usefulness of a particular legal AI product, what are we really talking about?
Benchmarks in this regard are a new frontier. You’ve got benchmarks written by computer scientists – which suggest model capability doubling every seven months – and you’ve got those that apply specifically to the legal domain, like this, this, this, this, and this.
Side-note here. Generally, when you talk about a benchmark or evaluation, you’re going to have a dataset that you’re going to use for the benchmark as well. Which raises several questions:
- What are benchmarks really?
- What do they measure?
- When are general benchmarks useful?
- When are legal-specific ones useful?
- Should we prefer one set to the other, or use them in tandem?
- How should we put together datasets?
- How often should we update these datasets?
- Should these datasets be generally available, or publicly available?
For a really good primer on evaluation and benchmarking, refer to Sebastian Raschka’s excellent Substack, here. The post is technical, but it’s worth Googling, or asking Claude about every term that you’re unsure of.
The next question is what exactly we are benchmarking or evaluating? Here are several levels:
- The ‘raw’ model – which you will rarely encounter in the wild. All the main large language models are extensively programmed with system prompts. Where the prompt engineering ends, and where actual model improvement begins (weight adjustment / training, etc.) is usually opaque to the user, if the model is not open source3If you are a lawyer, and you have no clue what open source means, I strongly suggest acquainting yourself with the concept. It’s been the driving force behind many major technological innovations in the past few decades, and will continue to be so. It’s definitely coming for the legal world, because the legal domain is replete with an insane amount of monopolistic and oligopolistic behaviour. Open source will be one of the directions from which this will be disrupted.. These models are usually trained on such large amounts of data, and using such sophisticated techniques and methodologies, that you are seldom going to get much, if any improvements, from doing further training on your own.
- The model behind an Application Programming Interface: what you get when you directly call the service, instead of through something like the chat bot on ChatGPT, or Gemini. Programmers or savvy lawyers are the most likely ones to be tapping on this directly, and it can be a real boon to your workflow4I personally use a mix of subscriptions to get more done, but use Claude Max plan, Cursor, and Warp as my main go-tos, especially since I use the command line and Markdown a lot, even for legal-related work..
- Particular system prompts that you can define. You have more control over this when interacting directly with the API rather than a user interface. System prompts can change model behaviour in significant ways. For a mix of both approaches, you might want to try things like the Consoles offered by the main LLM providers e.g. Claude Console.
- Workflows: stringing together and combining5One refers to the specific sequence, the other refers to the specific composition of the ‘team’ of models and agents. Sequence is important: whether you brush your teeth before or after you eat cake is very important. And every one knows how important team composition is, without any illustrative example. But if not, think of your worst group project (I hope you were not the cause of it). models and prompts in different ways. Again, orchestrating this well can affect model performance. n8n has many great examples of this, and its community workflows are a great example of how sectors outside of legal have grown exponentially more sophisticated in their self-deployment of AI than legal.
- Agents: giving models some degree of ability to work independently, and decide on which other tools they can call. Like searching the internet, and things like that. Usually also requires some measure of prompt engineering.
- User interface: a “wrapper” encompassing all of these things. Like a candy wrapper, packaging is important, but its substance is what you really care about. Unlike a sweet wrapper, most of our digital and actual lives are mediated by interfaces, so getting it right, and not just looking pretty, is supremely important. But in terms of general legal sector fluency with AI, we are still very much at the level of liking our candy wrappers, and how well they fit in between our main course and end-of-meal choice of drink6Which also means they’ll usually only try something out of that tried and tested flow of courses – suitably marked with prestige – when there’s enough of a critical mass of prestigious folks doing similarly, or if the good/service has become a Veblen good.. There are exceptions, but they prove the norm.
- General domain or domain specific AIs. It is usually going to be a combination of the above. If someone claims that they have a legal AI that’s been trained on legal domains, you should analyse:
- Is this a truly large language model? Or one of the smaller (but still relatively large, especially when compared to history) ones? This claim is more plausible in the latter case, and extremely suspicious when it comes to the former.
- Does the vendor / sales rep really mean fine-tuning instead? Here’s a excerpt from Claude’s documentation on the concept being applied to one of its models.

From Anthropic.
- Ask for the model training / model knowledge cut-off date, and see whether this matches any of the publicly known training cut-off dates for the model. Since I first started in the legaltech space in 2018, and during my consulting, I’ve heard at least one sales rep claim additional training of a truly large language model on legal documents, but the training cut off date for this training was the same as GPT-4. Disingenuous to say the least.
NOTE: the gap between general wrapper providers and legal-specific ones is quickly becoming a hotly-contested one for a very of reasons. I would argue most should keep a tab on the space.
- Use-case specific applications. Some use-cases are specialised and specific enough that they would benefit from a very specific combination of the above. Lawyers have the less ideal distinction of being one of the sectors that claim that every use case is like this, but still can’t help but love the universal chat modality.
- Practitioners and companies using and orchestrating a specific combination of the above. This is self-explanatory. You might place a truly AI-native company, extremely savvy solo practitioner, or big law firm that’s only claiming AI use, rather than actually using it all here.
These are all, to put it very briefly, levels of abstraction. Most stop and start their knowledge of legal AI at six, and some measure of seven. That’s why things seem so automagical and frustrating at the same time. It’s like a client who pays for a legal service and doesn’t want to worry about anything else.
Then there’s the how:
- Are you evaluating how well a model retrieves relevant information – like a search engine – or are you evaluating the entire product, which is usually a collection of retrieval models + generative models? Metrics like precision, recall, and accuracy are very important in this regard.
- Are you evaluating qualitatively, or quantitatively? And of course, it’s entirely possible to give qualitative assessments the look and feel of a quantitative (“Please rate my service on a scale of 1-5”, and attach explanatory notes to each level).
- And this one lawyers will love. We all know the discussion and literature on subjective vs objective tests in law. How does that apply here?
Other crucial things to consider, which have been described extensively elsewhere, so I will not repeat here. Generally, if you won’t go shouting about your client’s information in a lift, then don’t copy-paste it anywhere on the internet:
- Cybersecurity
- Privacy
- Confidentiality
- Likely longevity of the company. Again, intelligent risk-taking at best, a foolish gamble at worst. Remember there are startups, small businesses, and huge enterprises all operating in this area. Most are only talking about the huge startups and huge enterprises, and also assume that scale = longevity7Check out the bubbles and monopolies of history: tulips, dot-coms, and subprime mortgages come to mind.. That’s not necessarily the case8As a small business owner myself, I of course have a soft spot for small business (in the same vein as the family-run food stalls and restaurants that dot cities around the world), and think they have their place. I think too many make the mistake of small means < Microsoft. I think a healthy combination of all scales of companies is essential to the sector being antifragile.The legal sector is historically a very good example of this, but the same reasons (prestige, money, what counts as living a good life), are driving why everyone talks about huge legal companies and huge legal AI, despite the rest of the S&P 500 equivalent all playing hugely important roles in the fabric of society. So I think the legal sector is actually one of those sectors that are beginning to show very strong signs of fragility in this respect. The uneven deployment of legal AI is not helping..
- Pricing strategy or roadmap: is VC / some other source of money subsidising your use until you are locked in? Remember how Uber and Grab used to be so cheap?
- Benchmarking vs evals: are you looking to set your own bar exam for one of the levels of abstraction above, or looking to develop an evaluation framework that allows you to improve those things using a fast feedback loop, on a daily basis?
And finally there’s the all important question:
What are you really looking to do?
If you are looking for one tool to do everything, you are going to get hurt real bad. If you’re confusing knowledge management with contract review, you are also going to get hurt real bad. Or if you would like an agent to also organise your files and folders for you as they retrieve things. It’s a bad idea when asking humans to do it, and the same generally applies to AI: solve the problem of the needle in a haystack, one needle and one haystack at a time.
One of my favourite products is the email client, Superhuman, and they’ve pinpointed what exactly they would like their AI to do in the context of an email client. It would be almost ludicrous to have a Swiss army knife like ChatGPT purports to be operating within the context (“Please find this invoice from Party X … Oh help me find an aglio olio recipe while you’re at it” This is a symptom of ADHD, not a mark of an effective agent).
I hope this will help you in evaluating what you are paying for!
The current legal AI debate – wrapper vs. non-wrapper – usually asks the wrong question at the wrong level of abstraction. And most benchmarking is at the wrapper level: each vendor’s specific combination of what we have discussed above. Whether that’s a prudent approach is a much deeper topic.
Now you know what questions to ask at every level, but also how to get started putting together these pieces for yourself at a scale, cost, and mode that are appropriate to you.
Endnotes
| ↑1 | Phrased to avoid the implication that the startup was deliberately named after a very good looking but unethical lawyer. |
|---|---|
| ↑2 | This is deliberate. I do think many vendors need help and encouragement to be honest. |
| ↑3 | If you are a lawyer, and you have no clue what open source means, I strongly suggest acquainting yourself with the concept. It’s been the driving force behind many major technological innovations in the past few decades, and will continue to be so. It’s definitely coming for the legal world, because the legal domain is replete with an insane amount of monopolistic and oligopolistic behaviour. Open source will be one of the directions from which this will be disrupted. |
| ↑4 | I personally use a mix of subscriptions to get more done, but use Claude Max plan, Cursor, and Warp as my main go-tos, especially since I use the command line and Markdown a lot, even for legal-related work. |
| ↑5 | One refers to the specific sequence, the other refers to the specific composition of the ‘team’ of models and agents. Sequence is important: whether you brush your teeth before or after you eat cake is very important. And every one knows how important team composition is, without any illustrative example. But if not, think of your worst group project (I hope you were not the cause of it). |
| ↑6 | Which also means they’ll usually only try something out of that tried and tested flow of courses – suitably marked with prestige – when there’s enough of a critical mass of prestigious folks doing similarly, or if the good/service has become a Veblen good. |
| ↑7 | Check out the bubbles and monopolies of history: tulips, dot-coms, and subprime mortgages come to mind. |
| ↑8 | As a small business owner myself, I of course have a soft spot for small business (in the same vein as the family-run food stalls and restaurants that dot cities around the world), and think they have their place. I think too many make the mistake of small means < Microsoft. I think a healthy combination of all scales of companies is essential to the sector being antifragile.The legal sector is historically a very good example of this, but the same reasons (prestige, money, what counts as living a good life), are driving why everyone talks about huge legal companies and huge legal AI, despite the rest of the S&P 500 equivalent all playing hugely important roles in the fabric of society. So I think the legal sector is actually one of those sectors that are beginning to show very strong signs of fragility in this respect. The uneven deployment of legal AI is not helping. |

