
Legal AI is priced for perfection: your firm’s data is not
Keith Fenner, chief revenue officer at Morae, highlights why robust information management is key for law firms in the AI age as recent AI legislation introduces stringent data governance requirements
Record valuations assume legal AI will deploy flawlessly. Inside most law firms, the data it depends on is anything but flawless and for SME firms, that gap is now the difference between AI that pays back and AI that stalls.
Capital is betting on flawless execution
In the first half of 2026, capital moved into legal AI at a pace the category has never seen. Record funding rounds priced the sector’s most-backed platforms in the billions — legal technology has stopped being a niche market and become one of the fastest-repricing categories in enterprise software, priced not on what it earns today, but on near-flawless execution of everything it promises next.
I am not here to argue those numbers are wrong. Capital moves ahead of certainty, that is its job.
But every valuation is a claim about the future, and these valuations make a specific one: that legal AI can be deployed quickly, successfully and at scale inside regulated organisations. That claim rests on an assumption the funding narrative rarely examines, and it is the assumption that will decide whether your firm’s AI investment pays back.
What is the real constraint on legal AI?
The market is pricing legal AI as if model performance is the main risk. It is not. From the operating seat at Morae, working inside law firms and regulated organisations every day, the binding constraint on legal AI value is not model capability, that is improving quarterly and converging across vendors. The binding constraint is the state of the data estate the AI is pointed at.
The scale of the problem is documented. The International Data Corporation (IDC) estimates that around 90% of enterprise data is unstructured and largely inaccessible to AI systems. Gartner reported in 2025 that 63% of organisations either lack, or are unsure they have, the data management practices AI requires and predicted that through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data. A Harvard Business Review Analytic Services survey reported in April 2026 found only 15% of organisations currently possess the data foundation for agentic AI.
In law firms the problem is sharper than almost anywhere else, because the data estate is not merely messy. It is dark: decades of accumulated matter files, email archives, legacy document management systems, file shares and departed-fee-earner repositories that are unclassified, unowned, duplicated and contradictory. Content nobody has mapped, privilege nobody has traced, personal data nobody has minimised, and retention obligations nobody is enforcing. SME firms are not exempt from this: they simply have fewer people to absorb the consequences.
What happens when AI meets ungoverned data?
Point a large language model (LLM) at that estate and three things happen, in order. First, the outputs are unreliable, because the model faithfully synthesises the contradictions and obsolete versions it was fed. Second, trust collapses and in legal, trust collapses once. A partner burned by one confidently wrong answer will not return for a second. Third, the programme stalls, the investment is written down, and the firm concludes the technology failed. The technology did not fail. The data was never ready.
Why compliance just made this unavoidable
There was a period when data readiness could be deferred as an efficiency question. That period ended this month. The bulk of the EU AI Act’s obligations became applicable on 2 August 2026, including data governance requirements for high-risk AI systems: data must be relevant, sufficiently representative and, to the best extent possible, free of errors and complete for the intended purpose. UK firms with EU clients or matters are within its reach, and the direction of travel for UK regulators is the same.
Read that requirement against the reality of the average firm’s data estate and the gap is obvious. You cannot evidence the quality of data you have never mapped. You cannot demonstrate that privileged or personal content was excluded from AI processing if you do not know where it lives. Dark data was always a cost and breach-exposure problem. Under an active AI regulatory regime, it is a compliance issue and regulators, courts and opposing counsel will not give credit for effort alone.
Where should SME firms start?
Information governance is not the compliance tax you pay before the interesting AI work begins. It is the work that makes legal AI usable:
- Discovering and classifying the estate
- Defensibly disposing of redundant, obsolete and trivial content
- Mapping privilege and sensitive data
- Enforcing retention as policy in the systems where the data lives.
That discipline turns a dark, liability-bearing archive into a governed estate an AI system can be trusted and legally permitted to work from.
This is also where the structural gap in the market sits. AI vendors, brilliant as their models are, sell software; the state of your data estate is contractually and practically your problem. Traditional advisory can assess the estate but rarely stays to run the remediation. The work that unlocks the value requires governance expertise, technology and managed execution operating as one motion which is the ground Morae has chosen to hold, because it is where deployment reality lives.
The next three years in legal AI will not be decided by who licenses the best model; strong models will be table stakes. The winners will be the firms that governed their data first. AI readiness is not decided when the model arrives. It is decided years earlier, in the choices firms make about data governance, and in the consequences of the choices they defer.


