Last week we ran Case File 01 of the Audit Files: two mathematicians who fed their unpublished work into OpenAI’s tools, then watched the company announce results that looked strikingly like their own private techniques. OpenAI’s written position was that it could not rule out having trained on their conversations. Nobody outside OpenAI can audit the training run — not the mathematicians, not regulators, not even, it appears, the company itself. That case remains unproven, and may stay that way permanently, because the system has no way to answer the question.
This week, on the other end of the same problem, a law firm quietly showed what the institutional answer looks like: stop asking. Own the machine.
🔍 THE BOTTOM LINE
Latham & Watkins — the US’s second-largest law firm, with US$8.3 billion in revenue last year — has purchased several of its own Nvidia GPU servers and is fine-tuning open-weight models in-house. It is, per the Financial Times, the first public example of a major law firm buying its own AI hardware and customising models, and it sits on a pattern that is quickly becoming the shape of enterprise AI: Capital One fine-tunes open-weight models on its own data, Thomson Reuters spent US$40 million building its own legal model on an open base, and Netflix and Uber run their own inference infrastructure. The driving logic is not that open models beat frontier APIs — often they don’t. It is that a company with decades of confidential client work would rather own its compute, its data and its workflow than rent capability from a provider it cannot audit. If you don’t own it, you can lose it — or at least, you can never prove you haven’t.
What Latham actually bought
The verified facts, as reported by the Financial Times and Bloomberg Law:
- The firm has bought several servers, each holding multiple GPUs — the FT reports this has happened “in recent years,” with Bloomberg Law adding that Latham operates multiple Nvidia H200 GPUs and is evaluating newer Blackwell-generation systems.
- The machines sit in a third-party data centre that only Latham employees can access.
- Latham’s in-house machine learning engineers are fine-tuning Nvidia’s Nemotron 3 open-weight models — adapting open models to the firm’s own needs rather than building a foundation model from scratch.
- It is a hybrid, not an exit: Latham still uses commercial AI services such as Harvey. Chief information officer Rene Mendoza told the FT that some client information is sensitive enough that the firm does not want it with “any cloud vendor.”
So the accurate read is not “law firm abandons cloud AI.” It is workload placement: the firm decided which workloads it needs to control, and bought the hardware to control them.
Latham is not alone
The reason this single firm’s procurement decision matters is who is standing next to it.
Capital One built its multi-agent AI platform around open-weight models fine-tuned on its own proprietary data. “We view our data as a huge advantage and something that nobody else has, something that the general frontier models cannot provide,” the bank’s machine learning engineering lead told VentureBeat. Its customer-facing Chat Concierge runs on a customised Llama model, not an off-the-shelf frontier API.
Thomson Reuters spent about US$40 million building “Thomson,” its own legal model built on Alibaba’s open-weight Qwen, trained on its Westlaw and Practical Law content. The company’s stated reasons read like a procurement doctrine: fine-tuning closed frontier models tends to degrade them, you stay locked into the provider’s inference costs and roadmap, and — in the words of its chief technology officer — building in-house is “renting a house versus buying a house.” Every expert correction becomes training data on a model you own. Renting, that equity evaporates at the provider.
Netflix runs its full LLM serving stack in-house and describes graduating from a hosted model to a fine-tuned self-hosted one as “nearly seamless — for quality, latency, cost, or data privacy.” Uber fine-tunes Llama and Mixtral on its own on-premise GPU clusters, and reports fine-tuned models matching GPT-4-level performance on its own tasks at its own scale.
And on the supply side, Mistral’s Forge platform (launched March 2026) sells managed sovereign fine-tuning on customer-controlled infrastructure — with launch partners including ASML, Ericsson, the European Space Agency and two Singaporean defence agencies. When Europe’s largest industrial companies and Singapore’s defence establishment are the reference customers, “sovereign AI” has stopped being a government word.
The numbers behind the shift
This is anecdote-hunting without data, so here is the data. Deloitte’s State of AI in the Enterprise survey — 3,235 business and IT leaders across 24 countries — found the share of organisations using public cloud as their primary production AI environment fell from 56% to 41% in a single year, while 77% now factor an AI vendor’s country of origin into their selection decisions. Accenture research puts 62% of European organisations actively seeking sovereign AI solutions, higher in banking. In July, the CEOs of Meta, Microsoft, Nvidia, IBM, Dell, Hugging Face, Palantir and others signed an open letter on open weights and American AI leadership — a consensus statement, from the industry’s own leadership, that downloadable, inspectable, ownable models are strategic infrastructure rather than a threat to it.
None of these numbers say enterprises are abandoning OpenAI or Anthropic. They say something more structural: for a growing share of workloads, the default answer to “where does AI run?” is no longer “somewhere else.”
From the chat box to the server rack
Put last week’s case file and this week’s procurement decision side by side and the connection is hard to miss.
The mathematicians’ problem was that they had contributed value to a pipeline they could not see into. Their unpublished techniques went in as chat inputs; results came out; and the question “was our work used?” could not be answered by anyone, including the company that owned the pipeline. Whatever the truth is, the structure of the situation — not any finding of fact — is the scandal: trust was demanded where verification was impossible.
Latham’s server rack is the same problem solved by ownership. When the compute is yours, the models are open-weight and inspectable, the data never leaves your environment, and the only people with access are your employees, the question “is our data being used?” stops being a matter of trust. You can see the answer. You were there, and it isn’t.
That is why the “labs have no moat” framing — popular on X this week — is both overcooked and beside the point. The threat to frontier AI providers is not that open models beat closed ones on benchmarks; on most measures they still trail, and frontier access still wins the tail of hard tasks. The threat is quieter: their largest customers are discovering that they no longer need to trust them for the bulk of their work. A provider can survive losing a benchmark. Losing the assumption that customers must depend on you is a different kind of loss — and it is a decision each customer makes unilaterally, the day owning becomes cheaper than trusting.
The honest counterweight
Two things keep this story from being a victory lap. First, most enterprises that raced to buy GPUs are running them badly: Cast AI’s 2026 analysis of tens of thousands of production clusters found average GPU utilisation around 5% — roughly 20 times more capacity than the workloads use. Self-hosting is usually more expensive than renting, not less, unless volume is steady and high. The idle penalty, the DevOps salaries and the six-to-eight-weekly model refresh cycle eat the savings.
Second, Latham itself is not a cloud defector — it still uses Harvey and other commercial tools. The realistic enterprise pattern is routing, not choosing: own the compute for sensitive data and high-volume predictable workloads, rent frontier capability for the hard tail. The law firm bought an exit ramp, not a new highway. That is precisely why the move is significant rather than theatrical: it is boring, budgeted infrastructure planning by a firm with no appetite for experiments.
What this means for New Zealand
New Zealand’s economy runs on exactly the kind of expertise that makes this story matter: niche, deep, hard-won domain knowledge in agritech, geothermal engineering, earthquake resilience, marine science, health data. That kind of knowledge is small, concentrated and extremely valuable — which makes it precisely the fuel that the chat-input stage of the AI economy can quietly extract. A lone expert pasting crown-jewel work into a chat box has no leverage and no audit trail. We have made the case for New Zealand building sovereign AI infrastructure on open-weight models and renewable energy before; the Latham story adds the enterprise-scale version of the same argument. A country that only rents its AI capability can never verify what happened to the data it fed in — and its best industries are one pricing change or terms-of-service update away from a very bad quarter.
The practical takeaway for NZ firms is smaller than a server rack but the same idea. The questions that matter in 2026 are not “which model is smartest” but “who can see my data, where does it land, and what happens if the vendor changes the deal” — and the firms answering them best are the ones that kept the option to own.
❓ FAQ
Did OpenAI train on the mathematicians’ work? It has not been proven either way. OpenAI said one internal model “did not look up user data” and told one mathematician in writing that it could not rule out that his conversations contributed to training. No external audit of the training data is possible. That unresolvability is the point of this article.
Is Latham & Watkins abandoning cloud AI? No. It still uses commercial AI services including Harvey. It has bought its own Nvidia GPU servers to keep sensitive workloads and high-volume fine-tuning on infrastructure it controls — a hybrid strategy, per the FT and Bloomberg Law.
Why open-weight models instead of OpenAI or Anthropic? Open-weight models can be downloaded, fine-tuned and run on hardware the buyer controls, with no per-token API and no data leaving the environment. Companies including Thomson Reuters say fine-tuning closed frontier models also degrades general capability and locks them into one provider’s pricing and roadmap.
Is owning your own GPUs actually cheaper? Usually not at low volume. Industry data suggests most enterprise GPU capacity sits idle (average utilisation around 5%), and self-hosting carries staffing and refresh costs that API pricing avoids. The case for owning rests on data sensitivity, predictable volume, and independence — not unit cost alone.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.