Skip to content

28 September 2026

Fine-tuning LLMs in the Enterprise: Better Than a Frontier Model?

Julien DELAMOTTE

Julien DELAMOTTE

Head of the Data & AI Business Unit

The next battle in AI could be one of specialization

For the past two years, the race for generative AI seems to have followed a fairly simple pattern: with each new generation of models, we seek to determine which one is the most powerful. GPT, Claude, Gemini, and other so-called “frontier” models are constantly pushing the boundaries in areas such as reasoning, code generation, multimodality, and tool usage.

This race is exciting, but it sometimes leads us to ask the wrong question when it comes to deploying AI in business. Do we really need the best model in the world for every query? Or do we primarily need the best model to perform a specific business task, with a level of quality, latency, and cost compatible with large-scale deployment?

This distinction is essential. An extremely powerful general-purpose model is not necessarily the most cost-effective model for a clearly defined business process. Conversely, a properly specialized open-weight model can perform exceptionally well within a much narrower scope, to the point of competing with—or even outperforming—frontier models on certain tasks.

This is precisely what recent advances in fine-tuning and, more broadly, in post-training are beginning to demonstrate.

A general-purpose model is designed to be able to do almost anything

Frontier models are extraordinary because they can solve an impressive variety of problems. The same model can explain a scientific concept, analyze a contract, generate Python code, write a marketing campaign, translate a text, or analyze an image.

This versatility is immensely valuable. But it also means that we are drawing on a system designed to handle a vast range of problems, even when our specific need is much more limited.

Let’s consider a repetitive business process: analyzing contracts, qualifying customer requests, verifying document compliance, generating a financial summary in a specified format, or searching for information in a business document database. Once the process has been sufficiently well defined, the challenge is no longer about having universal intelligence. It is about producing an extremely reliable response for a specific task—potentially hundreds of thousands or millions of times.

In this context, consistently paying for a very powerful general-purpose model may amount to oversizing the computational infrastructure needed for the task.

The focus then shifts from the model’s raw power to its effectiveness.

Fine-tuning a model isn’t just about training it on documents

The term “fine-tuning” is still a source of confusion. We sometimes hear that all it takes is to take a company’s internal documents and feed them into a model to impart knowledge about the organization to it.

That’s usually not the right line of reasoning.

When information needs to be private, verifiable, frequently updated, or linked to a specific source, a RAG architecture is often the better choice. The model will retrieve the information it needs when generating its response.

Fine-tuning addresses a different issue. It is used more to teach the model how to perform a task.

How do you analyze a document? What elements are important? In what order should you look for information? When should you refrain from answering? How should you structure the results? How precise should citations be? What tools should be used? What criteria allow an expert to determine that the work has been done correctly?

That is why it is now more accurate to refer to “post-training.” Beyond traditional supervised fine-tuning, modern approaches can combine expert examples, synthetic data, model distillation, reinforcement learning, and training environments that replicate the situations the model will encounter in production.

The challenge, therefore, is no longer simply to feed knowledge into the model. It is to impart some of the domain expertise to it.

Harvey Tenet: Teaching a Model to Work Like a Lawyer

In this regard, the work recently published by Harvey is a particularly interesting case study. Harvey develops AI solutions for law firms and legal departments. In August 2026, the company introduced Harvey Tenet, a model based on Kimi K3 and fine-tuned with Fireworks to perform complex legal tasks over long time horizons.

This approach is interesting because Harvey did not simply train the model on legal texts. The training environments replicate real-world work situations: an instruction corresponding to a request from a partner, a client file containing the necessary documents, and an expert evaluation grid detailing the elements that high-quality work must include. The model also has the tools needed to search through documents, analyze them, and produce its output.

The corpus combines synthetic data, public legal data, and data produced or validated by human experts. Harvey also explicitly highlights the crucial role of this expert data in his post-training process. On the LAB benchmark validation tasks, the final model succeeded on nearly twice as many tasks as the initial Kimi K3 model. It also improved its performance on LAB Contracts and, according to Harvey, achieved state-of-the-art performance on this benchmark at the time of publication.

This result illustrates an important point: the original model has not fundamentally changed. It has not suddenly become smarter in every area. It has become much more effective in a specific environment because it was taught the behaviors that lead to high-quality legal output.

The true benchmark: quality per euro spent

The most interesting aspect of the Harvey experiment, however, is not just the improvement in quality. It is the ability to simultaneously improve both the quality and the cost-effectiveness of the system.

In a large-scale document analysis study, Harvey fine-tuned GLM-5.2 to better perform structured extraction tasks, produce more concise responses, and generate more accurate citations. Compared to the best-performing baselines tested, Harvey delivers improved response and citation quality at a cost per cell that is approximately ten times lower. In particular, the model learns to refrain from responding when the question does not apply to the document and to provide more precise evidence rather than generating numerous unnecessary citations.

Another result is perhaps even more revealing. To explore and leverage a firm’s accumulated knowledge, Harvey and Engram worked on a Qwen model with 27 billion parameters. Post-training teaches the model to better structure and leverage this knowledge in order to avoid the repetitive and inefficient searches performed by the base model.

According to the published results, this specialization reduces the total number of tokens used on completed trajectories by 58%. The model then achieves performance on par with the leading frontier models tested, while reducing the cost per query by approximately 90%.

This is probably where the paradigm shift lies.

We have become accustomed to benchmarking LLMs primarily based on their absolute performance levels. For a company, this perspective is incomplete. A much more meaningful metric is the quality produced per euro spent.

Invest in training to sustainably reduce operating costs

The business model then changes.

With a Frontier model accessed via an API, the upfront cost is low. You can get started quickly and immediately benefit from the capabilities of an excellent model. This is extremely effective for prototyping a use case, testing its adoption, or processing limited volumes.

The cost then increases with use.

Specializing an open-weight model shifts part of this effort to the earlier stages of the process. This involves building a corpus, selecting the right data, engaging domain experts, defining evaluation criteria, performing training, and scaling up inference.

This investment should not be downplayed. The Tenet project presented by Harvey is a full-fledged research program. Harvey states that he used approximately 150 NVIDIA B300 GPUs for two months to train the model. It would therefore be misleading to present this case as a simple fine-tuning task that can be accomplished with just a few thousand euros.

But obviously, not all specializations require this scale. Techniques such as LoRA and, more broadly, Parameter-Efficient Fine-Tuning make it possible to fine-tune only a limited fraction of a model’s parameters. They thus significantly reduce memory and computational requirements compared to full fine-tuning.

The equation therefore becomes particularly interesting for repetitive, high-volume tasks. We accept a larger initial investment to create a specialized model, but each subsequent run then becomes potentially much less expensive.

The higher the volume, the more this investment can be recouped.

The business corpus is becoming a true strategic asset

This development also shifts the source of differentiation.

Open-weight models are accessible to many companies. Two organizations can download the exact same model and have the same infrastructure. Yet, after a few months of work, they may achieve radically different performance levels—an issue that ties into the questions of digital sovereignty we’re seeing in other data and AI projects.

Why? Because they don’t have access to the same training data.

Let’s imagine a company that has tens of thousands of real-world cases—representative of its business—along with the decisions made by its top experts. It can distinguish between a merely correct answer and an excellent one. It is familiar with common errors, borderline cases, and situations in which human validation is required. It is also capable of transforming this knowledge into automatable evaluation criteria.

This company has considerable assets.

The challenge, then, is no longer simply about having access to the best LLM. It is about successfully structuring the expertise accumulated by the organization to transform it into data that the models can use.

In my view, this is a major development for data and AI strategies. Until now, many companies have invested in data collection and governance. They will now have to learn how to build a new category of assets: business learning data and internal benchmarks.

Quality examples, expert annotations, corrections made during production, preferences, decision rules, and evaluations are gradually becoming a form of intellectual property, which requires a dedicated AI governance framework, just like traditional data governance.

RAG, fine-tuning, and frontier models are not competing with one another

It would nevertheless be an overstatement to conclude that all companies must replace their frontier models with specialized open-weight models.

These approaches address different needs.

A frontier model remains extremely useful when the problem is new, complex, variable, or poorly defined. It is also an excellent way to get a project off the ground quickly without immediately investing in a training pipeline.

RAG remains essential when it is necessary to provide the system with proprietary, large-scale, or regularly updated knowledge.

Specialization becomes particularly worthwhile when the task is stable, repetitive, and large-scale enough to justify the investment required for learning.

The AI architecture of the future might therefore resemble less a company that has “chosen its LLM” and more a portfolio of models and strategies for accessing knowledge.

Frontier models will be able to handle the most complex problems. Specialized models will execute high-volume business processes. Much smaller models will handle classification, extraction, or routing. RAG will provide up-to-date knowledge. An orchestration layer will determine which model to use based on the task’s difficulty, the expected quality level, and the acceptable cost.

This approach is more complex to design, but it is also much closer to a truly industrial approach.

At DATASOLUTION, we start with the business rather than the model

This is precisely the approach we aim to promote at DATASOLUTION, as an AI and data agency.

The first question in an AI project shouldn’t be, “Which LLM are we going to use?” We need to start by understanding the business context. What task do we want to improve? How is the quality of this task currently measured? Which errors are acceptable or unacceptable? What volume will need to be processed? What level of latency do we expect? And most importantly, what economic value does an improvement of a few quality points generate?

From there, it becomes possible to build a true industry benchmark. We work with experts to create a golden dataset, then evaluate various frontier and open-weight models using the same criteria. We compare not only quality, but also cost, latency, robustness, and deployment constraints.

If a frontier model achieves 95% accuracy and an open-weight model achieves 70%, the matter is likely settled. On the other hand, if the open-weight model already achieves 90%, the question becomes much more interesting: how much should be invested to close the remaining five-point gap through fine-tuning or post-training, and how much will that investment subsequently save in production?

That is when fine-tuning ceases to be a technical experiment and becomes a genuine economic strategy—an approach to measuring impact that we recently discussed in detail at the RAISE Summit.

The next frontier may well be specialization

The race to develop frontier models will continue. Future generations will be smarter, faster, and more multimodal. They will remain essential for pushing the boundaries of what can be automated.

But another competition is emerging.

It’s no longer just about building a model that can do everything. It’s about building a model that can do one thing—something that really matters to the company—exceptionally well.

In this new equation, the open-weight model is just the starting point. The true competitive advantage is then built on business data, human expertise, evaluation systems, and the ability to organize post-training activities.

The goal, therefore, is not to demonstrate that an open-weight model is inherently better than GPT, Claude, or Gemini. That would be the wrong way to frame the problem.

The relevant question is much more practical: for a precisely defined business task, can we build a specialized model that is as good as—or better than—a frontier model, at a much lower operating cost?

More and more findings suggest that the answer may be yes.

And that may be where one of the next major sources of ROI for generative AI in business lies: investing in teaching a model about our line of work, so that we no longer have to pay for all the intelligence in the world with every query when what we really need is exceptional intelligence specific to our field.

Are you wondering which model is right for one of your business processes? Our AI and Data Agency experts can help you develop a business benchmark and work with you to assess the value of specialization through fine-tuning. Schedule an appointment with our experts.

FAQ

Further Reading

Find more insights on AI in business on the DATASOLUTION blog, and learn more about our AI and Data Agency services.

Sources

Harvey — Post-training update: Harvey Tenet

Hugging Face — PEFT: Overview of Methods (LoRA)

Discover the datasolution galaxy