
Are we witnessing an inflation of specialist expertise? What justified a dedicated domain model yesterday may be no more than a feature in one of the powerful foundation models tomorrow.
The distinction between the general and the specific defines what academic and specialist publishers do. For a long time, it seemed that this would remain true in the digital future. Many AI-enabled products present themselves as specialised artificial intelligence, while the underlying capability is simultaneously becoming a general-purpose commodity. Behind this lies a robust technical trend: scaling laws have shown for years that language models improve on average with more data, compute and model capacity. Specialist knowledge is explicitly no exception. Could the economic value of a good answer therefore decline because the largest generalists can provide the same answer as well – often without industry-specific training and without the premium pricing the industry expects?
At the same time, the price of capability is falling. According to the Stanford AI Index 2025, the cost of querying a system at GPT-3.5 performance fell from $20 to 7 cents per million tokens between November 2022 and October 2024 – a reduction of more than 280-fold. This creates pressure from both sides: frontier labs push the performance boundary upwards, while more efficient models drive down the cost per unit. Only the explosive growth in usage pushes total costs higher.
Is the specialist provider caught in a trap? Within months, a new API model may match its product on quality while a smaller open model undercuts it on price. One might call this AI inflation.
Medicine provides an instructive example. A comparative study recently published in Nature Medicine tested OpenEvidence and UpToDate Expert AI against GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6: 500 MedQA questions, 500 HealthBench tasks and 100 real clinical queries, with responses assessed in blinded review by twelve physicians. The result: the generalists led in all three parts. On MedQA, Gemini achieved 97.4 per cent, OpenEvidence 89.6 per cent and UpToDate 88.4 per cent. The three frontier models also formed the leading group on the real-world queries.
This is not an isolated case. An earlier study of biomedically fine-tuned open-source models frequently found no advantage for specialists on clinical tasks. The most dramatic example: OpenBioLLM-8B scored 30 per cent on NEJM cases, while its generalist base model, Llama-3-8B-Instruct, reached 64.3 per cent. Specialisation can apparently add knowledge while simultaneously damaging general capabilities.
But caution is warranted: strictly speaking, OpenEvidence is not a publicly documented “small model”. It is a specialised clinical system whose architecture, base model and training pipeline are proprietary. The comparison therefore does not simply prove that “large beats small”. It demonstrates something more important economically: even a vertical product with curated sources and retrieval can be caught by the next generation of horizontal models.
Even this result is not the final word. A later Real-POCQi preprint, which has not yet been peer reviewed, reaches the opposite conclusion. A total of 149 specialty-matched physicians compared responses to 620 real-world questions derived from OpenEvidence usage and 187 HealthBench questions. OpenEvidence received the highest ratings for accuracy, clinical utility, source quality, verifiability and completeness, with net win margins ranging from 25 to 39 percentage points. The study design was agreed with OpenEvidence and the questions came from its platform – context that must be disclosed alongside the paper’s preprint status.
The apparent contradiction is the real insight.
Benchmarks do not measure “intelligence” neutrally. They measure a particular task, user group and evaluation method. Exam questions reward broad knowledge and strong reasoning. Clinical practice additionally rewards current evidence, good citations, verifiability and a suitable workflow. Change the field of measurement and the winner may change with it.
The Nature result should therefore not be taken to mean that general-purpose chatbots can safely make clinical decisions without oversight. Nor does the Real-POCQi result establish that a specialist product is inherently superior. Both studies compare systems at a particular moment. Closed models change, products update their retrieval and prompts, and public tests can enter training data. In high-risk domains, prospective evaluation, error analysis and human responsibility therefore matter more than a single leaderboard score.
Small specialist models are under pressure, as they have been before, but they are far from finished. They can be superior for clearly bounded predictions: a 2026 preprint on hospital operations reports that, after fine-tuning, a clinically pretrained one-billion-parameter model outperformed larger generalists on operational tasks – including zero-shot models with as many as 671 times more parameters. Small models also have an advantage where cost, latency, privacy, local deployment and reproducible behaviour carry substantial weight.
But the moat cannot be: “We know the specialist vocabulary.” Foundation models will absorb that knowledge. What remains defensible is exclusive, continuously updated data; robust evaluation on the real use case; regulatory approval; integration into workflows; traceable sources; and demonstrably lower total costs. The combination appears especially promising: a large generalist orchestrates, while small specialists provide precise signals. A Nature Biomedical Engineering study demonstrates exactly this generalist–specialist principle across 32 medical datasets.
For founders and investors, this leads to a sober four-question test.
First: will the customer benefit still exist when the next foundation model can solve today’s core task without fine-tuning? Second: does the product own data or feedback loops that the platform provider cannot simply buy? Third: can its quality be demonstrated in real cases – including failures, not merely averages? Fourth: does the system measurably reduce time, risk or cost across the entire workflow? Anyone unable to answer these questions may be selling little more than a temporary wrapper around someone else’s intelligence.
My concern therefore remains justified, but it requires a qualification. Large models will not displace many small domain models because specialisation is useless. They will displace them when the specialisation resides only in the model itself. Those who instead own data, trust, evaluation and workflow no longer sell a model, but a reliable chain of decisions – or a complete solution. In a world of AI inflation, that is the difference between an interchangeable feature and an enduring product.
Sources
- Kaplan et al. (2020): Scaling Laws for Neural Language Models
- Stanford HAI (2025): AI Index Report — Research and Development
- Vishwanath et al. (2026): General-purpose large language models outperform specialized clinical AI tools on medical benchmarks, Nature Medicine
- Dorfner et al. (2024): Biomedical Large Language Models Seem not to be Superior to Generalist Models on Unseen Medical Data
- Feng et al. (2026): Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries (Preprint)
- Jiang et al. (2026): Generalist Foundation Models Are Not Clinical Enough for Hospital Operations (Preprint)
- He et al. (2026): Towards generalizable AI in medicine via Generalist–Specialist Collaboration, Nature Biomedical Engineering