Vals AI Wants to Be the Neutral Referee of AI Benchmarking

Sep 20, 2026 - 14:55
Updated: 20 days ago
0 4
Vals AI Wants to Be the Neutral Referee of AI Benchmarking
Analyst reviewing performance charts on a computer screen, illustrating independent evaluation of AI models.

Every few weeks, another lab announces a model that tops the charts. The charts, however, are often produced by the same companies selling the models. That conflict of interest sits at the center of a growing credibility problem in artificial intelligence, and it is the problem Vals AI has set out to solve.

The company positions itself as an independent evaluation firm: a third party that tests large language models against tasks drawn from real professional work, then publishes the results openly. Rather than asking whether a model can solve competition math problems or recall trivia, Vals AI asks whether it can draft a contract clause correctly, reconcile a financial statement, or answer a tax question the way a practitioner would.

Why self-reported scores fall short

Standard academic benchmarks have two persistent weaknesses. The first is contamination: once a test set circulates online, it can end up in training data, and a high score may reflect memorization rather than reasoning. The second is relevance. A model that excels at multiple-choice questions may still fail at the messy, document-heavy, multi-step tasks that define knowledge work.

Vendor-run evaluations add a third issue. Labs choose which benchmarks to report, which competitors to include, and which prompting strategies to use. None of that is necessarily dishonest, but it makes cross-model comparison difficult for buyers who need to make procurement decisions worth millions.

The domain-specific approach

Vals AI's answer has been to build benchmarks with professionals in the relevant field, covering areas such as legal reasoning, corporate finance, tax and healthcare. Test sets are kept private where possible to reduce contamination risk, and grading rubrics are designed with domain experts rather than crowdworkers. The results are published as public leaderboards, so a general counsel or CFO can see how competing systems rank on tasks that resemble their own.

That framing matters because the enterprise market has changed. In 2026, the question is rarely "should we use AI?" It is "which of these dozen near-identical options should we deploy, and how do we justify that choice to risk and compliance?" Neutral, reproducible evidence is the missing piece of that conversation.

The hard part: staying neutral

Independence is easier to claim than to maintain. An evaluation company still needs revenue, and much of that revenue tends to come from the enterprises—and sometimes the labs—whose products it measures. Credibility therefore depends on transparent methodology, disclosed funding relationships, and a willingness to publish unflattering results about well-known models.

There are also technical limits. Benchmarks age quickly as models improve, and any fixed test eventually becomes a target to optimize against. Keeping evaluations meaningful requires continuous refreshes, careful versioning, and honesty about what a score does and does not predict in production.

What it signals for the industry

The rise of independent evaluators mirrors what happened in other maturing sectors: crash-test ratings for cars, clinical trial registries for drugs, credit ratings for debt. Each emerged when the gap between marketing claims and real-world performance became commercially and legally significant.

Whether Vals AI or a competitor becomes the default reference point is unsettled. But the underlying demand is not. As AI systems move into regulated, consequential workflows, buyers will increasingly want performance claims verified by someone who is not selling the model—and that expectation is likely to outlast any single leaderboard.

Earn money for reading
Registered readers earn a reward for every article they read to the end. Log in or create a free account to start earning.
Free Android App
Read and earn on the go: get the Earnships app

Install in seconds and keep earning from your phone.

Download App

Frequently Asked Questions

Vals AI is an independent evaluation firm that tests large language models on tasks taken from real professional work rather than academic trivia. It measures things like drafting a contract clause, reconciling financial statements, or answering tax questions, and then publishes the outcomes as open leaderboards.

Labs decide which benchmarks to report, which rivals to include, and which prompting methods to use, which makes fair cross-model comparison difficult. This isn't necessarily dishonest, but it leaves enterprise buyers making multimillion-dollar procurement decisions without neutral evidence.

Contamination happens when a public test set circulates online and ends up inside a model's training data. When that occurs, a high score may simply reflect memorization instead of genuine reasoning, so Vals AI keeps test sets private where it can to limit the risk.

The company develops its evaluations together with practitioners in fields such as law, corporate finance, tax and healthcare. Grading rubrics are written by domain experts rather than crowdworkers, so results reflect how a professional would judge the work.

Neutrality is hard to sustain because revenue often comes from the enterprises, and sometimes the labs, whose products are being tested, so transparent methodology and disclosed funding are essential. There are technical limits too: benchmarks become outdated as models improve and any fixed test eventually becomes something to optimize against, requiring constant refreshes and versioning.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0

Comments (0)

User