Vals AI Wants to Be the Neutral Referee of AI Benchmarking
Every few weeks, another lab announces a model that tops the charts. The charts, however, are often produced by the same companies selling the models. That conflict of interest sits at the center of a growing credibility problem in artificial intelligence, and it is the problem Vals AI has set out to solve.
The company positions itself as an independent evaluation firm: a third party that tests large language models against tasks drawn from real professional work, then publishes the results openly. Rather than asking whether a model can solve competition math problems or recall trivia, Vals AI asks whether it can draft a contract clause correctly, reconcile a financial statement, or answer a tax question the way a practitioner would.
Why self-reported scores fall short
Standard academic benchmarks have two persistent weaknesses. The first is contamination: once a test set circulates online, it can end up in training data, and a high score may reflect memorization rather than reasoning. The second is relevance. A model that excels at multiple-choice questions may still fail at the messy, document-heavy, multi-step tasks that define knowledge work.
Vendor-run evaluations add a third issue. Labs choose which benchmarks to report, which competitors to include, and which prompting strategies to use. None of that is necessarily dishonest, but it makes cross-model comparison difficult for buyers who need to make procurement decisions worth millions.
The domain-specific approach
Vals AI's answer has been to build benchmarks with professionals in the relevant field, covering areas such as legal reasoning, corporate finance, tax and healthcare. Test sets are kept private where possible to reduce contamination risk, and grading rubrics are designed with domain experts rather than crowdworkers. The results are published as public leaderboards, so a general counsel or CFO can see how competing systems rank on tasks that resemble their own.
That framing matters because the enterprise market has changed. In 2026, the question is rarely "should we use AI?" It is "which of these dozen near-identical options should we deploy, and how do we justify that choice to risk and compliance?" Neutral, reproducible evidence is the missing piece of that conversation.
The hard part: staying neutral
Independence is easier to claim than to maintain. An evaluation company still needs revenue, and much of that revenue tends to come from the enterprises—and sometimes the labs—whose products it measures. Credibility therefore depends on transparent methodology, disclosed funding relationships, and a willingness to publish unflattering results about well-known models.
There are also technical limits. Benchmarks age quickly as models improve, and any fixed test eventually becomes a target to optimize against. Keeping evaluations meaningful requires continuous refreshes, careful versioning, and honesty about what a score does and does not predict in production.
What it signals for the industry
The rise of independent evaluators mirrors what happened in other maturing sectors: crash-test ratings for cars, clinical trial registries for drugs, credit ratings for debt. Each emerged when the gap between marketing claims and real-world performance became commercially and legally significant.
Whether Vals AI or a competitor becomes the default reference point is unsettled. But the underlying demand is not. As AI systems move into regulated, consequential workflows, buyers will increasingly want performance claims verified by someone who is not selling the model—and that expectation is likely to outlast any single leaderboard.
Install in seconds and keep earning from your phone.
Frequently Asked Questions
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)