COMPANY
Vals AI
Overview
Vals AI is a domain-evaluation startup building third-party benchmarks for enterprise AI buyers across law, finance, engineering, medicine, coding, cybersecurity and biosecurity. The company’s structural bet is contamination-resistance-over-auditability: test items are kept private so models can’t be trained against them, in explicit trade for the openness auditable public benchmarks provide. Vals is one of the vendors positioned to plug into the third-party evaluator surface the Sept 12 pacing coalition and emerging AEF-1 evaluator standard are converging on.
Timeline
- 2026-09-20-AI-Digest — TechCrunch profile surfaces this weekend as the substantive positioning piece on Vals AI’s $40M Series A led by Andreessen Horowitz at a
$400Mpost-money valuation, which actually closed in mid-August 2026 (per FinSMEs at the time). Round participants include 8VC, Pear VC, Bloomberg Beta, HRT Ventures and Next Ladder Ventures. Revenue is8×full-year 2025, team tripled in six months. Vals builds domain evaluations for law, finance, engineering and medicine — TechCrunch adds coding, cybersecurity and biosecurity — and keeps test items private so models can’t be trained against them. Load-bearing framing to carry: positioning isaspirationally "gold-standard independent evaluator", per TechCrunch’s own headline framing — not an achieved standard yet, and no specific enterprise contract wins are named in the coverage. Structural read: a benchmark can be auditable (open, publishable, contestable) or uncontamination-resistant (private, closed) but rarely both, and enterprise buyers increasingly want the latter. The Anthropic-Accenture embedded-evaluators arrangement (2026-09-19-AI-Digest) and the emerging AEF-1 evaluator standard both raise the market for third-party evaluators that can plug into a lab’s release process; Vals’s private-items approach solves the contamination problem those pacts have to negotiate. Whether it can also solve auditability is what will decide whether the “gold-standard” framing survives contact with the pacing coalition’s disclosure requirements. Log against MOC - Developer Tools and MOC - Agent Security.
Key Developments
-
a16z-Led $40M Series A at $400M Post-Money — Round Closed August, TechCrunch Positioning Profile This Weekend (August 2026 / Covered September 20): The commercial round closed in mid-August at a $400M post-money valuation with a16z lead and 8VC / Pear VC / Bloomberg Beta / HRT Ventures / Next Ladder Ventures participating; this weekend’s TechCrunch piece is the substantive positioning writeup — 8× revenue vs full-year 2025 and a team tripled in six months are the load-bearing operating disclosures. Load-bearing framing to carry: “gold-standard” is aspirational per TechCrunch’s own headline, not an achieved standard — no specific enterprise contract wins named in coverage. 30 / 60 / 90-day watch: whether Vals surfaces as a named third-party evaluator inside any frontier-lab release process; whether the private-items approach comes under pressure from the pacing-coalition’s evolving disclosure requirements; whether a competing evaluator surfaces with an auditability-plus-contamination-resistant hybrid.
-
Contamination-Resistance-Over-Auditability Is the Structural Bet: Keeping benchmark items private trades off the openness that lets external parties reproduce and contest results in exchange for uncontamination. Enterprise buyers increasingly want the latter; the frontier-lab pacing coalition and its AEF-1 evaluator standard raise the value of a plug-in evaluator that has already solved the contamination problem. The open question the corpus should carry: whether Vals can also solve auditability, or whether the AEF-1 standard’s disclosure requirements force a hybrid.
- 2026-09-21-AI-Digest — Vals AI reads more sharply against the AMI Labs / World Labs commercial-quiet-at-All-In backdrop: a neutral, private-items benchmark house is one of the few external handles on world-model claims once the labs’ own disclosures thin out. Yesterday’s (2026-09-20-AI-Digest) TechCrunch positioning profile on the $40M a16z Series A had Vals building domain evaluations for law, finance, engineering, medicine, coding, cybersecurity and biosecurity — none of them yet spatial-intelligence. Load-bearing framing to carry: the corpus watch is whether the a16z-Vals evaluation methodology grows a spatial-intelligence dimension over the next quarter — the more world-model labs go commercial-quiet on product terms, the more valuable a private-items evaluator with a defensible methodology becomes as an external anchor. Log against MOC - Major Companies and MOC - Developer Tools.
Related
See also: Anthropic, Accenture, Irregular, MOC - Developer Tools, MOC - Agent Security.