TOPIC
AI Observatory
topictopic-noteresearchevaluation
Overview
The AI Observatory is an independent multi-dataset research project led by Anka Reuel that aggregates real-world frontier-chatbot usage across consented datasets — the first independent-audit substrate positioned to pressure-test vendor self-reports on how people use AI assistants. Launched with 24,521 conversations and 92,493 exchange pairs across seven consented datasets covering Claude, Gemini, and other frontier assistants, with an accompanying paper at NeurIPS 2026.
Timeline
- 2026-08-19-AI-Digest — AI Observatory launches with 24,521 conversations / 92,493 exchange pairs from seven consented datasets, alongside NeurIPS 2026 paper. Early finding: usage patterns diverge sharply by vendor — people go to Anthropic for coding, Gemini for social/roleplay, and ChatGPT for homework — and personal / sensitive use is significantly higher than lab-published usage reports show. Narrow read: 7-dataset aggregation of ~24k conversations, not a global usage census — finding directions are load-bearing; absolute magnitudes are indicative. Structural read the corpus carries: first independent-audit dataset that can pressure-test vendor self-reports on usage. Reads alongside today’s Linear-AI-usage HN datapoint as two independent-of-vendor datasets landing within a week — a substrate shift for AI-adoption claims.
Key Developments
- First Independent Multi-Dataset Picture of Frontier Chatbot Usage (August 19, 2026): 24,521 conversations / 92,493 exchange pairs across seven consented datasets, NeurIPS 2026 paper. The load-bearing structural datum is independent-of-vendor rather than the specific breakdown — Anthropic has been the most public on usage transparency this year (the 46% Claude Code merge-rate disclosure from 2026-08-15-AI-Digest), and the Observatory’s early “coding dominates Claude use” finding corroborates the vendor narrative on that axis while surfacing a gap on personal / sensitive use volumes. 30 / 60 / 90-day watch: whether the Observatory publishes the vendor-by-use-case breakdown at conversation granularity; whether other academic groups fork the seven-dataset methodology; vendor responses (expect coordinated disclosure from Anthropic first given alignment with self-report).
Related
See also: Anthropic, Claude, Gemini, Linear, MOC - Major Companies, MOC - Agent Security.