MODEL
Faraday
modeltopic-noteagentresearch-replication
Overview
Faraday is Inherent‘s 27B-parameter research-replication agent, disclosed on 2026-08-22 as the first commercial ship of the week’s harness-and-scaffolding motion. Faraday is designed to reproduce published scientific papers end-to-end and uses GPT-5.5 Codex as its coding tool, orchestrated via Inherent’s own harness. Per Inherent’s disclosure, Faraday beats Claude Opus 4.8 and GPT-5.5 on the Replica paper-replication suite (310 tasks across 100 papers) at a fraction of the params — a specialist scaffolding result on Inherent’s own benchmark.
Timeline
- 2026-08-23-AI-Digest — Faraday publicly disclosed by Inherent as a 27B agent purpose-built to reproduce published scientific papers end-to-end, using GPT-5.5 Codex as its coding tool. Per Inherent’s own report, Faraday beats Claude Opus 4.8 and GPT-5.5 on the Replica suite (310 tasks / 100 papers). Ships alongside Inherent’s stealth exit ($50M Index-led seed, Radical participating). Narrow read the corpus carries: numerical delta is Inherent’s own report on Inherent’s own suite — Aider polyglot top-5 is still GPT-5 / o3-pro / Gemini with no specialist-agent entry, and SWE-bench Science has even the frontier stack below 50% on general scientific coding tasks. Correct read: “specialist scaffolding beats generalist frontier on the specialist’s own eval,” the historical shape of this beat. Structural read: Faraday is the first commercial instance of the harness-heavy motion landing the same day as EnvHarness (arXiv:2608.19880), FACET (arXiv:2608.18580), the Princeton/UCSD skills study, and two coding-agent practitioner posts — five same-day signals on the same axis. Frame to carry: procedural scaffolding, tool-use engineering, and verification are the actionable near-term surfaces; weight capability still sets the ceiling (Terminal-Bench 2.1: GPT-5.6 Sol 89.5% vs Claude Opus 5 89.1%; SWE-bench Pro: Opus 5 79.2% vs 64.6%).
Key Developments
- First Commercial Instance of the Harness-Heavy Motion — 27B Agent Orchestrating GPT-5.5 as Tool Beats Frontier on Inherent’s Own Replica Suite (August 23, 2026): Faraday’s disclosure is the first commercial ship pointing at the harness-and-scaffolding axis this week’s research beats have been landing on. Load-bearing framing to carry: take the specific structural claim seriously — a 27B model orchestrating a frontier tool and beating the frontier on a domain-specific benchmark is a real result if the eval matches — but the numerical delta is Inherent’s own report on Inherent’s own suite. Do NOT lift “harness > weights” from this — weights are still doing load-bearing work (Terminal-Bench 2.1 89.5% vs 89.1%; SWE-bench Pro 79.2% vs 64.6%). Structural read: compositional beat where the specialist stack won its own eval while the generalist frontier remains the ceiling. 30 / 60 / 90-day watch: independent third-party replication of the Replica benchmark result; whether Inherent opens the harness or the eval so the “harness on top of frontier tool” pattern can be reproduced without Inherent’s own infrastructure; whether a second scaffolding-heavy specialist ships within 30 days with a similar structure.
Related
See also: Inherent, GPT-5.5, Claude Opus 4.8, MOC - Agentic Coding, MOC - Open Source Models.