MODEL
DarwinX
Overview
DarwinX is a research artifact (arXiv:2608.07545, ▲70 as of publication) that treats agent self-improvement as population-level selection over harnesses — prompts, tools, skills, control flow — with a frozen base model. It uses a preserve-and-extend contract so variants can only be admitted if they extend coverage without regression, and reports a WebArena-Infinity jump from 43.5% → 93.0% on one evolution loop with additional gains on Terminal-Bench 2.1 (reported ~85%). The corpus tracks DarwinX as one of the load-bearing data points that harness search can turn eval compute into durable capability without touching weights — pairing directly with the Anthropic Auto Mode default-on rollout as two same-day signals that near-term agent-quality gains are landing at the harness layer, not the weights layer.
Timeline
- 2026-08-16-AI-Digest — DarwinX paper (arXiv:2608.07545, ▲70) lands as one of the day’s HuggingFace paper picks. Treats agent self-improvement as population-level selection over harnesses (prompts, tools, skills, control flow) with a frozen base model, using a preserve-and-extend contract so variants can only be admitted if they extend coverage without regression. Reports WebArena-Infinity 43.5% → 93.0% on one evolution loop, with additional gains on Terminal-Bench 2.1 (reported ~85%). Narrow read the digest carries: shows harness search can turn eval compute into durable capability without touching weights. Structural read the digest carries: pair with today’s Anthropic Auto Mode default-on flip as two independent same-day data points that near-term agent-quality gains are landing at the harness layer, not the weights layer — Auto Mode ships harness-level classifiers as the paid-tier default reporting 89% dangerous-command catch; DarwinX shows harness evolution jumps a coding-eval score by ~50 percentage points with a frozen model. Log against MOC - Agentic Coding and MOC - Developer Tools.
Key Developments
- Harness Evolution With a Frozen Base Model — WebArena-Infinity 43.5% → 93.0% on One Evolution Loop, Terminal-Bench 2.1 Reported ~85% (August 16, 2026): Population-level selection over harnesses (prompts, tools, skills, control flow) under a preserve-and-extend contract that only admits variants extending coverage without regression. Load-bearing corpus framing: harness search can turn eval compute into durable capability without touching weights — a category-different result from weights-side scaling or post-training. Pairs with today’s Anthropic Auto Mode default-on rollout on Pro / Max / Team as two independent same-day data points that near-term agent-quality gains are landing at the harness layer with a frozen base model. 30 / 60 / 90-day watch: whether the DarwinX harness-evolution recipe gets picked up by any lab as a shipped training loop rather than a research artifact; whether independent replications of harness-evolution-with-frozen-model gains land inside 90 days; whether the “harness search over frozen weights” framing becomes a distinct product line adjacent to weights-side pretraining and post-training.
Related
See also: Anthropic, Auto Mode, Claude Code, MOC - Agentic Coding, MOC - Developer Tools.