Open-source drop for "Mirages and Measurements: A Dual-Process Audit of ARC-AGI-2, with Negative Results" — ARC Prize 2026 Paper Track. MIT-0. Every number is a hidden-set leaderboard score with its live submission ref.
Find a file
DajaiStewart b41e248d7e band: 2.92 (n=4) -> 3.20 (n=9) from the live ledger, 2026-09-04
Three automated daily rerolls of the frozen v38 kernel (28.89 / 29.31 / 26.94) plus the two
earlier v37/v38 draws widen the identical-code rerun-variance band. Cover regenerated with all
nine draws; tools/band_from_ledger.py recomputes the band from Kaggle on submit day.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-04 05:27:23 -07:00
curriculum Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
hybrid Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
scheduler Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
symbolic Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
tinyarc Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
tools band: 2.92 (n=4) -> 3.20 (n=9) from the live ledger, 2026-09-04 2026-09-04 05:27:23 -07:00
verify Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
cover.png band: 2.92 (n=4) -> 3.20 (n=9) from the live ledger, 2026-09-04 2026-09-04 05:27:23 -07:00
LICENSE Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
PAPER_mirages_and_measurements.md band: 2.92 (n=4) -> 3.20 (n=9) from the live ledger, 2026-09-04 2026-09-04 05:27:23 -07:00
README.md Mirages and Measurements — ARC Prize 2026 paper track open-source drop (MIT-0) 2026-09-02 16:50:59 -07:00
REPRODUCE.md band: 2.92 (n=4) -> 3.20 (n=9) from the live ledger, 2026-09-04 2026-09-04 05:27:23 -07:00
WRITEUP.md band: 2.92 (n=4) -> 3.20 (n=9) from the live ledger, 2026-09-04 2026-09-04 05:27:23 -07:00

Mirages and Measurements — open-source code drop

Companion code for the ARC Prize 2026 Paper Track submission "Mirages and Measurements: A Dual-Process Audit of ARC-AGI-2, with Negative Results" (DAJAI Stewart, Hellcat Blondie LLC).

Competition: Kaggle arc-prize-2026-arc-agi-2. License: MIT-0 (MIT No Attribution) for every file. The ARC Prize 2026 winner obligation reads: "you hereby license and will license your winning Submission and the source code used to generate the Submission under [CC-BY-4.0 license]." MIT-0 is strictly more permissive than CC-BY-4.0 (it additionally waives the attribution condition), so this drop satisfies that obligation on its own terms. See LICENSE.

Everything here is code we authored. The third-party public community TTT notebook our System-1 leg descends from is not redistributed — our scheduler contribution ships as a unified diff against it instead (see scheduler/). That is deliberate: the paper's own limitations section says we did not originate the TTT pipeline, and this repo reflects that.


What supports what

Path Paper section Claim it backs
symbolic/twin_symbolic_solver.py §4.3, §5.2, §6.1 The 5,155-line pure-stdlib program-induction solver: 57 transformation-family generators, D8 handling, CPU-only, no internet. Submitted standalone as ref 54624414 → hidden LB 0.00 (the local-eval mirage).
symbolic/score_local.py §5.2 Local exact-match pass@2 scorer for the solver on the public splits. Reproduces 29/120 = 24.17% on ARC-AGI-2 public eval.
scheduler/time_master.diff §4.4, §5.4 The closed-loop, self-calibrating wall-time controller. Unified diff vs. the baseline notebook's cells 5 and 6. Kernel twin-timemaster-v1, submission ref 54748065 (2026-07-16) → hidden LB 29.44.
hybrid/hybrid_postprocess.py §4.5 The dual-process merge: neural submission.json in, symbolic two_attempts() per task, symbolic written to attempt_2 when it differs from attempt_1. Verbatim from kernel twin-lb33-hybrid-v1 v1 (submission ref 55112943).
curriculum/residual_hard_filter.py §7 The residual-hard partition: split any task set into symbolic-solved vs. residual-hard using the train-exact gate. Measured 30 / 90 / frac 0.75 / 12.6 s on the 120-task public eval split.
curriculum/expand_residual_rearc.py §7 RE-ARC-style expansion worklist generator over the residual-hard ids only. Worklist generator, no ML, no training.
curriculum/residual_hard_agi2eval_report.json §7 The shipped report artifact for the numbers above.
tinyarc/tiny_arc.py §4.6 TinyARC: 38,474 parameters, 86% of weights in the reasoner, per-color embedding + factored 2D positional encoding, K=16 shared recursive self-attention block, per-cell head. Per-task from scratch, no pretraining.
tinyarc/tiny_arc_solve.py, tiny_arc_solve_v2.py §5.5 Per-task trainer / solver: train-fit 0/3 at 300 steps, 3/3 at step 1800 (CE 1.56e6 → 0.31), test-match False.
tinyarc/tiny_arc_consensus.py §5.5 Null 1 Multi-restart (N=3–4) × 8-symmetry consensus voter. Result: 0/3 test matches, restart agreement 0.76–0.91.
tinyarc/tiny_arc_sota.py §4.6 Encoder→decoder variant for shape-changing tasks (the per-cell head can only emit same-shape outputs).
verify/verify_paper_numbers.py §5.2, §7 One command that re-measures every CPU-checkable number in the paper and prints PASS/FAIL.
verify/MEASURED_2026-07-30.txt — Its output as run on 2026-07-30. All three claims PASS.

The long-horizon "grokking" run (§5.5 Null 2 — 17,500 steps, AdamW weight decay 0.05, zero test-match transitions) was executed as a step-count sweep over tiny_arc_solve_v2.py; there is no separate script file for it. Stated plainly here rather than shipping a reconstructed one.


Reproducing the numbers

The ARC datasets are not vendored (they have their own licenses and they are large). Fetch:

python3 verify/verify_paper_numbers.py \
  --agi2-eval /path/to/arc-agi_agi2eval_challenges.json \
  --agi2-sol  /path/to/arc-agi_agi2eval_solutions.json \
  --agi1-dir  /path/to/ARC-AGI/data/evaluation

Measured on an Apple M-series CPU, 2026-07-30 (verify/MEASURED_2026-07-30.txt):

agi2 eval : 29/120 solved = 24.17%   |  train-exact gate hits 30/120, residual-hard 90/120 (0.75)
agi1 eval : 60/400 solved = 15.00%
PASS §5.2 agi2 eval 29/120
PASS §7   gate 30/120, residual-hard 90/120
PASS §5.2 agi1 eval 60/400

symbolic/score_local.py expects the datasets under refs/CompressARC/dataset/ relative to its parent directory; verify/verify_paper_numbers.py takes explicit paths and is the recommended entry point.

tinyarc/* requires Apple MLX (pip install mlx) and runs per-task on CPU/Metal. Nothing in this repository trains a large model, calls a network service, or needs a GPU except the scheduler/ diff, which only makes sense applied inside the Kaggle notebook it targets.


An honest note on the gate (read before citing §4.1)

The paper's architecture section describes the System-2 leg as train-gated: it should not emit unless an induced program reproduces every demonstration pair exactly. In this shipped code that gate is only half-implemented. fitting_programs() is genuinely train-exact — it returns only programs consistent with every demonstration pair. But two_attempts() pads its two slots with a most-common-training-output fallback and identity(test_input) when no fitting program exists, and hybrid/hybrid_postprocess.py writes that fallback into attempt_2 whenever it differs from attempt_1.

Measured consequence on the 120-task public eval split: the gate fires on 30 tasks; on the other 90 (75%) the leg emits ungated fallbacks rather than staying silent. Anyone building on this should add the silence path (return [] when cands is empty, and skip the merge) before claiming a structurally-zero false-positive rate. The paper's draft wording on this point is being corrected against this measurement rather than the other way round.


Not included, on purpose

No revenue, operational, credential, or personal-data code from the author's other systems is in this repository. It contains only ARC research code, and it has been scanned for key material.