- Python 100%
Three automated daily rerolls of the frozen v38 kernel (28.89 / 29.31 / 26.94) plus the two earlier v37/v38 draws widen the identical-code rerun-variance band. Cover regenerated with all nine draws; tools/band_from_ledger.py recomputes the band from Kaggle on submit day. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|---|---|---|
| curriculum | ||
| hybrid | ||
| scheduler | ||
| symbolic | ||
| tinyarc | ||
| tools | ||
| verify | ||
| cover.png | ||
| LICENSE | ||
| PAPER_mirages_and_measurements.md | ||
| README.md | ||
| REPRODUCE.md | ||
| WRITEUP.md | ||
Mirages and Measurements — open-source code drop
Companion code for the ARC Prize 2026 Paper Track submission "Mirages and Measurements: A Dual-Process Audit of ARC-AGI-2, with Negative Results" (DAJAI Stewart, Hellcat Blondie LLC).
Competition: Kaggle arc-prize-2026-arc-agi-2.
License: MIT-0 (MIT No Attribution) for every file. The ARC Prize 2026 winner obligation
reads: "you hereby license and will license your winning Submission and the source code used to
generate the Submission under [CC-BY-4.0 license]." MIT-0 is strictly more permissive than
CC-BY-4.0 (it additionally waives the attribution condition), so this drop satisfies that
obligation on its own terms. See LICENSE.
Everything here is code we authored. The third-party public community TTT notebook our
System-1 leg descends from is not redistributed — our scheduler contribution ships as a
unified diff against it instead (see scheduler/). That is deliberate: the paper's own
limitations section says we did not originate the TTT pipeline, and this repo reflects that.
What supports what
| Path | Paper section | Claim it backs |
|---|---|---|
symbolic/twin_symbolic_solver.py |
§4.3, §5.2, §6.1 | The 5,155-line pure-stdlib program-induction solver: 57 transformation-family generators, D8 handling, CPU-only, no internet. Submitted standalone as ref 54624414 → hidden LB 0.00 (the local-eval mirage). |
symbolic/score_local.py |
§5.2 | Local exact-match pass@2 scorer for the solver on the public splits. Reproduces 29/120 = 24.17% on ARC-AGI-2 public eval. |
scheduler/time_master.diff |
§4.4, §5.4 | The closed-loop, self-calibrating wall-time controller. Unified diff vs. the baseline notebook's cells 5 and 6. Kernel twin-timemaster-v1, submission ref 54748065 (2026-07-16) → hidden LB 29.44. |
hybrid/hybrid_postprocess.py |
§4.5 | The dual-process merge: neural submission.json in, symbolic two_attempts() per task, symbolic written to attempt_2 when it differs from attempt_1. Verbatim from kernel twin-lb33-hybrid-v1 v1 (submission ref 55112943). |
curriculum/residual_hard_filter.py |
§7 | The residual-hard partition: split any task set into symbolic-solved vs. residual-hard using the train-exact gate. Measured 30 / 90 / frac 0.75 / 12.6 s on the 120-task public eval split. |
curriculum/expand_residual_rearc.py |
§7 | RE-ARC-style expansion worklist generator over the residual-hard ids only. Worklist generator, no ML, no training. |
curriculum/residual_hard_agi2eval_report.json |
§7 | The shipped report artifact for the numbers above. |
tinyarc/tiny_arc.py |
§4.6 | TinyARC: 38,474 parameters, 86% of weights in the reasoner, per-color embedding + factored 2D positional encoding, K=16 shared recursive self-attention block, per-cell head. Per-task from scratch, no pretraining. |
tinyarc/tiny_arc_solve.py, tiny_arc_solve_v2.py |
§5.5 | Per-task trainer / solver: train-fit 0/3 at 300 steps, 3/3 at step 1800 (CE 1.56e6 → 0.31), test-match False. |
tinyarc/tiny_arc_consensus.py |
§5.5 Null 1 | Multi-restart (N=3–4) × 8-symmetry consensus voter. Result: 0/3 test matches, restart agreement 0.76–0.91. |
tinyarc/tiny_arc_sota.py |
§4.6 | Encoder→decoder variant for shape-changing tasks (the per-cell head can only emit same-shape outputs). |
verify/verify_paper_numbers.py |
§5.2, §7 | One command that re-measures every CPU-checkable number in the paper and prints PASS/FAIL. |
verify/MEASURED_2026-07-30.txt |
— | Its output as run on 2026-07-30. All three claims PASS. |
The long-horizon "grokking" run (§5.5 Null 2 — 17,500 steps, AdamW weight decay 0.05, zero
test-match transitions) was executed as a step-count sweep over tiny_arc_solve_v2.py; there is no
separate script file for it. Stated plainly here rather than shipping a reconstructed one.
Reproducing the numbers
The ARC datasets are not vendored (they have their own licenses and they are large). Fetch:
- ARC-AGI-2 public eval — https://github.com/arcprize/ARC-AGI-2 (
data/evaluation, or the Kaggle-formatarc-agi_agi2eval_challenges.json/_solutions.json) - ARC-AGI-1 eval — https://github.com/fchollet/ARC-AGI (
data/evaluation, 400 files)
python3 verify/verify_paper_numbers.py \
--agi2-eval /path/to/arc-agi_agi2eval_challenges.json \
--agi2-sol /path/to/arc-agi_agi2eval_solutions.json \
--agi1-dir /path/to/ARC-AGI/data/evaluation
Measured on an Apple M-series CPU, 2026-07-30 (verify/MEASURED_2026-07-30.txt):
agi2 eval : 29/120 solved = 24.17% | train-exact gate hits 30/120, residual-hard 90/120 (0.75)
agi1 eval : 60/400 solved = 15.00%
PASS §5.2 agi2 eval 29/120
PASS §7 gate 30/120, residual-hard 90/120
PASS §5.2 agi1 eval 60/400
symbolic/score_local.py expects the datasets under refs/CompressARC/dataset/ relative to its
parent directory; verify/verify_paper_numbers.py takes explicit paths and is the recommended
entry point.
tinyarc/* requires Apple MLX (pip install mlx) and runs per-task on CPU/Metal. Nothing in
this repository trains a large model, calls a network service, or needs a GPU except the
scheduler/ diff, which only makes sense applied inside the Kaggle notebook it targets.
An honest note on the gate (read before citing §4.1)
The paper's architecture section describes the System-2 leg as train-gated: it should not emit
unless an induced program reproduces every demonstration pair exactly. In this shipped code that
gate is only half-implemented. fitting_programs() is genuinely train-exact — it returns only
programs consistent with every demonstration pair. But two_attempts() pads its two slots with a
most-common-training-output fallback and identity(test_input) when no fitting program exists, and
hybrid/hybrid_postprocess.py writes that fallback into attempt_2 whenever it differs from
attempt_1.
Measured consequence on the 120-task public eval split: the gate fires on 30 tasks; on the
other 90 (75%) the leg emits ungated fallbacks rather than staying silent. Anyone building on
this should add the silence path (return [] when cands is empty, and skip the merge) before
claiming a structurally-zero false-positive rate. The paper's draft wording on this point is being
corrected against this measurement rather than the other way round.
Not included, on purpose
No revenue, operational, credential, or personal-data code from the author's other systems is in this repository. It contains only ARC research code, and it has been scanned for key material.