89% on ARC-AGI-1, at 2 cents a task, ARC Prize found.

Open weights at that cost let a university or ministry without frontier-API budgets run the model itself.

ARC tests novel visual puzzles no training set can contain, so the score approximates abstraction rather than memorised recall.

China’s DeepSeek has not published production costs; ARC-AGI-2 score sits lower, at 61.4%.

Sources: Hacker News