AI “solving” math may just be recall.

Whether models generalize decides if they’re safe on problems nobody has written before.

Training data often contains benchmark problems and solutions, so scores conflate memorization with reasoning unless problems postdate training.

Author Davide Piffer’s essay, not a peer-reviewed test, is drawing 551 Hacker News votes.

Sources: Hacker News