A 3-billion-parameter model matched one twice its size.

Reliability from design, not scale, narrows the gap between funded and cheap training runs.

Researchers tested 7 models from 1 to 7 billion parameters, where design explained a quarter of the performance gap.

A quarter of adversarial prompts still defeated every model regardless of size.

Sources: Nature