LondonA 3-billion-parameter model matched one twice its size.
Reliability from design, not scale, narrows the gap between funded and cheap training runs.
Researchers tested 7 models from 1 to 7 billion parameters, where design explained a quarter of the performance gap.
A quarter of adversarial prompts still defeated every model regardless of size.
Sources: Nature