Skip to content

Concept

Scaling laws

Version 1

The empirical regularity that model capability improves predictably with scale — more compute, more data, more parameters, more intelligence out. In this corpus the term is never derived, always invoked: it is the load-bearing premise under two otherwise different conversations, which is what promotes it from a word to a concept here.

The first use is the ex-NVIDIA engineer at [00:34:47] of the Invest Like The Best conversation, grounding it in the architecture: transformers are “such great sponges … you increase the compute available to a transformer by 10x and you’ll get some log improvement somewhere. And so far the scaling laws really work. They’re really quite beautiful.” The second is DHH at [04:09:08] of the Lex Fridman episode, reading the same regularity off the market: “so far the scaling laws are true, and the more billions are poured in, the more intelligence comes out” — offered as the explanation for the investment wave, with the hedge stated in the same breath: the LLMs “could … eventually plateau. We haven’t seen any evidence of it yet.”

Two boundaries keep the term honest. First, as used here it is the informal compute–capability relationship, not the fitted exponents of the scaling-law literature; nobody in either conversation cites a curve. Second, it is distinct from test-time compute scaling — giving one deployed agent more inference time for a better answer — which the same conversation defines separately at [00:05:14] and treats as its own trend. And note the tense both speakers reach for: “so far.” In this corpus a scaling law is an observed streak with an open end, not a law of nature.

Used by