“The amount of information per flop has been decreasing as you had to unroll 2 days worth of thinking in order to see if you even did something correctly.”
@Dwarkesh Podcast
RL's bits-per-flop problem pushes spend off the training update onto rollout/inference compute, where CoreWeave bills.
“in the fullness of time, compute kind of is the single most important determinant on how things work... as you scale up energy and compute and parameters,…”
@Dwarkesh Podcast
If intelligence falls out of raw scale, the compute owner compounds while efficiency tricks get arbitraged away.