As language models hit parity, AI builders must hand billions in copyright settlements to proprietary data owners like the New York Times.