Notes and talks
Write-ups of papers I have presented or read closely, and reports from course and research projects. Each one tries to say what the work establishes and where I think it stops short.
-
Scaling laws for neural language models
On Kaplan et al. (2020). The fitted power laws, the batch-size correction that the compute result depends on, and why the paper's most quoted conclusion rests on its least robust number.