Kyle Harrison
paper

Are Large Language Models a Threat to Digital Public Goods?

Maria del Rio-Chanona, Nadzeya Laurentsyeva, Johannes Wachs July 14, 2023 View original ↗

Are Large Language Models a Threat to Digital Public Goods?

arXiv preprint 2307.07367, submitted July 14, 2023. By Maria del Rio-Chanona, Nadzeya Laurentsyeva and Johannes Wachs.

Key Takeaways

  • The measured effect: a difference-in-differences model estimates a 16% decrease in weekly posts on Stack Overflow following ChatGPT’s release.
  • The identification is the interesting part. They compare Stack Overflow against Russian and Chinese equivalents (where ChatGPT access is limited) and against math forums (where the model performs worse) — so the control groups isolate the LLM effect rather than a general decline in forum use.
  • The decline intensified over time and was most pronounced for widely-used programming languages — exactly where the model is most competent, which is what you would predict if substitution is the mechanism.
  • Post quality stayed comparable, so this was not merely low-quality questions being filtered out. Real contribution left.
  • The recursive argument, and the reason to keep this paper: users now interact privately with a model instead of posting publicly, so the pool of human-generated content shrinks — undermining both future training data and the public learning resource itself. The commons that trained the model is starved by the model.

Connections

  • The empirical spine of Death of Knowledge. It is rare to have a number attached to this argument, and 16% inside months is a large one.
  • The recursive point connects to Open Source Knowledge and to the model-collapse literature: a system that consumes a commons faster than it replenishes it.
  • ⚠️ Note the date — July 2023, months after ChatGPT’s release. It measures the first shock, not the equilibrium. Anyone citing it should say so.