"Evaluating Large Language Models Trained on Code", arxiv.org/abs/2107.03374
A detailed description of early (pre-GitHub Copilot) versions of OpenAI Codex. This is the "paper of the year" so far: we finally have real progress in AI-assisted computer programming (and difficulties of computer programming form the key bottleneck limiting the speed of progress).
See comments in dmm.dreamwidth.org/44860.html for details.
JuliaCon 2021 starts on July 20 with 8 days of workshops followed by 3 days of main conference. JuliaCon 2020 was great, this is likely to be even better.
This is a fully virtual conference for the second year in a row; the registration is free and needed to access interactive features, poster sessions, and such, but the bulk of materials will be accessible via YouTube without registration. I created a post dmm.dreamwidth.org/46160.html with links and I'll keep populating it with various comments as the conference progresses.
Cross-post: anhinga-anhinga.livejournal.com/85003.html
===
EDIT (Aug 18, 2026): Recently we marked 5 years since the first Codex. That paper has been quite seminal; today we see that applications of AI to coding have been the most important category of AI applications so far and are becoming absolutely central to the overall AI progress.
The paper has a number of prominent authors from OpenAI and Anthropic (many of them are now elsewhere). Back then I noted:
>6 first authors listed as "equal contribution", and I don't really know any of them, although I do recognize some of the more senior authors
The first author, Mark Chen, is now the Chief Research Officer of OpenAI.
The second author, Jerry Tworek, has led the crucial development of OpenAI reasoning models from o1 to GPT-5 series, but has left in January to start a more ambitious effort (Core Automation) trying to find less standard solutions in architecture and training methods of AI systems.
Earlier today, a few comments has been posted under this post by another person, and I am leaving them there as food for thought. The only remark I'd like to make in connection with that material is that the paper by Burtsev and Turchin, "Evolution of cooperative strategies from first principles", Nature, volume 440, pages 1041–1044 (2006) can be downloaded from Petya Turchin's web site peterturchin.com/wp-content/uploads/2023/06/nature04470.pdf and that the first author of that paper, Mikhail Burtsev, has written a number of interesting papers since then and is mostly working in London in recent years: www.linkedin.com/in/mikhail-burtsev-85a47b9/ and scholar.google.com/citations?user=t_PLQakAAAAJ&hl=en (and lims.ac.uk/mikhail-burtsev/ and https://lims.ac.uk/papers/?aid=114).
A detailed description of early (pre-GitHub Copilot) versions of OpenAI Codex. This is the "paper of the year" so far: we finally have real progress in AI-assisted computer programming (and difficulties of computer programming form the key bottleneck limiting the speed of progress).
See comments in dmm.dreamwidth.org/44860.html for details.
JuliaCon 2021 starts on July 20 with 8 days of workshops followed by 3 days of main conference. JuliaCon 2020 was great, this is likely to be even better.
This is a fully virtual conference for the second year in a row; the registration is free and needed to access interactive features, poster sessions, and such, but the bulk of materials will be accessible via YouTube without registration. I created a post dmm.dreamwidth.org/46160.html with links and I'll keep populating it with various comments as the conference progresses.
Cross-post: anhinga-anhinga.livejournal.com/85003.html
===
EDIT (Aug 18, 2026): Recently we marked 5 years since the first Codex. That paper has been quite seminal; today we see that applications of AI to coding have been the most important category of AI applications so far and are becoming absolutely central to the overall AI progress.
The paper has a number of prominent authors from OpenAI and Anthropic (many of them are now elsewhere). Back then I noted:
>6 first authors listed as "equal contribution", and I don't really know any of them, although I do recognize some of the more senior authors
The first author, Mark Chen, is now the Chief Research Officer of OpenAI.
The second author, Jerry Tworek, has led the crucial development of OpenAI reasoning models from o1 to GPT-5 series, but has left in January to start a more ambitious effort (Core Automation) trying to find less standard solutions in architecture and training methods of AI systems.
Earlier today, a few comments has been posted under this post by another person, and I am leaving them there as food for thought. The only remark I'd like to make in connection with that material is that the paper by Burtsev and Turchin, "Evolution of cooperative strategies from first principles", Nature, volume 440, pages 1041–1044 (2006) can be downloaded from Petya Turchin's web site peterturchin.com/wp-content/uploads/2023/06/nature04470.pdf and that the first author of that paper, Mikhail Burtsev, has written a number of interesting papers since then and is mostly working in London in recent years: www.linkedin.com/in/mikhail-burtsev-85a47b9/ and scholar.google.com/citations?user=t_PLQakAAAAJ&hl=en (and lims.ac.uk/mikhail-burtsev/ and https://lims.ac.uk/papers/?aid=114).
