January 26, 2022

Crossref Releases Updated Public Data File with 120+ Million Metadata Records

From a Crossref Blog Post by Jennifer Kemp:

2020 wasn’t all bad. In April of last year, we released our first public data file. Though Crossref metadata is always openly available––and our board recently cemented this by voting to adopt the Principles of Open Scholarly Infrastructure (POSI)––we’ve decided to release an updated file. This will provide a more efficient way to get such a large volume of records. The file (JSON records, 102.6GB) is now available, with thanks once again to Academic Torrents.


We continue to see around 10% growth in records each year––and while journal articles account for most of the volume, preprints and book chapters are two of our fast-growing content types. In addition to the growth in the number of records, many of the records are getting bigger and better as members look at their participation report and understand the value of enriching metadata records for distribution throughout the scholarly ecosystem. Elsevier recently opened its references, enriching over 12 million records.  A number of members, including Royal Society, Sage, Emerald, OUP, World Scientific and more have started adding abstracts which now number over 9 million.

Learn More, Read the Complete Blog Post

About Gary Price

Gary Price (gprice@mediasourceinc.com) is a librarian, writer, consultant, and frequent conference speaker based in the Washington D.C. metro area. Before launching INFOdocket, Price and Shirl Kennedy were the founders and senior editors at ResourceShelf and DocuTicker for 10 years. From 2006-2009 he was Director of Online Information Services at Ask.com, and is currently a contributing editor at Search Engine Land.