Datasets: A Community Library for Natural Language Processing

  • 2021-09-07 03:59:22
  • Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander M. Rush, Thomas Wolf
  • 235

Abstract

The scale, variety, and quantity of publicly-available NLP datasets has grownrapidly as researchers propose new tasks, larger models, and novel benchmarks.Datasets is a community library for contemporary NLP designed to support thisecosystem. Datasets aims to standardize end-user interfaces, versioning, anddocumentation, while providing a lightweight front-end that behaves similarlyfor small datasets as for internet-scale corpora. The design of the libraryincorporates a distributed, community-driven approach to adding datasets anddocumenting usage. After a year of development, the library now includes morethan 650 unique datasets, has more than 250 contributors, and has helpedsupport a variety of novel cross-dataset research projects and shared tasks.The library is available at https://github.com/huggingface/datasets.

 

Quick Read (beta)

loading the full paper ...