arXivpreprint repositoryopen accessscientific publishinge-prints

ArXiv: The Open-Access Powerhouse of Scientific Preprints

arXiv: The Open-Access Powerhouse of Scientific Preprints In the traditional world of academic publishing, the journey from a completed study to a public journal article can take months o...

arXiv: The Open-Access Powerhouse of Scientific Preprints

In the traditional world of academic publishing, the journey from a completed study to a public journal article can take months or even years. arXiv (pronounced "archive," where the 'X' represents the Greek letter chi ⟨χ⟩) fundamentally changed this timeline. As an independent, open-access repository, arXiv allows researchers to share e-prints—electronic preprints and postprints—with the global community long before they undergo formal peer review.

Serving as a critical hub for mathematics, physics, astronomy, computer science, quantitative biology, statistics, mathematical finance, and economics, arXiv has become an essential tool for modern science. In many STEM fields, it is now standard practice for authors to self-archive their work on the platform prior to official journal publication.

A screenshot of the arXiv taken in 1994,[9] using the browser NCSA Mosaic. At the time, HTML forms were a new technology.
A screenshot of the arXiv taken in 1994,[9] using the browser NCSA Mosaic. At the time, HTML forms were a new technology.

Key Facts

  • Founded: August 14, 1991, by Paul Ginsparg.
  • Scale: Reached 2 million articles by the end of 2021; currently receives approximately 24,000 submissions per month (as of November 2024).
  • Status: An independent nonprofit organization (separated from Cornell University on July 1, 2026).
  • Core Purpose: Provides open access to scientific papers before or after peer review.
  • Recognition: Named one of the "10 computer codes that transformed science" by Nature in 2021.

The Evolution of arXiv

The birth of arXiv was driven by a practical problem: the limitations of email. Around 1990, physicist Joanne Cohn began emailing physics preprints as TeX files—a compact format that allowed scientific documents to be transmitted easily and rendered on the recipient's computer. However, the volume of papers soon overwhelmed email mailboxes.

Recognizing the need for a centralized system, Paul Ginsparg created a repository mailbox at the Los Alamos National Laboratory (LANL) in August 1991. The system evolved rapidly, adding FTP access in 1991, Gopher in 1992, and the World Wide Web in 1993. Originally known as the LANL preprint archive (with the domain xxx.lanl.gov), the service expanded beyond physics to include other scientific disciplines.

arXiv's yearly submission rate growth over 30 years since its beginning with topics labeled by the standard abbreviations used on arxiv.org[10]
arXiv's yearly submission rate growth over 30 years since its beginning with topics labeled by the standard abbreviations used on arxiv.org[10]

In 2001, the repository moved to Cornell University and was renamed arxiv.org. The name was a creative solution to the fact that "archive" was already taken as a domain; Ginsparg replaced the "chi" with an "X" and dropped the "e" for symmetry. For years, the project was managed by the Cornell University Library, funded by a mix of grants, institutional member fees, and university support.

How arXiv Works

Identification and Versioning

To keep track of millions of documents, arXiv uses a uniquely specific identifier system. Modern papers follow a YYMM.NNNNN format (e.g., 1507.00123), while older papers use category-based prefixes (e.g., hep-th/9901001). Because research is iterative, arXiv supports versioning; for instance, 1709.08980v1 denotes the first version of a paper. If no version is specified, the system defaults to the most recent update.

A screenshot of viewing a paper's abstract on arxiv.org in 2021
A screenshot of viewing a paper's abstract on arxiv.org in 2021

Moderation and Endorsement

While arXiv is not a peer-reviewed journal, it is not an unregulated upload site. It employs a moderation process to ensure content is relevant to the specified disciplines. Since 2004, an endorsement system has required new authors in certain categories to be vetted by an established arXiv author. While authors from recognized academic institutions often receive automatic endorsement, the system ensures that submissions are appropriate for their subject area.

Handling AI and Quality Control

The rise of AI-generated content has led to stricter policies. In November 2025, arXiv announced it would no longer accept computer science review articles or position papers unless they had been vetted by an academic journal or conference. This move was designed to combat the surge of AI-generated research. Additionally, while the site generally re-classifies dubious papers rather than deleting them, approximately 14,000 preprints have been withdrawn, primarily due to crucial errors.

Impact on Scientific Publishing

arXiv was a pioneer of the open access movement, advocating for the free exchange of scientific knowledge. Its influence is evidenced by high-profile cases, such as Grigori Perelman, who uploaded his proof of the Poincaré conjecture to arXiv in 2002 and declined to publish it in a traditional journal, despite being offered the Fields Medal and Clay Mathematics Millennium Prizes.

Today, arXiv is a top 10 global host of "green open access," with its metadata available via OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting). This allows the content to be indexed by major services like BASE, CORE, and Unpaywall.

Feature Details
Founder Paul Ginsparg
Launch Date August 14, 1991
Primary Fields Physics, Math, CS, Astronomy, Biology, Stats, Economics
Legal Status Independent Nonprofit (as of July 2026)
Identifier Format YYMM.NNNNN (Modern) / Category-based (Legacy)
Access Model Open Access (Green)

Frequently Asked Questions

Is arXiv a peer-reviewed journal?

No. arXiv is a repository for preprints and postprints. While it uses a moderation and endorsement system to ensure papers are relevant to their field, it does not perform the formal peer review process associated with academic journals.

What is the difference between a preprint and a postprint?

A preprint is a version of a scientific paper shared before it has undergone peer review. A postprint is the version of the paper after it has been peer-reviewed and accepted for publication, which some publishers allow authors to archive on arXiv.

How is arXiv funded?

Historically, it was funded by Cornell University Library, the Simons Foundation, and annual voluntary contributions from member institutions. These fees are tiered based on the institution's download usage, ranging from $1,000 to $4,400.

Why are some papers withdrawn from arXiv?

Papers are typically withdrawn due to the discovery of crucial errors by the authors or because the preprint has been subsumed by another publication.

Who owns the copyright of papers on arXiv?

Copyright varies by paper. Some are licensed under Creative Commons (CC BY-NC-SA or CC BY-SA), while most remain the copyright of the author, with arXiv holding a non-exclusive irrevocable license to distribute the work.