Corpus Data & Downloads

478,991 papers ingested from OpenAlex Computer Science, with 8,981,698 citation edges computed. All analytics exports below are generated from ClickHouse and available as raw JSON under CC-BY 4.0.

478,991
Papers ingested
8,981,698
Citation edges
511,313
Abstracts stored
1.5 GB
Compressed size

Sources

arxiv
399,902
78.2% of corpus
biorxiv
63,752
12.5% of corpus
openreview
30,327
5.9% of corpus
medrxiv
17,332
3.4% of corpus

106,595 peer reviews aggregated from OpenReview venues. Average paper size: 3.0 kB.

Papers per year (recent)

2026
380
2025
4,712
2024
25,662
2023
32,591
2022
32,436
2021
40,418
2020
50,744
2019
42,617

Top semantic communities

50 communities detected from citation graph structure. Each has anchor papers and topic labels.

cluster #0 24,393 papers
ADADELTA: An Adaptive Learning Rate Method
languagelearningneuralmodels
cluster #1 18,493 papers
Very Deep Convolutional Networks for Large-Scale Image Recognition
learningsegmentationobjectdeep
cluster #2 17,101 papers
Improving neural networks by preventing co-adaptation of feature detectors
learningneuraldeepnetworks
cluster #3 14,639 papers
Adam: A Method for Stochastic Optimization
imagelearninggenerativedeep
cluster #4 13,568 papers
Quantum computers
quantumqubitmmlsuperconducting
cluster #5 10,478 papers
Fast large-scale optimization by unifying stochastic gradient and quasi-Newton methods
learningfederatedfederated learningprivacy
cluster #6 9,992 papers
Playing Atari with Deep Reinforcement Learning
learningreinforcementreinforcement learningplanning
cluster #7 9,435 papers
Advances in quantum metrology
quantumentanglementstatesmml

Downloadable JSON exports

All files are served from /data/*.json and licensed CC-BY 4.0. Cite as: researchPapers OpenAlex CS Corpus Analytics, papers.highsignal.app/data.

File Description Download
summary.json Corpus ingest summary: paper counts, edge counts, storage size JSON
top_papers.json Top 200 papers by PageRank + Katz centrality JSON
top_cited_works.json Most-cited works across the corpus JSON
hot.json 100 hot papers: citation velocity × review rating × PageRank JSON
sleepers.json 100 reviewer-loved early papers with low citation counts JSON
communities.json 64 semantic communities with anchor papers and labels JSON
embedding_clusters.json Embedding cluster assignments from MiniLM JSON
temporal.json Papers per year, cites per year, top paper per year JSON
tag_rating.json Mean reviewer rating by extracted topic JSON
tag_cooccurrence.json Topic tag co-occurrence matrix JSON
review_venues.json Peer-review venue summaries and review counts JSON
review_rating_distribution.json Rating distribution per review venue JSON
ch_sources_summary.json Paper counts by source (arxiv, OpenReview, bioRxiv, medRxiv) JSON
top_authors.json Most prolific authors in the corpus JSON
top_hosts.json Top hosting domains for paper URLs JSON
cycles.json Citation cycle detection results JSON
abstract_clusters.json Abstract-level clustering results JSON
host_categories.json Host URL categorization JSON

Data provenance

Papers are sourced from OpenAlex (Computer Science, cited ≥1000). Citation edges are computed from OpenAlex reference lists. Peer-review signals come from OpenReview (ICLR, NeurIPS, ICML, COLM). Semantic search uses all-MiniLM-L6-v2 embeddings stored in ClickHouse. The full pipeline runs via scripts/ in the repository.