arXiv (Cornell University) · 2019 · 22 citations · 10 references
Data RepresentationEngineeringGraph DatabaseLarge-scale DatasetsGraph ProcessingText MiningNatural Language ProcessingInformation RetrievalData ScienceData MiningCitation AnalysisData ManagementBenchmark DatasetsNatural Graph StructureKnowledge DiscoveryComputer ScienceCitation GraphEdge Citation GraphExciting CandidateGraph TheoryBusinessData Modeling
The arXiv has collected 1.5 million pre-print articles over 28 years, hosting literature from scientific fields including Physics, Mathematics, and Computer Science. Each pre-print features text, figures, authors, citations, categories, and other metadata. These rich, multi-modal features, combined with the natural graph structure---created by citation, affiliation, and co-authorship---makes the arXiv an exciting candidate for benchmarking next-generation models. Here we take the first necessary steps toward this goal, by providing a pipeline which standardizes and simplifies access to the arXiv's publicly available data. We use this pipeline to extract and analyze a 6.7 million edge citation graph, with an 11 billion word corpus of full-text research articles. We present some baseline classification results, and motivate application of more exciting generative graph models.
10
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7
Relational inductive biases, deep learning, and graph networks
Peter Battaglia, Jessica B. Hamrick, Victor Bapst et al. · arXiv (Cornell University) · 2018 · 2.4K citations · Full text
Artificial Intelligence, Large Ai Model, Cognitive Science +14
Daniel Cer, Yinfei Yang, Sheng-yi Kong et al. · arXiv (Cornell University) · 2018 · 1.3K citations · Full text