Cluster ComputingEngineeringDistributed AlgorithmsNetwork AnalysisGraph DatabaseTrinity Graph EngineGraph ProcessingData ScienceGraph Query LanguageComputing SystemsParallel ComputingGraph AlgorithmsKnowledge DiscoveryEfficient Random AccessComputer ScienceData-intensive ComputingGraph AlgorithmGraph ExplorationGraph DatabasesRelational QueriesNetwork ScienceGraph TheoryCloud ComputingBusinessParallel Programming
Graph algorithms demand high random data access, which disk technology cannot efficiently provide, and memory‑based approaches fail to scale due to single‑machine capacity limits. The paper introduces Trinity, a general‑purpose graph engine built on a distributed memory cloud. Trinity achieves fast graph exploration and efficient parallel computing by optimizing memory management and network communication, leveraging graph access patterns, and offering a high‑level specification language (TSL) for schema and communication. Experiments show Trinity can process low‑latency graph queries and high‑throughput analytics on web‑scale, billion‑node graphs using only a few commodity machines.
Computations performed by graph algorithms are data driven, and require a high degree of random data access. Despite the great progresses made in disk technology, it still cannot provide the level of efficient random access required by graph computation. On the other hand, memory-based approaches usually do not scale due to the capacity limit of single machines. In this paper, we introduce Trinity, a general purpose graph engine over a distributed memory cloud. Through optimized memory management and network communication, Trinity supports fast graph exploration as well as efficient parallel computing. In particular, Trinity leverages graph access patterns in both online and offline computation to optimize memory and communication for best performance. These enable Trinity to support efficient online query processing and offline analytics on large graphs with just a few commodity machines. Furthermore, Trinity provides a high level specification language called TSL for users to declare data schema and communication protocols, which brings great ease-of-use for general purpose graph management and computing. Our experiments show Trinity’s performance in both low latency graph queries as well as high throughput graph analytics on web-scale, billion-node graphs.
21
Velvet: Algorithms for de novo short read assembly using de Bruijn graphs
Daniel R. Zerbino, Ewan Birney · Genome Research · 2008 · 9.6K citations · Full text
Spark: cluster computing with working sets
Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin et al. · UC Berkeley · 2010 · 4.2K citations
Grzegorz Malewicz, Matthew H. Austern, Aart J. C. Bik et al. · 2010 · 3.5K citations
Giuseppe DeCandia, Deniz Hastorun, Madan Jampani et al. · 2007 · 3.4K citations
Fay W. Chang, Sanjay Ghemawat, Wilson C. Hsieh et al. · ACM Transactions on Computer Systems · 2008 · 3.4K citations