随笔分类 -  Apache Spark

Apache Spark is a fast and general engine for large-scale data processing.
摘要:Basic Functionssc.parallelize(List(1,2,3,4,5,6)).map(_ * 2).filter(_ > 5).collect()*** res: Array[Int] = Array(6, 8, 10, 12) ***val rdd = sc.parallelize(List(1,2,3,4,5,6,7,8,9,10))rdd.reduce(_+_)*** r... 阅读全文
posted @ 2016-01-06 15:18 linbojin 阅读(175) 评论(0) 推荐(0)
摘要:Apache Spark is an open source cluster computing system that aims to make data analytics fast — both fast to run and fast to write.BDAS, the Berkeley Data Analytics Stack, is an open source software s... 阅读全文
posted @ 2016-01-04 18:37 linbojin 阅读(390) 评论(0) 推荐(0)