spark - Mirror of Apache Spark

	Commit message (Collapse)	Author	Age	Files	Lines
*	[SPARK-6518][MLLIB][EXAMPLE][DOC] Add example code and user guide for ↵	Yu ISHIKAWA	2015-12-16	2	-0/+36
\| \| \| \| \| \| \| \| \| \| \|	bisecting k-means This PR includes only an example code in order to finish it quickly. I'll send another PR for the docs soon. Author: Yu ISHIKAWA <yuu.ishikawa@gmail.com> Closes #9952 from yu-iskw/SPARK-6518.
*	[SPARK-12215][ML][DOC] User guide section for KMeans in spark.ml	Yu ISHIKAWA	2015-12-16	1	-0/+71
\| \| \| \| \| \| \| \|	cc jkbradley Author: Yu ISHIKAWA <yuu.ishikawa@gmail.com> Closes #10244 from yu-iskw/SPARK-12215.
*	[SPARK-12318][SPARKR] Save mode in SparkR should be error by default	Jeff Zhang	2015-12-16	1	-1/+8
\| \| \| \| \| \| \| \|	shivaram Please help review. Author: Jeff Zhang <zjffdu@apache.org> Closes #10290 from zjffdu/SPARK-12318.
*	[SPARK-12324][MLLIB][DOC] Fixes the sidebar in the ML documentation	Timothy Hunter	2015-12-16	3	-33/+141
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This fixes the sidebar, using a pure CSS mechanism to hide it when the browser's viewport is too narrow. Credit goes to the original author Titan-C (mentioned in the NOTICE). Note that I am not a CSS expert, so I can only address comments up to some extent. Default view: <img width="936" alt="screen shot 2015-12-14 at 12 46 39 pm" src="https://cloud.githubusercontent.com/assets/7594753/11793597/6d1d6eda-a261-11e5-836b-6eb2054e9054.png"> When collapsed manually by the user: <img width="1004" alt="screen shot 2015-12-14 at 12 54 02 pm" src="https://cloud.githubusercontent.com/assets/7594753/11793669/c991989e-a261-11e5-8bf6-aecf3bdb6319.png"> Disappears when column is too narrow: <img width="697" alt="screen shot 2015-12-14 at 12 47 22 pm" src="https://cloud.githubusercontent.com/assets/7594753/11793607/7754dbcc-a261-11e5-8b15-e0d074b0e47c.png"> Can still be opened by the user if necessary: <img width="651" alt="screen shot 2015-12-14 at 12 51 15 pm" src="https://cloud.githubusercontent.com/assets/7594753/11793612/7bf82968-a261-11e5-9cc3-e827a7a6b2b0.png"> Author: Timothy Hunter <timhunter@databricks.com> Closes #10297 from thunterdb/12324.
*	[SPARK-10123][DEPLOY] Support specifying deploy mode from configuration	jerryshao	2015-12-15	1	-3/+12
\| \| \| \| \| \| \| \|	Please help to review, thanks a lot. Author: jerryshao <sshao@hortonworks.com> Closes #10195 from jerryshao/SPARK-10123.
*	[SPARK-12351][MESOS] Add documentation about submitting Spark with mesos ↵	Timothy Chen	2015-12-15	2	-6/+35
\| \| \| \| \| \| \| \| \| \|	cluster mode. Adding more documentation about submitting jobs with mesos cluster mode. Author: Timothy Chen <tnachen@gmail.com> Closes #10086 from tnachen/mesos_supervise_docs.
*	[MINOR][DOC] Fix broken word2vec link	BenFradet	2015-12-14	1	-1/+1
\| \| \| \| \| \| \| \|	Follow-up of [SPARK-12199](https://issues.apache.org/jira/browse/SPARK-12199) and #10193 where a broken link has been left as is. Author: BenFradet <benjamin.fradet@gmail.com> Closes #10282 from BenFradet/SPARK-12199.
*	[SPARK-12199][DOC] Follow-up: Refine example code in ml-features.md	Xusen Yin	2015-12-12	1	-11/+11
\| \| \| \| \| \| \| \| \| \| \| \|	https://issues.apache.org/jira/browse/SPARK-12199 Follow-up PR of SPARK-11551. Fix some errors in ml-features.md mengxr Author: Xusen Yin <yinxusen@gmail.com> Closes #10193 from yinxusen/SPARK-12199.
*	[SPARK-12217][ML] Document invalid handling for StringIndexer	BenFradet	2015-12-11	1	-0/+36
\| \| \| \| \| \| \| \| \| \|	Added a paragraph regarding StringIndexer#setHandleInvalid to the ml-features documentation. I wonder if I should also add a snippet to the code example, input welcome. Author: BenFradet <benjamin.fradet@gmail.com> Closes #10257 from BenFradet/SPARK-12217.
*	[SPARK-11964][DOCS][ML] Add in Pipeline Import/Export Documentation	anabranch	2015-12-11	1	-0/+13
\| \| \| \| \| \| \| \| \|	Adding in Pipeline Import and Export Documentation. Author: anabranch <wac.chambers@gmail.com> Author: Bill Chambers <wchambers@ischool.berkeley.edu> Closes #10179 from anabranch/master.
*	[STREAMING][DOC][MINOR] Update the description of direct Kafka stream doc	jerryshao	2015-12-10	1	-1/+1
\| \| \| \| \| \| \| \| \| \|	With the merge of [SPARK-8337](https://issues.apache.org/jira/browse/SPARK-8337), now the Python API has the same functionalities compared to Scala/Java, so here changing the description to make it more precise. zsxwing tdas , please review, thanks a lot. Author: jerryshao <sshao@hortonworks.com> Closes #10246 from jerryshao/direct-kafka-doc-update.
*	[SPARK-12251] Document and improve off-heap memory configurations	Josh Rosen	2015-12-10	1	-0/+16
\| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds documentation for Spark configurations that affect off-heap memory and makes some naming and validation improvements for those configs. - Change `spark.memory.offHeapSize` to `spark.memory.offHeap.size`. This is fine because this configuration has not shipped in any Spark release yet (it's new in Spark 1.6). - Deprecated `spark.unsafe.offHeap` in favor of a new `spark.memory.offHeap.enabled` configuration. The motivation behind this change is to gather all memory-related configurations under the same prefix. - Add a check which prevents users from setting `spark.memory.offHeap.enabled=true` when `spark.memory.offHeap.size == 0`. After SPARK-11389 (#9344), which was committed in Spark 1.6, Spark enforces a hard limit on the amount of off-heap memory that it will allocate to tasks. As a result, enabling off-heap execution memory without setting `spark.memory.offHeap.size` will lead to immediate OOMs. The new configuration validation makes this scenario easier to diagnose, helping to avoid user confusion. - Document these configurations on the configuration page. Author: Josh Rosen <joshrosen@databricks.com> Closes #10237 from JoshRosen/SPARK-12251.
*	[SPARK-11563][CORE][REPL] Use RpcEnv to transfer REPL-generated classes.	Marcelo Vanzin	2015-12-10	2	-16/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This avoids bringing up yet another HTTP server on the driver, and instead reuses the file server already managed by the driver's RpcEnv. As a bonus, the repl now inherits the security features of the network library. There's also a small change to create the directory for storing classes under the root temp dir for the application (instead of directly under java.io.tmpdir). Author: Marcelo Vanzin <vanzin@cloudera.com> Closes #9923 from vanzin/SPARK-11563.
*	[SPARK-12212][ML][DOC] Clarifies the difference between spark.ml, ↵	Timothy Hunter	2015-12-10	31	-1793/+149
\| \| \| \| \| \| \| \| \| \| \| \|	spark.mllib and mllib in the documentation. Replaces a number of occurences of `MLlib` in the documentation that were meant to refer to the `spark.mllib` package instead. It should clarify for new users the difference between `spark.mllib` (the package) and MLlib (the umbrella project for ML in spark). It also removes some files that I forgot to delete with #10207 Author: Timothy Hunter <timhunter@databricks.com> Closes #10234 from thunterdb/12212.
*	[SPARK-11678][SQL][DOCS] Document basePath in the programming guide.	Yin Huai	2015-12-09	1	-0/+7
\| \| \| \| \| \| \| \| \| \| \| \| \|	This PR adds document for `basePath`, which is a new parameter used by `HadoopFsRelation`. The compiled doc is shown below. ![image](https://cloud.githubusercontent.com/assets/2072857/11673132/1ba01192-9dcb-11e5-98d9-ac0b4e92e98c.png) JIRA: https://issues.apache.org/jira/browse/SPARK-11678 Author: Yin Huai <yhuai@databricks.com> Closes #10211 from yhuai/basePathDoc.
*	[SPARK-12211][DOC][GRAPHX] Fix version number in graphx doc for migration ↵	Andrew Ray	2015-12-09	1	-1/+1
\| \| \| \| \| \| \| \| \| \|	from 1.1 Migration from 1.1 section added to the GraphX doc in 1.2.0 (see https://spark.apache.org/docs/1.2.0/graphx-programming-guide.html#migrating-from-spark-11) uses \{{site.SPARK_VERSION}} as the version where changes were introduced, it should be just 1.2. Author: Andrew Ray <ray.andrew@gmail.com> Closes #10206 from aray/graphx-doc-1.1-migration.
*	[SPARK-11551][DOC] Replace example code in ml-features.md using include_example	Xusen Yin	2015-12-09	1	-1061/+51
\| \| \| \| \| \| \| \| \|	PR on behalf of somideshmukh, thanks! Author: Xusen Yin <yinxusen@gmail.com> Author: somideshmukh <somilde@us.ibm.com> Closes #10219 from yinxusen/SPARK-11551.
*	[SPARK-8517][ML][DOC] Reorganizes the spark.ml user guide	Timothy Hunter	2015-12-08	8	-81/+1752
\| \| \| \| \| \| \| \| \| \|	This PR moves pieces of the spark.ml user guide to reflect suggestions in SPARK-8517. It does not introduce new content, as requested. <img width="192" alt="screen shot 2015-12-08 at 11 36 00 am" src="https://cloud.githubusercontent.com/assets/7594753/11666166/e82b84f2-9d9f-11e5-8904-e215424d8444.png"> Author: Timothy Hunter <timhunter@databricks.com> Closes #10207 from thunterdb/spark-8517.
*	[SPARK-12069][SQL] Update documentation with Datasets	Michael Armbrust	2015-12-08	3	-100/+172
\| \| \| \| \| \|	Author: Michael Armbrust <michael@databricks.com> Closes #10060 from marmbrus/docs.
*	[SPARK-12159][ML] Add user guide section for IndexToString transformer	BenFradet	2015-12-08	1	-16/+88
\| \| \| \| \| \| \| \|	Documentation regarding the `IndexToString` label transformer with code snippets in Scala/Java/Python. Author: BenFradet <benjamin.fradet@gmail.com> Closes #10166 from BenFradet/SPARK-12159.
*	[SPARK-11551][DOC][EXAMPLE] Revert PR #10002	Cheng Lian	2015-12-08	1	-51/+1058
\| \| \| \| \| \| \| \| \| \|	This reverts PR #10002, commit 78209b0ccaf3f22b5e2345dfb2b98edfdb746819. The original PR wasn't tested on Jenkins before being merged. Author: Cheng Lian <lian@databricks.com> Closes #10200 from liancheng/revert-pr-10002.
*	[SPARK-11958][SPARK-11957][ML][DOC] SQLTransformer user guide and example code	Yanbo Liang	2015-12-07	1	-0/+59
\| \| \| \| \| \| \| \|	Add ```SQLTransformer``` user guide, example code and make Scala API doc more clear. Author: Yanbo Liang <ybliang8@gmail.com> Closes #10006 from yanboliang/spark-11958.
*	[SPARK-11551][DOC][EXAMPLE] Replace example code in ml-features.md using ↵	somideshmukh	2015-12-07	1	-1058/+51
\| \| \| \| \| \| \| \| \| \| \| \| \|	include_example Made new patch contaning only markdown examples moved to exmaple/folder. Ony three java code were not shfted since they were contaning compliation error ,these classes are 1)StandardScale 2)NormalizerExample 3)VectorIndexer Author: Xusen Yin <yinxusen@gmail.com> Author: somideshmukh <somilde@us.ibm.com> Closes #10002 from somideshmukh/SomilBranch1.33.
*	[SPARK-11963][DOC] Add docs for QuantileDiscretizer	Xusen Yin	2015-12-07	1	-0/+65
\| \| \| \| \| \| \| \|	https://issues.apache.org/jira/browse/SPARK-11963 Author: Xusen Yin <yinxusen@gmail.com> Closes #9962 from yinxusen/SPARK-11963.
*	[SPARK-12080][CORE] Kryo - Support multiple user registrators	rotems	2015-12-04	1	-2/+2
\| \| \| \| \| \|	Author: rotems <roter> Closes #10078 from Botnaim/KryoMultipleCustomRegistrators.
*	[SPARK-12116][SPARKR][DOCS] document how to workaround function name ↵	felixcheung	2015-12-03	1	-1/+2
\| \| \| \| \| \| \| \| \| \|	conflicts with dplyr shivaram Author: felixcheung <felixcheung_m@hotmail.com> Closes #10119 from felixcheung/rdocdplyrmasked.
*	[DOCUMENTATION][MLLIB] typo in mllib doc	Jeff Zhang	2015-12-03	1	-1/+1
\| \| \| \| \| \| \| \|	\cc mengxr Author: Jeff Zhang <zjffdu@apache.org> Closes #10093 from zjffdu/mllib_typo.
*	[SPARK-12081] Make unified memory manager work with small heaps	Andrew Or	2015-12-01	2	-3/+3
\| \| \| \| \| \| \| \| \| \|	The existing `spark.memory.fraction` (default 0.75) gives the system 25% of the space to work with. For small heaps, this is not enough: e.g. default 1GB leaves only 250MB system memory. This is especially a problem in local mode, where the driver and executor are crammed in the same JVM. Members of the community have reported driver OOM's in such cases. New proposal. We now reserve 300MB before taking the 75%. For 1GB JVMs, this leaves `(1024 - 300) * 0.75 = 543MB` for execution and storage. This is proposal (1) listed in the [JIRA](https://issues.apache.org/jira/browse/SPARK-12081). Author: Andrew Or <andrew@databricks.com> Closes #10081 from andrewor14/unified-memory-small-heaps.
*	[SPARK-11961][DOC] Add docs of ChiSqSelector	Xusen Yin	2015-12-01	1	-0/+50
\| \| \| \| \| \| \| \|	https://issues.apache.org/jira/browse/SPARK-11961 Author: Xusen Yin <yinxusen@gmail.com> Closes #9965 from yinxusen/SPARK-11961.
*	[SPARK-11821] Propagate Kerberos keytab for all environments	woj-i	2015-12-01	2	-5/+6
\| \| \| \| \| \| \| \| \|	andrewor14 the same PR as in branch 1.5 harishreedharan Author: woj-i <wojciechindyk@gmail.com> Closes #9859 from woj-i/master.
*	[HOTFIX][SPARK-12000] Add missing quotes in Jekyll API docs plugin.	Josh Rosen	2015-11-30	1	-1/+1
\| \| \| \|	I accidentally omitted these as part of #10049.
*	[SPARK-12035] Add more debug information in include_example tag of Jekyll	Xusen Yin	2015-11-30	1	-4/+6
\| \| \| \| \| \| \| \| \| \|	https://issues.apache.org/jira/browse/SPARK-12035 When we debuging lots of example code files, like in https://github.com/apache/spark/pull/10002, it's hard to know which file causes errors due to limited information in `include_example.rb`. With their filenames, we can locate bugs easily. Author: Xusen Yin <yinxusen@gmail.com> Closes #10026 from yinxusen/SPARK-12035.
*	[SPARK-12000] Fix API doc generation issues	Josh Rosen	2015-11-30	1	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This pull request fixes multiple issues with API doc generation. - Modify the Jekyll plugin so that the entire doc build fails if API docs cannot be generated. This will make it easy to detect when the doc build breaks, since this will now trigger Jenkins failures. - Change how we handle the `-target` compiler option flag in order to fix `javadoc` generation. - Incorporate doc changes from thunterdb (in #10048). Closes #10048. Author: Josh Rosen <joshrosen@databricks.com> Author: Timothy Hunter <timhunter@databricks.com> Closes #10049 from JoshRosen/fix-doc-build.
*	[SPARK-11960][MLLIB][DOC] User guide for streaming tests	Feynman Liang	2015-11-30	2	-0/+26
\| \| \| \| \| \| \| \|	CC jkbradley mengxr josepablocam Author: Feynman Liang <feynman.liang@gmail.com> Closes #10005 from feynmanliang/streaming-test-user-guide.
*	[SPARK-11689][ML] Add user guide and example code for LDA under spark.ml	Yuhao Yang	2015-11-30	3	-1/+34
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	jira: https://issues.apache.org/jira/browse/SPARK-11689 Add simple user guide for LDA under spark.ml and example code under examples/. Use include_example to include example code in the user guide markdown. Check SPARK-11606 for instructions. Original PR is reverted due to document build error. https://github.com/apache/spark/pull/9722 mengxr feynmanliang yinxusen Sorry for the troubling. Author: Yuhao Yang <hhbyyh@gmail.com> Closes #9974 from hhbyyh/ldaMLExample.
*	[MINOR][DOCS] fixed list display in ml-ensembles	BenFradet	2015-11-30	1	-0/+1
\| \| \| \| \| \| \| \| \| \| \| \|	The list in ml-ensembles.md wasn't properly formatted and, as a result, was looking like this: ![old](http://i.imgur.com/2ZhELLR.png) This PR aims to make it look like this: ![new](http://i.imgur.com/0Xriwd2.png) Author: BenFradet <benjamin.fradet@gmail.com> Closes #10025 from BenFradet/ml-ensembles-doc.
*	doc typo: "classificaion" -> "classification"	muxator	2015-11-26	1	-1/+1
\| \| \| \| \| \|	Author: muxator <muxator@users.noreply.github.com> Closes #10008 from muxator/patch-1.
*	[DOCUMENTATION] Fix minor doc error	Jeff Zhang	2015-11-25	1	-1/+1
\| \| \| \| \| \|	Author: Jeff Zhang <zjffdu@apache.org> Closes #9956 from zjffdu/dev_typo.
*	[MINOR] Remove unnecessary spaces in `include_example.rb`	Yu ISHIKAWA	2015-11-25	1	-4/+4
\| \| \| \| \| \|	Author: Yu ISHIKAWA <yuu.ishikawa@gmail.com> Closes #9960 from yu-iskw/minor-remove-spaces.
*	Updated sql programming guide to include jdbc fetch size	Stephen Samuel	2015-11-23	1	-0/+8
\| \| \| \| \| \|	Author: Stephen Samuel <sam@sksamuel.com> Closes #9377 from sksamuel/master.
*	[SPARK-11140][CORE] Transfer files using network lib when using NettyRpcEnv.	Marcelo Vanzin	2015-11-23	2	-2/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This change abstracts the code that serves jars / files to executors so that each RpcEnv can have its own implementation; the akka version uses the existing HTTP-based file serving mechanism, while the netty versions uses the new stream support added to the network lib, which makes file transfers benefit from the easier security configuration of the network library, and should also reduce overhead overall. The change includes a small fix to TransportChannelHandler so that it propagates user events to downstream handlers. Author: Marcelo Vanzin <vanzin@cloudera.com> Closes #9530 from vanzin/SPARK-11140.
*	[SPARK-11910][STREAMING][DOCS] Update twitter4j dependency version	Luciano Resende	2015-11-23	1	-1/+1
\| \| \| \| \| \|	Author: Luciano Resende <lresende@apache.org> Closes #9892 from lresende/SPARK-11910.
*	[SPARK-7173][YARN] Add label expression support for application master	jerryshao	2015-11-23	1	-0/+9
\| \| \| \| \| \| \| \| \| \|	Add label expression support for AM to restrict it runs on the specific set of nodes. I tested it locally and works fine. sryza and vanzin please help to review, thanks a lot. Author: jerryshao <sshao@hortonworks.com> Closes #9800 from jerryshao/SPARK-7173.
*	[SPARK-11835] Adds a sidebar menu to MLlib's documentation	Timothy Hunter	2015-11-22	6	-8/+163
\| \| \| \| \| \| \| \| \| \|	This PR adds a sidebar menu when browsing the user guide of MLlib. It uses a YAML file to describe the structure of the documentation. It should be trivial to adapt this to the other projects. ![screen shot 2015-11-18 at 4 46 12 pm](https://cloud.githubusercontent.com/assets/7594753/11259591/a55173f4-8e17-11e5-9340-0aed79d66262.png) Author: Timothy Hunter <timhunter@databricks.com> Closes #9826 from thunterdb/spark-11835.
*	Revert "[SPARK-11689][ML] Add user guide and example code for LDA under ↵	Xiangrui Meng	2015-11-20	3	-33/+1
\| \| \| \| \| \|	spark.ml" This reverts commit e359d5dcf5bd300213054ebeae9fe75c4f7eb9e7.
*	[SPARK-11549][DOCS] Replace example code in mllib-evaluation-metrics.md ↵	Vikas Nelamangala	2015-11-20	1	-925/+15
\| \| \| \| \| \| \| \|	using include_example Author: Vikas Nelamangala <vikasnelamangala@Vikass-MacBook-Pro.local> Closes #9689 from vikasnp/master.
*	[SPARK-11689][ML] Add user guide and example code for LDA under spark.ml	Yuhao Yang	2015-11-20	3	-1/+33
\| \| \| \| \| \| \| \| \| \|	jira: https://issues.apache.org/jira/browse/SPARK-11689 Add simple user guide for LDA under spark.ml and example code under examples/. Use include_example to include example code in the user guide markdown. Check SPARK-11606 for instructions. Author: Yuhao Yang <hhbyyh@gmail.com> Closes #9722 from hhbyyh/ldaMLExample.
*	[SPARK-11339][SPARKR] Document the list of functions in R base package that ↵	felixcheung	2015-11-18	1	-1/+36
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	are masked by functions with same name in SparkR Added tests for function that are reported as masked, to make sure the base:: or stats:: function can be called. For those we can't call, added them to SparkR programming guide. It would seem to me `table, sample, subset, filter, cov` not working are not actually expected - I investigated/experimented with them but couldn't get them to work. It looks like as they are defined in base or stats they are missing the S3 generic, eg. ``` > methods("transform") [1] transform,ANY-method transform.data.frame [3] transform,DataFrame-method transform.default see '?methods' for accessing help and source code > methods("subset") [1] subset.data.frame subset,DataFrame-method subset.default [4] subset.matrix see '?methods' for accessing help and source code Warning message: In .S3methods(generic.function, class, parent.frame()) : function 'subset' appears not to be S3 generic; found functions that look like S3 methods ``` Any idea? More information on masking: http://www.ats.ucla.edu/stat/r/faq/referencing_objects.htm http://www.sfu.ca/~sweldon/howTo/guide4.pdf This is what the output doc looks like (minus css): ![image](https://cloud.githubusercontent.com/assets/8969467/11229714/2946e5de-8d4d-11e5-94b0-dda9696b6fdd.png) Author: felixcheung <felixcheung_m@hotmail.com> Closes #9785 from felixcheung/rmasked.
*	[SPARK-11684][R][ML][DOC] Update SparkR glm API doc, user guide and example ↵	Yanbo Liang	2015-11-18	1	-8/+42
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	codes This PR includes: * Update SparkR:::glm, SparkR:::summary API docs. * Update SparkR machine learning user guide and example codes to show: * supporting feature interaction in R formula. * summary for gaussian GLM model. * coefficients for binomial GLM model. mengxr Author: Yanbo Liang <ybliang8@gmail.com> Closes #9727 from yanboliang/spark-11684.
*	[SPARK-11809] Switch the default Mesos mode to coarse-grained mode	Reynold Xin	2015-11-18	2	-11/+18
\| \| \| \| \| \| \| \|	Based on my conversions with people, I believe the consensus is that the coarse-grained mode is more stable and easier to reason about. It is best to use that as the default rather than the more flaky fine-grained mode. Author: Reynold Xin <rxin@databricks.com> Closes #9795 from rxin/SPARK-11809.