diff options
author | Joseph K. Bradley <joseph.kurata.bradley@gmail.com> | 2014-10-01 01:03:24 -0700 |
---|---|---|
committer | Xiangrui Meng <meng@databricks.com> | 2014-10-01 01:03:24 -0700 |
commit | 7bf6cc9701cbb0f77fb85a412e387fb92274fca5 (patch) | |
tree | 21d38a426534826700f9f94b8f8d81034f55ea9b /extras | |
parent | eb43043f411b87b7b412ee31e858246bd93fdd04 (diff) | |
download | spark-7bf6cc9701cbb0f77fb85a412e387fb92274fca5.tar.gz spark-7bf6cc9701cbb0f77fb85a412e387fb92274fca5.tar.bz2 spark-7bf6cc9701cbb0f77fb85a412e387fb92274fca5.zip |
[SPARK-3751] [mllib] DecisionTree: example update + print options
DecisionTreeRunner functionality additions:
* Allow user to pass in a test dataset
* Do not print full model if the model is too large.
As part of this, modify DecisionTreeModel and RandomForestModel to allow printing less info. Proposed updates:
* toString: prints model summary
* toDebugString: prints full model (named after RDD.toDebugString)
Similar update to Python API:
* __repr__() now prints a model summary
* toDebugString() now prints the full model
CC: mengxr chouqin manishamde codedeft Small update (whomever can take a look). Thanks!
Author: Joseph K. Bradley <joseph.kurata.bradley@gmail.com>
Closes #2604 from jkbradley/dtrunner-update and squashes the following commits:
b2b3c60 [Joseph K. Bradley] re-added python sql doc test, temporarily removed before
07b1fae [Joseph K. Bradley] repr() now prints a model summary toDebugString() now prints the full model
1d0d93d [Joseph K. Bradley] Updated DT and RF to print less when toString is called. Added toDebugString for verbose printing.
22eac8c [Joseph K. Bradley] Merge remote-tracking branch 'upstream/master' into dtrunner-update
e007a95 [Joseph K. Bradley] Updated DecisionTreeRunner to accept a test dataset.
Diffstat (limited to 'extras')
0 files changed, 0 insertions, 0 deletions