[SPARK-6055] [PySpark] fix incorrect __eq__ of DataType

The _eq_ of DataType is not correct, class cache is not use correctly (created class can not be find by dataType), then it will create lots of classes (saved in _cached_cls), never released. Also, all same DataType have same hash code, there will be many object in a dict with the same hash code, end with hash attach, it's very slow to access this dict (depends on the implementation of CPython). This PR also improve the performance of inferSchema (avoid the unnecessary converter of object). cc pwendell JoshRosen Author: Davies Liu <davies@databricks.com> Closes #4808 from davies/leak and squashes the following commits: 6a322a4 [Davies Liu] tests refactor 3da44fc [Davies Liu] fix __eq__ of Singleton 534ac90 [Davies Liu] add more checks 46999dc [Davies Liu] fix tests d9ae973 [Davies Liu] fix memory leak in sql
author: Davies Liu <davies@databricks.com> 2015-02-27 20:07:17 -0800
committer: Josh Rosen <joshrosen@databricks.com> 2015-02-27 20:07:17 -0800
commit: e0e64ba4b1b8eb72e856286f756c65fa22ab0a36 (patch)
tree: ca358052d6b572756ecbbf98133a093db3f4cc83 /python/pyspark/sql/dataframe.py
parent: 8c468a6600e0deb5464990df60148212e64fdecd (diff)
download: spark-e0e64ba4b1b8eb72e856286f756c65fa22ab0a36.tar.gz
spark-e0e64ba4b1b8eb72e856286f756c65fa22ab0a36.tar.bz2
spark-e0e64ba4b1b8eb72e856286f756c65fa22ab0a36.zip
1 files changed, 3 insertions, 1 deletions
diff --git a/python/pyspark/sql/dataframe.py b/python/pyspark/sql/dataframe.py
index aec99017fb..5c3b7377c3 100644
--- a/python/pyspark/sql/dataframe.py
+++ b/python/pyspark/sql/dataframe.py
@@ -1025,10 +1025,12 @@ class Column(object):
             ssql_ctx = sc._jvm.SQLContext(sc._jsc.sc())
             jdt = ssql_ctx.parseDataType(dataType.json())
             jc = self._jc.cast(jdt)
+        else:
+            raise TypeError("unexpected type: %s" % type(dataType))
         return Column(jc)
 
     def __repr__(self):
-        return 'Column<%s>' % self._jdf.toString().encode('utf8')
+        return 'Column<%s>' % self._jc.toString().encode('utf8')
 
 
 def _test():
author	Davies Liu <davies@databricks.com>	2015-02-27 20:07:17 -0800
committer	Josh Rosen <joshrosen@databricks.com>	2015-02-27 20:07:17 -0800
commit	e0e64ba4b1b8eb72e856286f756c65fa22ab0a36 (patch)
tree	ca358052d6b572756ecbbf98133a093db3f4cc83 /python/pyspark/sql/dataframe.py
parent	8c468a6600e0deb5464990df60148212e64fdecd (diff)
download	spark-e0e64ba4b1b8eb72e856286f756c65fa22ab0a36.tar.gz spark-e0e64ba4b1b8eb72e856286f756c65fa22ab0a36.tar.bz2 spark-e0e64ba4b1b8eb72e856286f756c65fa22ab0a36.zip