[SPARK-3713][SQL] Uses JSON to serialize DataType objects - spark

diff options

author	Cheng Lian <lian.cs.zju@gmail.com>	2014-10-08 17:04:49 -0700
committer	Michael Armbrust <michael@databricks.com>	2014-10-08 17:04:49 -0700
commit	a42cc08d219c579019f613faa8d310e6069c06fe (patch)
tree	47adb5abf147cd477a88e33524de43b29379b990 /sql/README.md
parent	a85f24accd3266e0f97ee04d03c22b593d99c062 (diff)
download	spark-a42cc08d219c579019f613faa8d310e6069c06fe.tar.gz spark-a42cc08d219c579019f613faa8d310e6069c06fe.tar.bz2 spark-a42cc08d219c579019f613faa8d310e6069c06fe.zip

[SPARK-3713][SQL] Uses JSON to serialize DataType objects

This PR uses JSON instead of `toString` to serialize `DataType`s. The latter is not only hard to parse but also flaky in many cases. Since we already write schema information to Parquet metadata in the old style, we have to reserve the old `DataType` parser and ensure downward compatibility. The old parser is now renamed to `CaseClassStringParser` and moved into `object DataType`. JoshRosen davies Please help review PySpark related changes, thanks! Author: Cheng Lian <lian.cs.zju@gmail.com> Closes #2563 from liancheng/datatype-to-json and squashes the following commits: fc92eb3 [Cheng Lian] Reverts debugging code, simplifies primitive type JSON representation 438c75f [Cheng Lian] Refactors PySpark DataType JSON SerDe per comments 6b6387b [Cheng Lian] Removes debugging code 6a3ee3a [Cheng Lian] Addresses per review comments dc158b5 [Cheng Lian] Addresses PEP8 issues 99ab4ee [Cheng Lian] Adds compatibility est case for Parquet type conversion a983a6c [Cheng Lian] Adds PySpark support f608c6e [Cheng Lian] De/serializes DataType objects from/to JSON

Diffstat (limited to 'sql/README.md')

0 files changed, 0 insertions, 0 deletions


context:
space:
mode: