spark - Mirror of Apache Spark

	Commit message (Collapse)	Author	Age	Files	Lines
*	Preparing Spark release v1.2.0-rc2	Patrick Wendell	2014-12-04	4	-4/+4
\|
*	Revert "Preparing Spark release v1.2.0-rc1"	Patrick Wendell	2014-12-04	4	-4/+4
\| \| \| \|	This reverts commit 1056e9ec13203d0c51564265e94d77a054498fdb.
*	Revert "Preparing development version 1.2.1-SNAPSHOT"	Patrick Wendell	2014-12-04	4	-4/+4
\| \| \| \|	This reverts commit 00316cc87983b844f6603f351a8f0b84fe1f6035.
*	[SQL] Minor: Avoid calling Seq#size in a loop	Aaron Davidson	2014-12-04	1	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \|	Just found this instance while doing some jstack-based profiling of a Spark SQL job. It is very unlikely that this is causing much of a perf issue anywhere, but it is unnecessarily suboptimal. Author: Aaron Davidson <aaron@databricks.com> Closes #3593 from aarondav/seq-opt and squashes the following commits: 962cdfc [Aaron Davidson] [SQL] Minor: Avoid calling Seq#size in a loop (cherry picked from commit c6c7165e7ecf1690027d6bd4e0620012cd0d2310) Signed-off-by: Reynold Xin <rxin@databricks.com>
*	[SPARK-4552][SQL] Avoid exception when reading empty parquet data through Hive	Michael Armbrust	2014-12-03	3	-45/+62
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	This is a very small fix that catches one specific exception and returns an empty table. #3441 will address this in a more principled way. Author: Michael Armbrust <michael@databricks.com> Closes #3586 from marmbrus/fixEmptyParquet and squashes the following commits: 2781d9f [Michael Armbrust] Handle empty lists for newParquet 04dd376 [Michael Armbrust] Avoid exception when reading empty parquet data through Hive (cherry picked from commit 513ef82e85661552e596d0b483b645ac24e86d4d) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4695][SQL] Get result using executeCollect	wangfei	2014-12-02	1	-1/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Using ```executeCollect``` to collect the result, because executeCollect is a custom implementation of collect in spark sql which better than rdd's collect Author: wangfei <wangfei1@huawei.com> Closes #3547 from scwf/executeCollect and squashes the following commits: a5ab68e [wangfei] Revert "adding debug info" a60d680 [wangfei] fix test failure 0db7ce8 [wangfei] adding debug info 184c594 [wangfei] using executeCollect instead collect (cherry picked from commit 3ae0cda83c5106136e90d59c20e61db345a5085f) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4670] [SQL] wrong symbol for bitwise not	Daoyuan Wang	2014-12-02	2	-10/+25
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	We should use `~` instead of `-` for bitwise NOT. Author: Daoyuan Wang <daoyuan.wang@intel.com> Closes #3528 from adrian-wang/symbol and squashes the following commits: affd4ad [Daoyuan Wang] fix code gen test case 56efb79 [Daoyuan Wang] ensure bitwise NOT over byte and short persist data type f55fbae [Daoyuan Wang] wrong symbol for bitwise not (cherry picked from commit 1f5ddf17e831ad9717f0f4b60a727a3381fad4f9) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4593][SQL] Return null when denominator is 0	Daoyuan Wang	2014-12-02	4	-5/+83
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	SELECT max(1/0) FROM src would return a very large number, which is obviously not right. For hive-0.12, hive would return `Infinity` for 1/0, while for hive-0.13.1, it is `NULL` for 1/0. I think it is better to keep our behavior with newer Hive version. This PR ensures that when the divider is 0, the result of expression should be NULL, same with hive-0.13.1 Author: Daoyuan Wang <daoyuan.wang@intel.com> Closes #3443 from adrian-wang/div and squashes the following commits: 2e98677 [Daoyuan Wang] fix code gen for divide 0 85c28ba [Daoyuan Wang] temp 36236a5 [Daoyuan Wang] add test cases 6f5716f [Daoyuan Wang] fix comments cee92bd [Daoyuan Wang] avoid evaluation 2 times 22ecd9a [Daoyuan Wang] fix style cf28c58 [Daoyuan Wang] divide fix 2dfe50f [Daoyuan Wang] return null when divider is 0 of Double type (cherry picked from commit f6df609dcc4f4a18c0f1c74b1ae0800cf09fa7ae) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4676][SQL] JavaSchemaRDD.schema may throw NullType MatchError if sql ↵	YanTangZhai	2014-12-02	5	-0/+59
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	has null val jsc = new org.apache.spark.api.java.JavaSparkContext(sc) val jhc = new org.apache.spark.sql.hive.api.java.JavaHiveContext(jsc) val nrdd = jhc.hql("select null from spark_test.for_test") println(nrdd.schema) Then the error is thrown as follows: scala.MatchError: NullType (of class org.apache.spark.sql.catalyst.types.NullType$) at org.apache.spark.sql.types.util.DataTypeConversions$.asJavaDataType(DataTypeConversions.scala:43) Author: YanTangZhai <hakeemzhai@tencent.com> Author: yantangzhai <tyz0303@163.com> Author: Michael Armbrust <michael@databricks.com> Closes #3538 from YanTangZhai/MatchNullType and squashes the following commits: e052dff [yantangzhai] [SPARK-4676] [SQL] JavaSchemaRDD.schema may throw NullType MatchError if sql has null 4b4bb34 [yantangzhai] [SPARK-4676] [SQL] JavaSchemaRDD.schema may throw NullType MatchError if sql has null 896c7b7 [yantangzhai] fix NullType MatchError in JavaSchemaRDD when sql has null 6e643f8 [YanTangZhai] Merge pull request #11 from apache/master e249846 [YanTangZhai] Merge pull request #10 from apache/master d26d982 [YanTangZhai] Merge pull request #9 from apache/master 76d4027 [YanTangZhai] Merge pull request #8 from apache/master 03b62b0 [YanTangZhai] Merge pull request #7 from apache/master 8a00106 [YanTangZhai] Merge pull request #6 from apache/master cbcba66 [YanTangZhai] Merge pull request #3 from apache/master cdef539 [YanTangZhai] Merge pull request #1 from apache/master (cherry picked from commit 10664276007beca3843638e558f504cad44b1fb3) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4663][sql]add finally to avoid resource leak	baishuo	2014-12-02	1	-4/+7
\| \| \| \| \| \| \| \| \| \| \| \| \|	Author: baishuo <vc_java@hotmail.com> Closes #3526 from baishuo/master-trycatch and squashes the following commits: d446e14 [baishuo] correct the code style b36bf96 [baishuo] correct the code style ae0e447 [baishuo] add finally to avoid resource leak (cherry picked from commit 69b6fed206565ecb0173d3757bcb5110422887c3) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4536][SQL] Add sqrt and abs to Spark SQL DSL	Kousuke Saruta	2014-12-02	4	-1/+74
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Spark SQL has embeded sqrt and abs but DSL doesn't support those functions. Author: Kousuke Saruta <sarutak@oss.nttdata.co.jp> Closes #3401 from sarutak/dsl-missing-operator and squashes the following commits: 07700cf [Kousuke Saruta] Modified Literal(null, NullType) to Literal(null) in DslQuerySuite 8f366f8 [Kousuke Saruta] Merge branch 'master' of git://git.apache.org/spark into dsl-missing-operator 1b88e2e [Kousuke Saruta] Merge branch 'master' of git://git.apache.org/spark into dsl-missing-operator 0396f89 [Kousuke Saruta] Added sqrt and abs to Spark SQL DSL (cherry picked from commit e75e04f980281389b881df76f59ba1adc6338629) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4529] [SQL] support view with column alias	Daoyuan Wang	2014-12-01	2	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Support view definition like CREATE VIEW view3(valoo) TBLPROPERTIES ("fear" = "factor") AS SELECT upper(value) FROM src WHERE key=86; [valoo as the alias of upper(value)]. This is missing part of SPARK-4239, for a fully view support. Author: Daoyuan Wang <daoyuan.wang@intel.com> Closes #3396 from adrian-wang/viewcolumn and squashes the following commits: 4d001d0 [Daoyuan Wang] support view with column alias (cherry picked from commit 4df60a8cbc58f2877787245c2a83b2de85579c82) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SQL] Minor fix for doc and comment	wangfei	2014-12-01	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \|	Author: wangfei <wangfei1@huawei.com> Closes #3533 from scwf/sql-doc1 and squashes the following commits: 962910b [wangfei] doc and comment fix (cherry picked from commit 7b79957879db4dfcc7c3601cb40ac4fd576259a5) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4658][SQL] Code documentation issue in DDL of datasource API	ravipesala	2014-12-01	2	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \|	Author: ravipesala <ravindra.pesala@huawei.com> Closes #3516 from ravipesala/ddl_doc and squashes the following commits: d101fdf [ravipesala] Style issues fixed d2238cd [ravipesala] Corrected documentation (cherry picked from commit bc353819cc86c3b0ad75caf81b47744bfc2aeeb3) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4650][SQL] Supporting multi column support in countDistinct function ↵	ravipesala	2014-12-01	2	-1/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	like count(distinct c1,c2..) in Spark SQL Supporting multi column support in countDistinct function like count(distinct c1,c2..) in Spark SQL Author: ravipesala <ravindra.pesala@huawei.com> Author: Michael Armbrust <michael@databricks.com> Closes #3511 from ravipesala/countdistinct and squashes the following commits: cc4dbb1 [ravipesala] style 070e12a [ravipesala] Supporting multi column support in count(distinct c1,c2..) in Spark SQL (cherry picked from commit 6a9ff19dc06745144d5b311d4f87073c81d53a8f) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4358][SQL] Let BigDecimal do checking type compatibility	Liang-Chi Hsieh	2014-12-01	1	-8/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Remove hardcoding max and min values for types. Let BigDecimal do checking type compatibility. Author: Liang-Chi Hsieh <viirya@gmail.com> Closes #3208 from viirya/more_numericLit and squashes the following commits: e9834b4 [Liang-Chi Hsieh] Remove byte and short types for number literal. 1bd1825 [Liang-Chi Hsieh] Fix Indentation and make the modification clearer. cf1a997 [Liang-Chi Hsieh] Modified for comment to add a rule of analysis that adds a cast. 91fe489 [Liang-Chi Hsieh] add Byte and Short. 1bdc69d [Liang-Chi Hsieh] Let BigDecimal do checking type compatibility. (cherry picked from commit b57365a1ec89e31470f424ff37d5ebc7c90a39d8) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SQL] add @group tab in limit() and count()	Jacky Li	2014-12-01	1	-0/+4
\| \| \| \| \| \| \| \| \| \| \| \| \|	group tab is missing for scaladoc Author: Jacky Li <jacky.likun@gmail.com> Closes #3458 from jackylk/patch-7 and squashes the following commits: 0121a70 [Jacky Li] add @group tab in limit() and count() (cherry picked from commit bafee67ebad01f7aea2cd393a70b57eb8345eeb0) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4661][Core] Minor code and docs cleanup	zsxwing	2014-12-01	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \|	Author: zsxwing <zsxwing@gmail.com> Closes #3521 from zsxwing/SPARK-4661 and squashes the following commits: 03cbe3f [zsxwing] Minor code and docs cleanup (cherry picked from commit 30a86acdefd5428af6d6264f59a037e0eefd74b4) Signed-off-by: Reynold Xin <rxin@databricks.com>
*	Preparing development version 1.2.1-SNAPSHOT	Patrick Wendell	2014-11-28	4	-4/+4
\|
*	Preparing Spark release v1.2.0-rc1	Patrick Wendell	2014-11-28	4	-4/+4
\|
*	Revert "Preparing Spark release v1.2.0-rc1"	Patrick Wendell	2014-11-28	4	-4/+4
\| \| \| \|	This reverts commit 39c7d1c1f9a7785285cf4c20dfbffd96f72d5634.
*	Revert "Preparing development version 1.2.1-SNAPSHOT"	Patrick Wendell	2014-11-28	4	-4/+4
\| \| \| \|	This reverts commit fc7bff00ac731d2632213a98cd92dc5e84ce7dcd.
*	Preparing development version 1.2.1-SNAPSHOT	Patrick Wendell	2014-11-28	4	-4/+4
\|
*	Preparing Spark release v1.2.0-rc1	Patrick Wendell	2014-11-28	4	-4/+4
\|
*	[SPARK-4645][SQL] Disables asynchronous execution in Hive 0.13.1 ↵	Cheng Lian	2014-11-28	1	-100/+39
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	HiveThriftServer2 This PR disables HiveThriftServer2 asynchronous execution by setting `runInBackground` argument in `ExecuteStatementOperation` to `false`, and reverting `SparkExecuteStatementOperation.run` in Hive 13 shim to Hive 12 version. This change makes Simba ODBC driver v1.0.0.1000 work. <!-- Reviewable:start --> [<img src="https://reviewable.io/review_button.png" height=40 alt="Review on Reviewable"/>](https://reviewable.io/reviews/apache/spark/3506) <!-- Reviewable:end --> Author: Cheng Lian <lian@databricks.com> Closes #3506 from liancheng/disable-async-exec and squashes the following commits: 593804d [Cheng Lian] Disables asynchronous execution in Hive 0.13.1 HiveThriftServer2
*	[SPARK-4308][SQL] Sets SQL operation state to ERROR when exception is thrown	Cheng Lian	2014-11-28	3	-29/+21
\| \| \| \| \| \| \| \| \| \|	In `HiveThriftServer2`, when an exception is thrown during a SQL execution, the SQL operation state should be set to `ERROR`, but now it remains `RUNNING`. This affects the result of the `GetOperationStatus` Thrift API. Author: Cheng Lian <lian@databricks.com> Closes #3175 from liancheng/fix-op-state and squashes the following commits: 6d4c1fe [Cheng Lian] Sets SQL operation state to ERROR when exception is thrown
*	Revert "Preparing Spark release v1.2.0-rc1"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit cc2c05e4ee81d2f34873a2ebb9a5272867cb65c2.
*	Revert "Preparing development version 1.2.1-SNAPSHOT"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit 380eba5f49eca1dbd4084e6c84e19866fffd4efa.
*	Preparing development version 1.2.1-SNAPSHOT	Patrick Wendell	2014-11-26	4	-4/+4
\|
*	Preparing Spark release v1.2.0-rc1	Patrick Wendell	2014-11-26	4	-4/+4
\|
*	Revert "Preparing Spark release v1.2.0-rc1"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit 5247dd859b95a440baa562b9827bdeb26aa6530e.
*	Revert "Preparing development version 1.2.1-SNAPSHOT"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit 79df6b43ae762263a8120f423ddb4a0811dd4b6f.
*	Preparing development version 1.2.1-SNAPSHOT	Patrick Wendell	2014-11-26	4	-4/+4
\|
*	Preparing Spark release v1.2.0-rc1	Patrick Wendell	2014-11-26	4	-4/+4
\|
*	Revert "Preparing Spark release v1.2.0-rc1"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit db7f4a898af22a02b36428507f8ef2b429d78dc1.
*	Revert "Preparing development version 1.2.1-SNAPSHOT"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit d7b1ecb25676d228deb6fe05efdb4e2ab9c3e30b.
*	Preparing development version 1.2.1-SNAPSHOT	Ubuntu	2014-11-26	4	-4/+4
\|
*	Preparing Spark release v1.2.0-rc1	Ubuntu	2014-11-26	4	-4/+4
\|
*	Revert "Preparing Spark release v1.2.0-snapshot1"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit 38c1fbd9694430cefd962c90bc36b0d108c6124b.
*	Revert "Preparing development version 1.2.1-SNAPSHOT"	Patrick Wendell	2014-11-26	4	-4/+4
\| \| \| \|	This reverts commit d7ac6013483e83caff8ea54c228f37aeca159db8.
*	[SQL] Compute timeTaken correctly	w00228970	2014-11-24	1	-7/+4
\| \| \| \| \| \| \| \| \| \| \| \| \|	```timeTaken``` should not count the time of printing result. Author: w00228970 <wangfei1@huawei.com> Closes #3423 from scwf/time-taken-bug and squashes the following commits: da7e102 [w00228970] compute time taken correctly (cherry picked from commit 723be60e233d0f85944d948efd06845ef546c9f5) Signed-off-by: Reynold Xin <rxin@databricks.com>
*	[SPARK-4548] []SPARK-4517] improve performance of python broadcast	Davies Liu	2014-11-24	2	-3/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Re-implement the Python broadcast using file: 1) serialize the python object using cPickle, write into disks. 2) Create a wrapper in JVM (for the dumped file), it read data from during serialization 3) Using TorrentBroadcast or HttpBroadcast to transfer the data (compressed) into executors 4) During deserialization, writing the data into disk. 5) Passing the path into Python worker, read data from disk and unpickle it into python object, until the first access. It fixes the performance regression introduced in #2659, has similar performance as 1.1, but support object larger than 2G, also improve the memory efficiency (only one compressed copy in driver and executor). Testing with a 500M broadcast and 4 tasks (excluding the benefit from reused worker in 1.2): name \| 1.1 \| 1.2 with this patch \| improvement ---------\|--------\|---------\|-------- python-broadcast-w-bytes \| 25.20 \| 9.33 \| 170.13% \| python-broadcast-w-set \| 4.13 \| 4.50 \| -8.35% \| Testing with 100 tasks (16 CPUs): name \| 1.1 \| 1.2 with this patch \| improvement ---------\|--------\|---------\|-------- python-broadcast-w-bytes \| 38.16 \| 8.40 \| 353.98% python-broadcast-w-set \| 23.29 \| 9.59 \| 142.80% Author: Davies Liu <davies@databricks.com> Closes #3417 from davies/pybroadcast and squashes the following commits: 50a58e0 [Davies Liu] address comments b98de1d [Davies Liu] disable gc while unpickle e5ee6b9 [Davies Liu] support large string 09303b8 [Davies Liu] read all data into memory dde02dd [Davies Liu] improve performance of python broadcast (cherry picked from commit 6cf507685efd01df77d663145ae08e48c7f92948) Signed-off-by: Josh Rosen <joshrosen@databricks.com>
*	[SPARK-4487][SQL] Fix attribute reference resolution error when using ORDER BY.	Kousuke Saruta	2014-11-24	2	-1/+8
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When we use ORDER BY clause, at first, attributes referenced by projection are resolved (1). And then, attributes referenced at ORDER BY clause are resolved (2). But when resolving attributes referenced at ORDER BY clause, the resolution result generated in (1) is discarded so for example, following query fails. SELECT c1 + c2 FROM mytable ORDER BY c1; The query above fails because when resolving the attribute reference 'c1', the resolution result of 'c2' is discarded. Author: Kousuke Saruta <sarutak@oss.nttdata.co.jp> Closes #3363 from sarutak/SPARK-4487 and squashes the following commits: fd314f3 [Kousuke Saruta] Fixed attribute resolution logic in Analyzer 6e60c20 [Kousuke Saruta] Fixed conflicts cb5b7e9 [Kousuke Saruta] Added test case for SPARK-4487 282d529 [Kousuke Saruta] Fixed attributes reference resolution error b6123e6 [Kousuke Saruta] Merge branch 'master' of git://git.apache.org/spark into concat-feature 317b7fb [Kousuke Saruta] WIP (cherry picked from commit dd1c9cb36cde8202cede8014b5641ae8a0197812) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SQL] Fix comment in HiveShim	Daniel Darabos	2014-11-24	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \| \|	This file is for Hive 0.13.1 I think. Author: Daniel Darabos <darabos.daniel@gmail.com> Closes #3432 from darabos/patch-2 and squashes the following commits: 4fd22ed [Daniel Darabos] Fix comment. This file is for Hive 0.13.1. (cherry picked from commit d5834f0732b586731034a7df5402c25454770fc5) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4479][SQL] Avoids unnecessary defensive copies when sort based ↵	Cheng Lian	2014-11-24	1	-1/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	shuffle is on This PR is a workaround for SPARK-4479. Two changes are introduced: when merge sort is bypassed in `ExternalSorter`, 1. also bypass RDD elements buffering as buffering is the reason that `MutableRow` backed row objects must be copied, and 2. avoids defensive copies in `Exchange` operator <!-- Reviewable:start --> [<img src="https://reviewable.io/review_button.png" height=40 alt="Review on Reviewable"/>](https://reviewable.io/reviews/apache/spark/3422) <!-- Reviewable:end --> Author: Cheng Lian <lian@databricks.com> Closes #3422 from liancheng/avoids-defensive-copies and squashes the following commits: 591f2e9 [Cheng Lian] Passes all shuffle suites 0c3c91e [Cheng Lian] Fixes shuffle write metrics when merge sort is bypassed ed5df3c [Cheng Lian] Fixes styling changes f75089b [Cheng Lian] Avoids unnecessary defensive copies when sort based shuffle is on (cherry picked from commit a6d7b61f92dc7c1f9632cecb232afa8040ab2b4d) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4522][SQL] Parse schema with missing metadata.	Michael Armbrust	2014-11-20	1	-0/+6
\| \| \| \| \| \| \| \| \| \| \| \| \|	This is just a quick fix for 1.2. SPARK-4523 describes a more complete solution. Author: Michael Armbrust <michael@databricks.com> Closes #3392 from marmbrus/parquetMetadata and squashes the following commits: bcc6626 [Michael Armbrust] Parse schema with missing metadata. (cherry picked from commit 90a6a46bd11030672597f015dd443d954107123a) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4413][SQL] Parquet support through datasource API	Michael Armbrust	2014-11-20	5	-79/+458
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Goals: - Support for accessing parquet using SQL but not requiring Hive (thus allowing support of parquet tables with decimal columns) - Support for folder based partitioning with automatic discovery of available partitions - Caching of file metadata See scaladoc of `ParquetRelation2` for more details. Author: Michael Armbrust <michael@databricks.com> Closes #3269 from marmbrus/newParquet and squashes the following commits: 1dd75f1 [Michael Armbrust] Pass all paths for FileInputFormat at once. 645768b [Michael Armbrust] Review comments. abd8e2f [Michael Armbrust] Alternative implementation of parquet based on the datasources API. 938019e [Michael Armbrust] Add an experimental interface to data sources that exposes catalyst expressions. e9d2641 [Michael Armbrust] logging / formatting improvements. (cherry picked from commit 02ec058efe24348cdd3691b55942e6f0ef138732) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-4244] [SQL] Support Hive Generic UDFs with constant object inspector ↵	Cheng Hao	2014-11-20	4	-8/+17
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	parameters Query `SELECT named_struct(lower("AA"), "12", lower("Bb"), "13") FROM src LIMIT 1` will throw exception, some of the Hive Generic UDF/UDAF requires the input object inspector is `ConstantObjectInspector`, however, we won't get that before the expression optimization executed. (Constant Folding). This PR is a work around to fix this. (As ideally, the `output` of LogicalPlan should be identical before and after Optimization). Author: Cheng Hao <hao.cheng@intel.com> Closes #3109 from chenghao-intel/optimized and squashes the following commits: 487ff79 [Cheng Hao] rebase to the latest master & update the unittest (cherry picked from commit 84d79ee9ec47465269f7b0a7971176da93c96f3f) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SQL] fix function description mistake	Jacky Li	2014-11-20	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \| \|	Sample code in the description of SchemaRDD.where is not correct Author: Jacky Li <jacky.likun@gmail.com> Closes #3344 from jackylk/patch-6 and squashes the following commits: 62cd126 [Jacky Li] [SQL] fix function description mistake (cherry picked from commit ad5f1f3ca240473261162c06ffc5aa70d15a5991) Signed-off-by: Michael Armbrust <michael@databricks.com>
*	[SPARK-2918] [SQL] Support the CTAS in EXPLAIN command	Cheng Hao	2014-11-20	2	-1/+41
\| \| \| \| \| \| \| \| \| \| \| \| \|	Hive supports the `explain` the CTAS, which was supported by Spark SQL previously, however, seems it was reverted after the code refactoring in HiveQL. Author: Cheng Hao <hao.cheng@intel.com> Closes #3357 from chenghao-intel/explain and squashes the following commits: 7aace63 [Cheng Hao] Support the CTAS in EXPLAIN command (cherry picked from commit 6aa0fc9f4d95f09383cbcb5f79166c60697e6683) Signed-off-by: Michael Armbrust <michael@databricks.com>