[SPARK-16839][SQL] redundant aliases after cleanupAliases

## What changes were proposed in this pull request? Simplify struct creation, especially the aspect of `CleanupAliases` which missed some aliases when handling trees created by `CreateStruct`. This PR includes: 1. A failing test (create struct with nested aliases, some of the aliases survive `CleanupAliases`). 2. A fix that transforms `CreateStruct` into a `CreateNamedStruct` constructor, effectively eliminating `CreateStruct` from all expression trees. 3. A `NamePlaceHolder` used by `CreateStruct` when column names cannot be extracted from unresolved `NamedExpression`. 4. A new Analyzer rule that resolves `NamePlaceHolder` into a string literal once the `NamedExpression` is resolved. 5. `CleanupAliases` code was simplified as it no longer has to deal with `CreateStruct`'s top level columns. ## How was this patch tested? running all tests-suits in package org.apache.spark.sql, especially including the analysis suite, making sure added test initially fails, after applying suggested fix rerun the entire analysis package successfully. modified few tests that expected `CreateStruct` which is now transformed into `CreateNamedStruct`. Credit goes to hvanhovell for assisting with this PR. Author: eyal farago <eyal farago> Author: eyal farago <eyal.farago@gmail.com> Author: Herman van Hovell <hvanhovell@databricks.com> Author: Eyal Farago <eyal.farago@actimize.com> Author: Hyukjin Kwon <gurwls223@gmail.com> Author: eyalfa <eyal.farago@gmail.com> Closes #14444 from eyalfa/SPARK-16839_redundant_aliases_after_cleanupAliases.
author: eyal farago <eyal farago> 2016-11-01 17:12:20 +0100
committer: Herman van Hovell <hvanhovell@databricks.com> 2016-11-01 17:12:20 +0100
commit: 5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975 (patch)
tree: 54c00373573510b26265761cdb86846c01a704a2 /R/pkg/inst/tests
parent: f7c145d8ce14b23019099c509d5a2b6dfb1fe62c (diff)
download: spark-5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975.tar.gz
spark-5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975.tar.bz2
spark-5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975.zip
1 files changed, 6 insertions, 6 deletions
diff --git a/R/pkg/inst/tests/testthat/test_sparkSQL.R b/R/pkg/inst/tests/testthat/test_sparkSQL.R
index 9289db57b6..5002655fc0 100644
--- a/R/pkg/inst/tests/testthat/test_sparkSQL.R
+++ b/R/pkg/inst/tests/testthat/test_sparkSQL.R
@@ -1222,16 +1222,16 @@ test_that("column functions", {
   # Test struct()
   df <- createDataFrame(list(list(1L, 2L, 3L), list(4L, 5L, 6L)),
                         schema = c("a", "b", "c"))
-  result <- collect(select(df, struct("a", "c")))
+  result <- collect(select(df, alias(struct("a", "c"), "d")))
   expected <- data.frame(row.names = 1:2)
-  expected$"struct(a, c)" <- list(listToStruct(list(a = 1L, c = 3L)),
-                                 listToStruct(list(a = 4L, c = 6L)))
+  expected$"d" <- list(listToStruct(list(a = 1L, c = 3L)),
+                      listToStruct(list(a = 4L, c = 6L)))
   expect_equal(result, expected)
 
-  result <- collect(select(df, struct(df$a, df$b)))
+  result <- collect(select(df, alias(struct(df$a, df$b), "d")))
   expected <- data.frame(row.names = 1:2)
-  expected$"struct(a, b)" <- list(listToStruct(list(a = 1L, b = 2L)),
-                                 listToStruct(list(a = 4L, b = 5L)))
+  expected$"d" <- list(listToStruct(list(a = 1L, b = 2L)),
+                      listToStruct(list(a = 4L, b = 5L)))
   expect_equal(result, expected)
 
   # Test encode(), decode()
author	eyal farago <eyal farago>	2016-11-01 17:12:20 +0100
committer	Herman van Hovell <hvanhovell@databricks.com>	2016-11-01 17:12:20 +0100
commit	5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975 (patch)
tree	54c00373573510b26265761cdb86846c01a704a2 /R/pkg/inst/tests
parent	f7c145d8ce14b23019099c509d5a2b6dfb1fe62c (diff)
download	spark-5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975.tar.gz spark-5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975.tar.bz2 spark-5441a6269e00e3903ae6c1ea8deb4ddf3d2e9975.zip