[SPARK-16698][SQL] Field names having dots should be allowed for datasources based on FileFormat

## What changes were proposed in this pull request? It seems this is a regression assuming from https://issues.apache.org/jira/browse/SPARK-16698. Field name having dots throws an exception. For example the codes below: ```scala val path = "/tmp/path" val json =""" {"a.b":"data"}""" spark.sparkContext .parallelize(json :: Nil) .saveAsTextFile(path) spark.read.json(path).collect() ``` throws an exception as below: ``` Unable to resolve a.b given [a.b]; org.apache.spark.sql.AnalysisException: Unable to resolve a.b given [a.b]; at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolve$1$$anonfun$apply$5.apply(LogicalPlan.scala:134) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolve$1$$anonfun$apply$5.apply(LogicalPlan.scala:134) at scala.Option.getOrElse(Option.scala:121) ``` This problem was introduced in https://github.com/apache/spark/commit/17eec0a71ba8713c559d641e3f43a1be726b037c#diff-27c76f96a7b2733ecfd6f46a1716e153R121 When extracting the data columns, it does not count that it can contains dots in field names. Actually, it seems the fields name are not expected as quoted when defining schema. So, It not have to consider whether this is wrapped with quotes because the actual schema (inferred or user-given schema) would not have the quotes for fields. For example, this throws an exception. (**Loading JSON from RDD is fine**) ```scala val json =""" {"a.b":"data"}""" val rdd = spark.sparkContext.parallelize(json :: Nil) spark.read.schema(StructType(Seq(StructField("`a.b`", StringType, true)))) .json(rdd).select("`a.b`").printSchema() ``` as below: ``` cannot resolve '```a.b```' given input columns: [`a.b`]; org.apache.spark.sql.AnalysisException: cannot resolve '```a.b```' given input columns: [`a.b`]; at org.apache.spark.sql.catalyst.analysis.package$AnalysisErrorAt.failAnalysis(package.scala:42) ``` ## How was this patch tested? Unit tests in `FileSourceStrategySuite`. Author: hyukjinkwon <gurwls223@gmail.com> Closes #14339 from HyukjinKwon/SPARK-16698-regression.
author: hyukjinkwon <gurwls223@gmail.com> 2016-07-25 22:51:30 +0800
committer: Cheng Lian <lian@databricks.com> 2016-07-25 22:51:30 +0800
commit: 79826f3c7936ee27457d030c7115d5cac69befd7 (patch)
tree: 40df96cafba2d7db1690487cb61e61775605db85 /sql/catalyst
parent: d6a52176ade92853f37167ad27631977dc79bc76 (diff)
download: spark-79826f3c7936ee27457d030c7115d5cac69befd7.tar.gz
spark-79826f3c7936ee27457d030c7115d5cac69befd7.tar.bz2
spark-79826f3c7936ee27457d030c7115d5cac69befd7.zip
1 files changed, 1 insertions, 1 deletions
diff --git a/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/LogicalPlan.scala b/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/LogicalPlan.scala
index d0b2b5d7b2..6d7799151d 100644
--- a/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/LogicalPlan.scala
+++ b/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/LogicalPlan.scala
@@ -127,7 +127,7 @@ abstract class LogicalPlan extends QueryPlan[LogicalPlan] with Logging {
    */
   def resolve(schema: StructType, resolver: Resolver): Seq[Attribute] = {
     schema.map { field =>
-      resolveQuoted(field.name, resolver).map {
+      resolve(field.name :: Nil, resolver).map {
         case a: AttributeReference => a
         case other => sys.error(s"can not handle nested schema yet...  plan $this")
       }.getOrElse {
author	hyukjinkwon <gurwls223@gmail.com>	2016-07-25 22:51:30 +0800
committer	Cheng Lian <lian@databricks.com>	2016-07-25 22:51:30 +0800
commit	79826f3c7936ee27457d030c7115d5cac69befd7 (patch)
tree	40df96cafba2d7db1690487cb61e61775605db85 /sql/catalyst
parent	d6a52176ade92853f37167ad27631977dc79bc76 (diff)
download	spark-79826f3c7936ee27457d030c7115d5cac69befd7.tar.gz spark-79826f3c7936ee27457d030c7115d5cac69befd7.tar.bz2 spark-79826f3c7936ee27457d030c7115d5cac69befd7.zip