diff options
author | Xiao Li <gatorsmile@gmail.com> | 2017-04-10 09:15:04 -0700 |
---|---|---|
committer | Xiao Li <gatorsmile@gmail.com> | 2017-04-10 09:15:04 -0700 |
commit | fd711ea13e558f0e7d3e01f08e01444d394499a6 (patch) | |
tree | c4c9815b52cbeec31e06bb873a926171fb44d26c /core/src | |
parent | 5acaf8c0c685e47ec619fbdfd353163721e1cf50 (diff) | |
download | spark-fd711ea13e558f0e7d3e01f08e01444d394499a6.tar.gz spark-fd711ea13e558f0e7d3e01f08e01444d394499a6.tar.bz2 spark-fd711ea13e558f0e7d3e01f08e01444d394499a6.zip |
[SPARK-20273][SQL] Disallow Non-deterministic Filter push-down into Join Conditions
## What changes were proposed in this pull request?
```
sql("SELECT t1.b, rand(0) as r FROM cachedData, cachedData t1 GROUP BY t1.b having r > 0.5").show()
```
We will get the following error:
```
Job aborted due to stage failure: Task 1 in stage 4.0 failed 1 times, most recent failure: Lost task 1.0 in stage 4.0 (TID 8, localhost, executor driver): java.lang.NullPointerException
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate.eval(Unknown Source)
at org.apache.spark.sql.execution.joins.BroadcastNestedLoopJoinExec$$anonfun$org$apache$spark$sql$execution$joins$BroadcastNestedLoopJoinExec$$boundCondition$1.apply(BroadcastNestedLoopJoinExec.scala:87)
at org.apache.spark.sql.execution.joins.BroadcastNestedLoopJoinExec$$anonfun$org$apache$spark$sql$execution$joins$BroadcastNestedLoopJoinExec$$boundCondition$1.apply(BroadcastNestedLoopJoinExec.scala:87)
at scala.collection.Iterator$$anon$13.hasNext(Iterator.scala:463)
```
Filters could be pushed down to the join conditions by the optimizer rule `PushPredicateThroughJoin`. However, Analyzer [blocks users to add non-deterministics conditions](https://github.com/apache/spark/blob/master/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckAnalysis.scala#L386-L395) (For details, see the PR https://github.com/apache/spark/pull/7535).
We should not push down non-deterministic conditions; otherwise, we need to explicitly initialize the non-deterministic expressions. This PR is to simply block it.
### How was this patch tested?
Added a test case
Author: Xiao Li <gatorsmile@gmail.com>
Closes #17585 from gatorsmile/joinRandCondition.
Diffstat (limited to 'core/src')
0 files changed, 0 insertions, 0 deletions