spark - Mirror of Apache Spark

	Commit message (Collapse)	Author	Age	Files	Lines
*	[MINOR][DOCS] Use multi-line JavaDoc comments in Scala code.	Dongjoon Hyun	2016-04-02	3	-8/+8
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR aims to fix all Scala-Style multiline comments into Java-Style multiline comments in Scala codes. (All comment-only changes over 77 files: +786 lines, −747 lines) ## How was this patch tested? Manual. Author: Dongjoon Hyun <dongjoon@apache.org> Closes #12130 from dongjoon-hyun/use_multiine_javadoc_comments.
*	[HOTFIX] Disable StateStoreSuite.maintenance	Reynold Xin	2016-04-02	1	-1/+1
\|
*	[HOTFIX] Fix compilation break.	Reynold Xin	2016-04-02	2	-5/+4
\|
*	[MINOR][SQL] Fix comments styl and correct several styles and nits in CSV ↵	hyukjinkwon	2016-04-01	1	-5/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	data source ## What changes were proposed in this pull request? While trying to create a PR (which was not an issue at the end), I just corrected some style nits. So, I removed the changes except for some coding style corrections. - According to the [scala-style-guide#documentation-style](https://github.com/databricks/scala-style-guide#documentation-style), Scala style comments are discouraged. >```scala >/** This is a correct one-liner, short description. / > >/* > * This is correct multi-line JavaDoc comment. And > * this is my second line, and if I keep typing, this would be > * my third line. > / > >/* In Spark, we don't use the ScalaDoc style so this > * is not correct. > */ >``` - Double newlines between consecutive methods was removed. According to [scala-style-guide#blank-lines-vertical-whitespace](https://github.com/databricks/scala-style-guide#blank-lines-vertical-whitespace), single newline appears when >Between consecutive members (or initializers) of a class: fields, constructors, methods, nested classes, static initializers, instance initializers. - Remove uesless parentheses in tests - Use `mapPartitions` instead of `mapPartitionsWithIndex()`. ## How was this patch tested? Unit tests were used and `dev/run_tests` for style tests. Author: hyukjinkwon <gurwls223@gmail.com> Closes #12109 from HyukjinKwon/SPARK-14271.
*	[SPARK-14285][SQL] Implement common type-safe aggregate functions	Reynold Xin	2016-04-01	3	-101/+140
\| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? In the Dataset API, it is fairly difficult for users to perform simple aggregations in a type-safe way at the moment because there are no aggregators that have been implemented. This pull request adds a few common aggregate functions in expressions.scala.typed package, and also creates the expressions.java.typed package without implementation. The java implementation should probably come as a separate pull request. One challenge there is to resolve the type difference between Scala primitive types and Java boxed types. ## How was this patch tested? Added unit tests for them. Author: Reynold Xin <rxin@databricks.com> Closes #12077 from rxin/SPARK-14285.
*	[SPARK-14251][SQL] Add SQL command for printing out generated code for debugging	Dongjoon Hyun	2016-04-01	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR implements `EXPLAIN CODEGEN` SQL command which returns generated codes like `debugCodegen`. In `spark-shell`, we don't need to `import debug` module. In `spark-sql`, we can use this SQL command now. Before ``` scala> import org.apache.spark.sql.execution.debug._ scala> sql("select 'a' as a group by 1").debugCodegen() Found 2 WholeStageCodegen subtrees. == Subtree 1 / 2 == ... Generated code: ... == Subtree 2 / 2 == ... Generated code: ... ``` After ``` scala> sql("explain extended codegen select 'a' as a group by 1").collect().foreach(println) [Found 2 WholeStageCodegen subtrees.] [== Subtree 1 / 2 ==] ... [] [Generated code:] ... [] [== Subtree 2 / 2 ==] ... [] [Generated code:] ... ``` ## How was this patch tested? Pass the Jenkins tests (including new testcases) Author: Dongjoon Hyun <dongjoon@apache.org> Closes #12099 from dongjoon-hyun/SPARK-14251.
*	[SPARK-14138] [SQL] [MASTER] Fix generated SpecificColumnarIterator code can ↵	Kazuaki Ishizaki	2016-04-01	1	-0/+10
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	exceed JVM size limit for cached DataFrames ## What changes were proposed in this pull request? This PR reduces Java byte code size of method in ```SpecificColumnarIterator``` by using a approach to make a group for lot of ```ColumnAccessor``` instantiations or method calls (more than 200) into a method ## How was this patch tested? Added a new unit test, which includes large instantiations and method calls, to ```InMemoryColumnarQuerySuite``` Author: Kazuaki Ishizaki <ishizaki@jp.ibm.com> Closes #12108 from kiszk/SPARK-14138-master.
*	[SPARK-14255][SQL] Streaming Aggregation	Michael Armbrust	2016-04-01	5	-113/+278
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This PR adds the ability to perform aggregations inside of a `ContinuousQuery`. In order to implement this feature, the planning of aggregation has augmented with a new `StatefulAggregationStrategy`. Unlike batch aggregation, stateful-aggregation uses the `StateStore` (introduced in #11645) to persist the results of partial aggregation across different invocations. The resulting physical plan performs the aggregation using the following progression: - Partial Aggregation - Shuffle - Partial Merge (now there is at most 1 tuple per group) - StateStoreRestore (now there is 1 tuple from this batch + optionally one from the previous) - Partial Merge (now there is at most 1 tuple per group) - StateStoreSave (saves the tuple for the next batch) - Complete (output the current result of the aggregation) The following refactoring was also performed to allow us to plug into existing code: - The get/put implementation is taken from #12013 - The logic for breaking down and de-duping the physical execution of aggregation has been move into a new pattern `PhysicalAggregation` - The `AttributeReference` used to identify the result of an `AggregateFunction` as been moved into the `AggregateExpression` container. This change moves the reference into the same object as the other intermediate references used in aggregation and eliminates the need to pass around a `Map[(AggregateFunction, Boolean), Attribute]`. Further clean up (using a different aggregation container for logical/physical plans) is deferred to a followup. - Some planning logic is moved from the `SessionState` into the `QueryExecution` to make it easier to override in the streaming case. - The ability to write a `StreamTest` that checks only the output of the last batch has been added to simulate the future addition of output modes. Author: Michael Armbrust <michael@databricks.com> Closes #12048 from marmbrus/statefulAgg.
*	[SPARK-14316][SQL] StateStoreCoordinator should extend ThreadSafeRpcEndpoint	Shixiong Zhu	2016-04-01	1	-5/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? RpcEndpoint is not thread safe and allows multiple messages to be processed at the same time. StateStoreCoordinator should use ThreadSafeRpcEndpoint. ## How was this patch tested? Existing unit tests. Author: Shixiong Zhu <shixiong@databricks.com> Closes #12100 from zsxwing/fix-StateStoreCoordinator.
*	[SPARK-13674] [SQL] Add wholestage codegen support to Sample	Liang-Chi Hsieh	2016-04-01	1	-0/+25
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	JIRA: https://issues.apache.org/jira/browse/SPARK-13674 ## What changes were proposed in this pull request? Sample operator doesn't support wholestage codegen now. This pr is to add support to it. ## How was this patch tested? A test is added into `BenchmarkWholeStageCodegen`. Besides, all tests should be passed. Author: Liang-Chi Hsieh <simonh@tw.ibm.com> Author: Liang-Chi Hsieh <viirya@gmail.com> Closes #11517 from viirya/add-wholestage-sample.
*	[SPARK-14160] Time Windowing functions for Datasets	Burak Yavuz	2016-04-01	1	-0/+242
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR adds the function `window` as a column expression. `window` can be used to bucket rows into time windows given a time column. With this expression, performing time series analysis on batch data, as well as streaming data should become much more simpler. ### Usage Assume the following schema: `sensor_id, measurement, timestamp` To average 5 minute data every 1 minute (window length of 5 minutes, slide duration of 1 minute), we will use: ```scala df.groupBy(window("timestamp", “5 minutes”, “1 minute”), "sensor_id") .agg(mean("measurement").as("avg_meas")) ``` This will generate windows such as: ``` 09:00:00-09:05:00 09:01:00-09:06:00 09:02:00-09:07:00 ... ``` Intervals will start at every `slideDuration` starting at the unix epoch (1970-01-01 00:00:00 UTC). To start intervals at a different point of time, e.g. 30 seconds after a minute, the `startTime` parameter can be used. ```scala df.groupBy(window("timestamp", “5 minutes”, “1 minute”, "30 second"), "sensor_id") .agg(mean("measurement").as("avg_meas")) ``` This will generate windows such as: ``` 09:00:30-09:05:30 09:01:30-09:06:30 09:02:30-09:07:30 ... ``` Support for Python will be made in a follow up PR after this. ## How was this patch tested? This patch has some basic unit tests for the `TimeWindow` expression testing that the parameters pass validation, and it also has some unit/integration tests testing the correctness of the windowing and usability in complex operations (multi-column grouping, multi-column projections, joins). Author: Burak Yavuz <brkyvz@gmail.com> Author: Michael Armbrust <michael@databricks.com> Closes #12008 from brkyvz/df-time-window.
*	[SPARK-14184][SQL] Support native execution of SHOW DATABASE command and fix ↵	Dilip Biswal	2016-04-01	2	-0/+83
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	SHOW TABLE to use table identifier pattern ## What changes were proposed in this pull request? This PR addresses the following 1. Supports native execution of SHOW DATABASES command 2. Fixes SHOW TABLES to apply the identifier_with_wildcards pattern if supplied. SHOW TABLE syntax ``` SHOW TABLES [IN database_name] ['identifier_with_wildcards']; ``` SHOW DATABASES syntax ``` SHOW (DATABASES\|SCHEMAS) [LIKE 'identifier_with_wildcards']; ``` ## How was this patch tested? Tests added in SQLQuerySuite (both hive and sql contexts) and DDLCommandSuite Note: Since the table name pattern was not working , tests are added in both SQLQuerySuite to verify the application of the table pattern. Author: Dilip Biswal <dbiswal@us.ibm.com> Closes #11991 from dilipbiswal/dkb_show_database.
*	[SPARK-14304][SQL][TESTS] Fix tests that don't create temp files in the ↵	Shixiong Zhu	2016-03-31	6	-22/+24
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	`java.io.tmpdir` folder ## What changes were proposed in this pull request? If I press `CTRL-C` when running these tests, the temp files will be left in `sql/core` folder and I need to delete them manually. It's annoying. This PR just moves the temp files to the `java.io.tmpdir` folder and add a name prefix for them. ## How was this patch tested? Existing Jenkins tests Author: Shixiong Zhu <shixiong@databricks.com> Closes #12093 from zsxwing/temp-file.
*	[SPARK-14182][SQL] Parse DDL Command: Alter View	gatorsmile	2016-03-31	1	-38/+137
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This PR is to provide native parsing support for DDL commands: `Alter View`. Since its AST trees are highly similar to `Alter Table`. Thus, both implementation are integrated into the same one. Based on the Hive DDL document: https://cwiki.apache.org/confluence/display/Hive/LanguageManual+DDL and https://cwiki.apache.org/confluence/display/Hive/PartitionedViews Syntax: ```SQL ALTER VIEW view_name RENAME TO new_view_name ``` - to change the name of a view to a different name Syntax: ```SQL ALTER VIEW view_name SET TBLPROPERTIES ('comment' = new_comment); ``` - to add metadata to a view Syntax: ```SQL ALTER VIEW view_name UNSET TBLPROPERTIES [IF EXISTS] ('comment', 'key') ``` - to remove metadata from a view Syntax: ```SQL ALTER VIEW view_name ADD [IF NOT EXISTS] PARTITION spec1[, PARTITION spec2, ...] ``` - to add the partitioning metadata for a view. - the syntax of partition spec in `ALTER VIEW` is identical to `ALTER TABLE`, EXCEPT that it is ILLEGAL to specify a `LOCATION` clause. Syntax: ```SQL ALTER VIEW view_name DROP [IF EXISTS] PARTITION spec1[, PARTITION spec2, ...] ``` - to drop the related partition metadata for a view. Added the related test cases to `DDLCommandSuite` Author: gatorsmile <gatorsmile@gmail.com> Author: xiaoli <lixiao1983@gmail.com> Author: Xiao Li <xiaoli@Xiaos-MacBook-Pro.local> Closes #11987 from gatorsmile/parseAlterView.
*	[SPARK-14263][SQL] Benchmark Vectorized HashMap for GroupBy Aggregates	Sameer Agarwal	2016-03-31	1	-10/+35
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR proposes a new data-structure based on a vectorized hashmap that can be potentially _codegened_ in `TungstenAggregate` to speed up aggregates with group by. Micro-benchmarks show a 10x improvement over the current `BytesToBytes` aggregation map. ## How was this patch tested? Intel(R) Core(TM) i7-4960HQ CPU 2.60GHz BytesToBytesMap: Best/Avg Time(ms) Rate(M/s) Per Row(ns) Relative ------------------------------------------------------------------------------------------- hash 108 / 119 96.9 10.3 1.0X fast hash 63 / 70 166.2 6.0 1.7X arrayEqual 70 / 73 150.8 6.6 1.6X Java HashMap (Long) 141 / 200 74.3 13.5 0.8X Java HashMap (two ints) 145 / 185 72.3 13.8 0.7X Java HashMap (UnsafeRow) 499 / 524 21.0 47.6 0.2X BytesToBytesMap (off Heap) 483 / 548 21.7 46.0 0.2X BytesToBytesMap (on Heap) 485 / 562 21.6 46.2 0.2X Vectorized Hashmap 54 / 60 193.7 5.2 2.0X Author: Sameer Agarwal <sameer@databricks.com> Closes #12055 from sameeragarwal/vectorized-hashmap.
*	[SPARK-14211][SQL] Remove ANTLR3 based parser	Herman van Hovell	2016-03-31	1	-1/+12
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	### What changes were proposed in this pull request? This PR removes the ANTLR3 based parser, and moves the new ANTLR4 based parser into the `org.apache.spark.sql.catalyst.parser package`. ### How was this patch tested? Existing unit tests. cc rxin andrewor14 yhuai Author: Herman van Hovell <hvanhovell@questtec.nl> Closes #12071 from hvanhovell/SPARK-14211.
*	[SPARK-14206][SQL] buildReader() implementation for CSV	Cheng Lian	2016-03-30	1	-2/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? Major changes: 1. Implement `FileFormat.buildReader()` for the CSV data source. 1. Add an extra argument to `FileFormat.buildReader()`, `physicalSchema`, which is basically the result of `FileFormat.inferSchema` or user specified schema. This argument is necessary because the CSV data source needs to know all the columns of the underlying files to read the file. ## How was this patch tested? Existing tests should do the work. Author: Cheng Lian <lian@databricks.com> Closes #12002 from liancheng/spark-14206-csv-build-reader.
*	[SPARK-14081][SQL] - Preserve DataFrame column types when filling nulls.	Travis Crawford	2016-03-30	1	-20/+30
\| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This change resolves an issue where `DataFrameNaFunctions.fill` changes a `FloatType` column to a `DoubleType`. We also clarify the contract that replacement values will be cast to the column data type, which may change the replacement value when casting to a lower precision type. ## How was this patch tested? This patch has associated unit tests. Author: Travis Crawford <travis@medium.com> Closes #11967 from traviscrawford/SPARK-14081-dataframena.
*	[SPARK-14259][SQL] Add a FileSourceStrategy option for limiting #files in a ↵	Takeshi YAMAMURO	2016-03-30	1	-0/+47
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	partition ## What changes were proposed in this pull request? This pr is to add a config to control the maximum number of files as even small files have a non-trivial fixed cost. The current packing can put a lot of small files together which cases straggler tasks. ## How was this patch tested? I added tests to check if many files get split into partitions in FileSourceStrategySuite. Author: Takeshi YAMAMURO <linguin.m.s@gmail.com> Closes #12068 from maropu/SPARK-14259.
*	[SPARK-14268][SQL] rename toRowExpressions and fromRowExpression to ↵	Wenchen Fan	2016-03-30	1	-2/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	serializer and deserializer in ExpressionEncoder ## What changes were proposed in this pull request? In `ExpressionEncoder`, we use `constructorFor` to build `fromRowExpression` as the `deserializer` in `ObjectOperator`. It's kind of confusing, we should make the name consistent. ## How was this patch tested? existing tests. Author: Wenchen Fan <wenchen@databricks.com> Closes #12058 from cloud-fan/rename.
*	[SPARK-14124][SQL] Implement Database-related DDL Commands	gatorsmile	2016-03-29	2	-27/+166
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	#### What changes were proposed in this pull request? This PR is to implement the following four Database-related DDL commands: - `CREATE DATABASE\|SCHEMA [IF NOT EXISTS] database_name` - `DROP DATABASE [IF EXISTS] database_name [RESTRICT\|CASCADE]` - `DESCRIBE DATABASE [EXTENDED] db_name` - `ALTER (DATABASE\|SCHEMA) database_name SET DBPROPERTIES (property_name=property_value, ...)` Another PR will be submitted to handle the unsupported commands. In the Database-related DDL commands, we will issue an error exception for `ALTER (DATABASE\|SCHEMA) database_name SET OWNER [USER\|ROLE] user_or_role`. cc yhuai andrewor14 rxin Could you review the changes? Is it in the right direction? Thanks! #### How was this patch tested? Added a few test cases in `command/DDLSuite.scala` for testing DDL command execution in `SQLContext`. Since `HiveContext` also shares the same implementation, the existing test cases in `\hive` also verifies the correctness of these commands. Author: gatorsmile <gatorsmile@gmail.com> Author: xiaoli <lixiao1983@gmail.com> Author: Xiao Li <xiaoli@Xiaos-MacBook-Pro.local> Closes #12009 from gatorsmile/dbDDL.
*	[SPARK-14227][SQL] Add method for printing out generated code for debugging	Eric Liang	2016-03-29	1	-0/+7
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This adds `debugCodegen` to the debug package for query execution. ## How was this patch tested? Unit and manual testing. Output example: ``` scala> import org.apache.spark.sql.execution.debug._ import org.apache.spark.sql.execution.debug._ scala> sqlContext.range(100).groupBy("id").count().orderBy("id").debugCodegen() Found 3 WholeStageCodegen subtrees. == Subtree 1 / 3 == WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) : +- Range 0, 1, 1, 100, [id#0L] Generated code: /* 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } / 004 / / 005 / /* Codegened pipeline for: /* 006 / TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) /* 007 / +- Range 0, 1, 1, 100, [id#0L] / 008 / / /* 009 / final class GeneratedIterator extends org.apache.spark.sql.execution.BufferedRowIterator { / 010 / private Object[] references; / 011 / private boolean agg_initAgg; / 012 / private org.apache.spark.sql.execution.aggregate.TungstenAggregate agg_plan; / 013 / private org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap agg_hashMap; / 014 / private org.apache.spark.sql.execution.UnsafeKVExternalSorter agg_sorter; / 015 / private org.apache.spark.unsafe.KVIterator agg_mapIter; / 016 / private org.apache.spark.sql.execution.metric.LongSQLMetric range_numOutputRows; / 017 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue range_metricValue; / 018 / private boolean range_initRange; / 019 / private long range_partitionEnd; / 020 / private long range_number; / 021 / private boolean range_overflow; / 022 / private scala.collection.Iterator range_input; / 023 / private UnsafeRow range_result; / 024 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder range_holder; / 025 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter range_rowWriter; / 026 / private UnsafeRow agg_result; / 027 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder agg_holder; / 028 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter agg_rowWriter; / 029 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowJoiner agg_unsafeRowJoiner; / 030 / private org.apache.spark.sql.execution.metric.LongSQLMetric wholestagecodegen_numOutputRows; / 031 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue wholestagecodegen_metricValue; / 032 / / 033 / public GeneratedIterator(Object[] references) { / 034 / this.references = references; / 035 / } / 036 / / 037 / public void init(scala.collection.Iterator inputs[]) { / 038 / agg_initAgg = false; / 039 / this.agg_plan = (org.apache.spark.sql.execution.aggregate.TungstenAggregate) references[0]; / 040 / agg_hashMap = agg_plan.createHashMap(); / 041 / / 042 / this.range_numOutputRows = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[1]; / 043 / range_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) range_numOutputRows.localValue(); / 044 / range_initRange = false; / 045 / range_partitionEnd = 0L; / 046 / range_number = 0L; / 047 / range_overflow = false; / 048 / range_input = inputs[0]; / 049 / range_result = new UnsafeRow(1); / 050 / this.range_holder = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(range_result, 0); / 051 / this.range_rowWriter = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(range_holder, 1); / 052 / agg_result = new UnsafeRow(1); / 053 / this.agg_holder = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(agg_result, 0); / 054 / this.agg_rowWriter = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(agg_holder, 1); / 055 / agg_unsafeRowJoiner = agg_plan.createUnsafeJoiner(); / 056 / this.wholestagecodegen_numOutputRows = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[2]; / 057 / wholestagecodegen_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) wholestagecodegen_numOutputRows.localValue(); / 058 / } / 059 / / 060 / private void agg_doAggregateWithKeys() throws java.io.IOException { / 061 / /** PRODUCE: Range 0, 1, 1, 100, [id#0L] / / 062 / / 063 / // initialize Range / 064 / if (!range_initRange) { / 065 / range_initRange = true; / 066 / if (range_input.hasNext()) { / 067 / initRange(((InternalRow) range_input.next()).getInt(0)); / 068 / } else { / 069 / return; / 070 / } / 071 / } / 072 / / 073 / while (!range_overflow && range_number < range_partitionEnd) { / 074 / long range_value = range_number; / 075 / range_number += 1L; / 076 / if (range_number < range_value ^ 1L < 0) { / 077 / range_overflow = true; / 078 / } / 079 / / 080 / /** CONSUME: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) / / 081 / / 082 / // generate grouping key / 083 / agg_rowWriter.write(0, range_value); / 084 / / hash(input[0, bigint], 42) / / 085 / int agg_value1 = 42; / 086 / / 087 / agg_value1 = org.apache.spark.unsafe.hash.Murmur3_x86_32.hashLong(range_value, agg_value1); / 088 / UnsafeRow agg_aggBuffer = null; / 089 / if (true) { / 090 / // try to get the buffer from hash map / 091 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 092 / } / 093 / if (agg_aggBuffer == null) { / 094 / if (agg_sorter == null) { / 095 / agg_sorter = agg_hashMap.destructAndCreateExternalSorter(); / 096 / } else { / 097 / agg_sorter.merge(agg_hashMap.destructAndCreateExternalSorter()); / 098 / } / 099 / / 100 / // the hash map had be spilled, it should have enough memory now, / 101 / // try to allocate buffer again. / 102 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 103 / if (agg_aggBuffer == null) { / 104 / // failed to allocate the first page / 105 / throw new OutOfMemoryError("No enough memory for aggregation"); / 106 / } / 107 / } / 108 / / 109 / // evaluate aggregate function / 110 / / (input[0, bigint] + 1) / / 111 / / input[0, bigint] / / 112 / long agg_value4 = agg_aggBuffer.getLong(0); / 113 / / 114 / long agg_value3 = -1L; / 115 / agg_value3 = agg_value4 + 1L; / 116 / // update aggregate buffer / 117 / agg_aggBuffer.setLong(0, agg_value3); / 118 / / 119 / if (shouldStop()) return; / 120 / } / 121 / / 122 / agg_mapIter = agg_plan.finishAggregate(agg_hashMap, agg_sorter); / 123 / } / 124 / / 125 / private void initRange(int idx) { / 126 / java.math.BigInteger index = java.math.BigInteger.valueOf(idx); / 127 / java.math.BigInteger numSlice = java.math.BigInteger.valueOf(1L); / 128 / java.math.BigInteger numElement = java.math.BigInteger.valueOf(100L); / 129 / java.math.BigInteger step = java.math.BigInteger.valueOf(1L); / 130 / java.math.BigInteger start = java.math.BigInteger.valueOf(0L); / 131 / / 132 / java.math.BigInteger st = index.multiply(numElement).divide(numSlice).multiply(step).add(start); / 133 / if (st.compareTo(java.math.BigInteger.valueOf(Long.MAX_VALUE)) > 0) { / 134 / range_number = Long.MAX_VALUE; / 135 / } else if (st.compareTo(java.math.BigInteger.valueOf(Long.MIN_VALUE)) < 0) { / 136 / range_number = Long.MIN_VALUE; / 137 / } else { / 138 / range_number = st.longValue(); / 139 / } / 140 / / 141 / java.math.BigInteger end = index.add(java.math.BigInteger.ONE).multiply(numElement).divide(numSlice) / 142 / .multiply(step).add(start); / 143 / if (end.compareTo(java.math.BigInteger.valueOf(Long.MAX_VALUE)) > 0) { / 144 / range_partitionEnd = Long.MAX_VALUE; / 145 / } else if (end.compareTo(java.math.BigInteger.valueOf(Long.MIN_VALUE)) < 0) { / 146 / range_partitionEnd = Long.MIN_VALUE; / 147 / } else { / 148 / range_partitionEnd = end.longValue(); / 149 / } / 150 / / 151 / range_metricValue.add((range_partitionEnd - range_number) / 1L); / 152 / } / 153 / / 154 / protected void processNext() throws java.io.IOException { / 155 / /** PRODUCE: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) / / 156 / / 157 / if (!agg_initAgg) { / 158 / agg_initAgg = true; / 159 / agg_doAggregateWithKeys(); / 160 / } / 161 / / 162 / // output the result / 163 / while (agg_mapIter.next()) { / 164 / wholestagecodegen_metricValue.add(1); / 165 / UnsafeRow agg_aggKey = (UnsafeRow) agg_mapIter.getKey(); / 166 / UnsafeRow agg_aggBuffer1 = (UnsafeRow) agg_mapIter.getValue(); / 167 / / 168 / UnsafeRow agg_resultRow = agg_unsafeRowJoiner.join(agg_aggKey, agg_aggBuffer1); / 169 / / 170 / /** CONSUME: WholeStageCodegen / / 171 / / 172 / append(agg_resultRow); / 173 / / 174 / if (shouldStop()) return; / 175 / } / 176 / / 177 / agg_mapIter.close(); / 178 / if (agg_sorter == null) { / 179 / agg_hashMap.free(); / 180 / } / 181 / } / 182 / } == Subtree 2 / 3 == WholeStageCodegen : +- Sort [id#0L ASC], true, 0 : +- INPUT +- Exchange rangepartitioning(id#0L ASC, 200), None +- WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) : +- INPUT +- Exchange hashpartitioning(id#0L, 200), None +- WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) : +- Range 0, 1, 1, 100, [id#0L] Generated code: / 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } / 004 / / 005 / /* Codegened pipeline for: /* 006 / Sort [id#0L ASC], true, 0 /* 007 / +- INPUT / 008 / / /* 009 / final class GeneratedIterator extends org.apache.spark.sql.execution.BufferedRowIterator { / 010 / private Object[] references; / 011 / private boolean sort_needToSort; / 012 / private org.apache.spark.sql.execution.Sort sort_plan; / 013 / private org.apache.spark.sql.execution.UnsafeExternalRowSorter sort_sorter; / 014 / private org.apache.spark.executor.TaskMetrics sort_metrics; / 015 / private scala.collection.Iterator<UnsafeRow> sort_sortedIter; / 016 / private scala.collection.Iterator inputadapter_input; / 017 / private org.apache.spark.sql.execution.metric.LongSQLMetric sort_dataSize; / 018 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue sort_metricValue; / 019 / private org.apache.spark.sql.execution.metric.LongSQLMetric sort_spillSize; / 020 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue sort_metricValue1; / 021 / / 022 / public GeneratedIterator(Object[] references) { / 023 / this.references = references; / 024 / } / 025 / / 026 / public void init(scala.collection.Iterator inputs[]) { / 027 / sort_needToSort = true; / 028 / this.sort_plan = (org.apache.spark.sql.execution.Sort) references[0]; / 029 / sort_sorter = sort_plan.createSorter(); / 030 / sort_metrics = org.apache.spark.TaskContext.get().taskMetrics(); / 031 / / 032 / inputadapter_input = inputs[0]; / 033 / this.sort_dataSize = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[1]; / 034 / sort_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) sort_dataSize.localValue(); / 035 / this.sort_spillSize = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[2]; / 036 / sort_metricValue1 = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) sort_spillSize.localValue(); / 037 / } / 038 / / 039 / private void sort_addToSorter() throws java.io.IOException { / 040 / /** PRODUCE: INPUT / / 041 / / 042 / while (inputadapter_input.hasNext()) { / 043 / InternalRow inputadapter_row = (InternalRow) inputadapter_input.next(); / 044 / /** CONSUME: Sort [id#0L ASC], true, 0 / / 045 / / 046 / sort_sorter.insertRow((UnsafeRow)inputadapter_row); / 047 / if (shouldStop()) return; / 048 / } / 049 / / 050 / } / 051 / / 052 / protected void processNext() throws java.io.IOException { / 053 / /** PRODUCE: Sort [id#0L ASC], true, 0 / / 054 / if (sort_needToSort) { / 055 / sort_addToSorter(); / 056 / Long sort_spillSizeBefore = sort_metrics.memoryBytesSpilled(); / 057 / sort_sortedIter = sort_sorter.sort(); / 058 / sort_metricValue.add(sort_sorter.getPeakMemoryUsage()); / 059 / sort_metricValue1.add(sort_metrics.memoryBytesSpilled() - sort_spillSizeBefore); / 060 / sort_metrics.incPeakExecutionMemory(sort_sorter.getPeakMemoryUsage()); / 061 / sort_needToSort = false; / 062 / } / 063 / / 064 / while (sort_sortedIter.hasNext()) { / 065 / UnsafeRow sort_outputRow = (UnsafeRow)sort_sortedIter.next(); / 066 / / 067 / /** CONSUME: WholeStageCodegen / / 068 / / 069 / append(sort_outputRow); / 070 / / 071 / if (shouldStop()) return; / 072 / } / 073 / } / 074 / } == Subtree 3 / 3 == WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) : +- INPUT +- Exchange hashpartitioning(id#0L, 200), None +- WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) : +- Range 0, 1, 1, 100, [id#0L] Generated code: / 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } / 004 / / 005 / /* Codegened pipeline for: /* 006 / TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) /* 007 / +- INPUT / 008 / / /* 009 / final class GeneratedIterator extends org.apache.spark.sql.execution.BufferedRowIterator { / 010 / private Object[] references; / 011 / private boolean agg_initAgg; / 012 / private org.apache.spark.sql.execution.aggregate.TungstenAggregate agg_plan; / 013 / private org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap agg_hashMap; / 014 / private org.apache.spark.sql.execution.UnsafeKVExternalSorter agg_sorter; / 015 / private org.apache.spark.unsafe.KVIterator agg_mapIter; / 016 / private scala.collection.Iterator inputadapter_input; / 017 / private UnsafeRow agg_result; / 018 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder agg_holder; / 019 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter agg_rowWriter; / 020 / private UnsafeRow agg_result1; / 021 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder agg_holder1; / 022 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter agg_rowWriter1; / 023 / private org.apache.spark.sql.execution.metric.LongSQLMetric wholestagecodegen_numOutputRows; / 024 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue wholestagecodegen_metricValue; / 025 / / 026 / public GeneratedIterator(Object[] references) { / 027 / this.references = references; / 028 / } / 029 / / 030 / public void init(scala.collection.Iterator inputs[]) { / 031 / agg_initAgg = false; / 032 / this.agg_plan = (org.apache.spark.sql.execution.aggregate.TungstenAggregate) references[0]; / 033 / agg_hashMap = agg_plan.createHashMap(); / 034 / / 035 / inputadapter_input = inputs[0]; / 036 / agg_result = new UnsafeRow(1); / 037 / this.agg_holder = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(agg_result, 0); / 038 / this.agg_rowWriter = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(agg_holder, 1); / 039 / agg_result1 = new UnsafeRow(2); / 040 / this.agg_holder1 = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(agg_result1, 0); / 041 / this.agg_rowWriter1 = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(agg_holder1, 2); / 042 / this.wholestagecodegen_numOutputRows = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[1]; / 043 / wholestagecodegen_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) wholestagecodegen_numOutputRows.localValue(); / 044 / } / 045 / / 046 / private void agg_doAggregateWithKeys() throws java.io.IOException { / 047 / /** PRODUCE: INPUT / / 048 / / 049 / while (inputadapter_input.hasNext()) { / 050 / InternalRow inputadapter_row = (InternalRow) inputadapter_input.next(); / 051 / /** CONSUME: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) / / 052 / / input[0, bigint] / / 053 / long inputadapter_value = inputadapter_row.getLong(0); / 054 / / input[1, bigint] / / 055 / long inputadapter_value1 = inputadapter_row.getLong(1); / 056 / / 057 / // generate grouping key / 058 / agg_rowWriter.write(0, inputadapter_value); / 059 / / hash(input[0, bigint], 42) / / 060 / int agg_value1 = 42; / 061 / / 062 / agg_value1 = org.apache.spark.unsafe.hash.Murmur3_x86_32.hashLong(inputadapter_value, agg_value1); / 063 / UnsafeRow agg_aggBuffer = null; / 064 / if (true) { / 065 / // try to get the buffer from hash map / 066 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 067 / } / 068 / if (agg_aggBuffer == null) { / 069 / if (agg_sorter == null) { / 070 / agg_sorter = agg_hashMap.destructAndCreateExternalSorter(); / 071 / } else { / 072 / agg_sorter.merge(agg_hashMap.destructAndCreateExternalSorter()); / 073 / } / 074 / / 075 / // the hash map had be spilled, it should have enough memory now, / 076 / // try to allocate buffer again. / 077 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 078 / if (agg_aggBuffer == null) { / 079 / // failed to allocate the first page / 080 / throw new OutOfMemoryError("No enough memory for aggregation"); / 081 / } / 082 / } / 083 / / 084 / // evaluate aggregate function / 085 / / (input[0, bigint] + input[2, bigint]) / / 086 / / input[0, bigint] / / 087 / long agg_value4 = agg_aggBuffer.getLong(0); / 088 / / 089 / long agg_value3 = -1L; / 090 / agg_value3 = agg_value4 + inputadapter_value1; / 091 / // update aggregate buffer / 092 / agg_aggBuffer.setLong(0, agg_value3); / 093 / if (shouldStop()) return; / 094 / } / 095 / / 096 / agg_mapIter = agg_plan.finishAggregate(agg_hashMap, agg_sorter); / 097 / } / 098 / / 099 / protected void processNext() throws java.io.IOException { / 100 / /** PRODUCE: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) / / 101 / / 102 / if (!agg_initAgg) { / 103 / agg_initAgg = true; / 104 / agg_doAggregateWithKeys(); / 105 / } / 106 / / 107 / // output the result / 108 / while (agg_mapIter.next()) { / 109 / wholestagecodegen_metricValue.add(1); / 110 / UnsafeRow agg_aggKey = (UnsafeRow) agg_mapIter.getKey(); / 111 / UnsafeRow agg_aggBuffer1 = (UnsafeRow) agg_mapIter.getValue(); / 112 / / 113 / / input[0, bigint] / / 114 / long agg_value6 = agg_aggKey.getLong(0); / 115 / / input[0, bigint] / / 116 / long agg_value7 = agg_aggBuffer1.getLong(0); / 117 / / 118 / /** CONSUME: WholeStageCodegen / / 119 / / 120 / agg_rowWriter1.write(0, agg_value6); / 121 / / 122 / agg_rowWriter1.write(1, agg_value7); / 123 / append(agg_result1); / 124 / / 125 / if (shouldStop()) return; / 126 / } / 127 / / 128 / agg_mapIter.close(); / 129 / if (agg_sorter == null) { / 130 / agg_hashMap.free(); / 131 / } / 132 / } / 133 */ } ``` rxin Author: Eric Liang <ekl@databricks.com> Closes #12025 from ericl/spark-14227.
*	[SPARK-14205][SQL] remove trait Queryable	Wenchen Fan	2016-03-28	1	-5/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? After DataFrame and Dataset are merged, the trait `Queryable` becomes unnecessary as it has only one implementation. We should remove it. ## How was this patch tested? existing tests. Author: Wenchen Fan <wenchen@databricks.com> Closes #12001 from cloud-fan/df-ds.
*	[SPARK-14086][SQL] Add DDL commands to ANTLR4 parser	Herman van Hovell	2016-03-28	1	-3/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	#### What changes were proposed in this pull request? This PR adds all the current Spark SQL DDL commands to the new ANTLR 4 based SQL parser. I have found a few inconsistencies in the current commands: - Function has an alias field. This is actually the class name of the function. - Partition specifications should contain nulls in some commands, and contain `None`s in others. - `AlterTableSkewedLocation`: Should defines which columns have skewed values, and should allow us to define storage for each skewed combination of values. We currently only allow one value per field. - `AlterTableSetFileFormat`: Should only have one file format, it currently supports both. I have implemented all these comments like they were, and I propose to improve them in follow-up PRs. #### How was this patch tested? The existing DDLCommandSuite. cc rxin andrewor14 yhuai Author: Herman van Hovell <hvanhovell@questtec.nl> Closes #12011 from hvanhovell/SPARK-14086.
*	[SPARK-14052] [SQL] build a BytesToBytesMap directly in HashedRelation	Davies Liu	2016-03-28	2	-7/+34
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? Currently, for the key that can not fit within a long, we build a hash map for UnsafeHashedRelation, it's converted to BytesToBytesMap after serialization and deserialization. We should build a BytesToBytesMap directly to have better memory efficiency. In order to do that, BytesToBytesMap should support multiple (K,V) pair with the same K, Location.putNewKey() is renamed to Location.append(), which could append multiple values for the same key (same Location). `Location.newValue()` is added to find the next value for the same key. ## How was this patch tested? Existing tests. Added benchmark for broadcast hash join with duplicated keys. Author: Davies Liu <davies@databricks.com> Closes #11870 from davies/map2.
*	[SPARK-13713][SQL] Migrate parser from ANTLR3 to ANTLR4	Herman van Hovell	2016-03-28	2	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	### What changes were proposed in this pull request? The current ANTLR3 parser is quite complex to maintain and suffers from code blow-ups. This PR introduces a new parser that is based on ANTLR4. This parser is based on the [Presto's SQL parser](https://github.com/facebook/presto/blob/master/presto-parser/src/main/antlr4/com/facebook/presto/sql/parser/SqlBase.g4). The current implementation can parse and create Catalyst and SQL plans. Large parts of the HiveQl DDL and some of the DML functionality is currently missing, the plan is to add this in follow-up PRs. This PR is a work in progress, and work needs to be done in the following area's: - [x] Error handling should be improved. - [x] Documentation should be improved. - [x] Multi-Insert needs to be tested. - [ ] Naming and package locations. ### How was this patch tested? Catalyst and SQL unit tests. Author: Herman van Hovell <hvanhovell@questtec.nl> Closes #11557 from hvanhovell/ngParser.
*	[SPARK-14177][SQL] Native Parsing for DDL Command "Describe Database" and ↵	gatorsmile	2016-03-26	1	-0/+38
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	"Alter Database" #### What changes were proposed in this pull request? This PR is to provide native parsing support for two DDL commands: ```Describe Database``` and ```Alter Database Set Properties``` Based on the Hive DDL document: https://cwiki.apache.org/confluence/display/Hive/LanguageManual+DDL ##### 1. ALTER DATABASE Syntax: ```SQL ALTER (DATABASE\|SCHEMA) database_name SET DBPROPERTIES (property_name=property_value, ...) ``` - `ALTER DATABASE` is to add new (key, value) pairs into `DBPROPERTIES` ##### 2. DESCRIBE DATABASE Syntax: ```SQL DESCRIBE DATABASE [EXTENDED] db_name ``` - `DESCRIBE DATABASE` shows the name of the database, its comment (if one has been set), and its root location on the filesystem. When `extended` is true, it also shows the database's properties #### How was this patch tested? Added the related test cases to `DDLCommandSuite` Author: gatorsmile <gatorsmile@gmail.com> Author: xiaoli <lixiao1983@gmail.com> Author: Xiao Li <xiaoli@Xiaos-MacBook-Pro.local> This patch had conflicts when merged, resolved by Committer: Yin Huai <yhuai@databricks.com> Closes #11977 from gatorsmile/parseAlterDatabase.
*	[SPARK-14157][SQL] Parse Drop Function DDL command	Liang-Chi Hsieh	2016-03-26	1	-1/+41
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? JIRA: https://issues.apache.org/jira/browse/SPARK-14157 We only parse create function command. In order to support native drop function command, we need to parse it too. From Hive [manual](https://cwiki.apache.org/confluence/display/Hive/LanguageManual+DDL#LanguageManualDDL-Create/Drop/ReloadFunction), the drop function command has syntax as: DROP [TEMPORARY] FUNCTION [IF EXISTS] function_name; ## How was this patch tested? Added test into `DDLCommandSuite`. Author: Liang-Chi Hsieh <simonh@tw.ibm.com> Closes #11959 from viirya/parse-drop-func.
*	[SPARK-14161][SQL] Native Parsing for DDL Command Drop Database	gatorsmile	2016-03-26	1	-0/+57
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	### What changes were proposed in this pull request? Based on the Hive DDL document https://cwiki.apache.org/confluence/display/Hive/LanguageManual+DDL The syntax of DDL command for Drop Database is ```SQL DROP (DATABASE\|SCHEMA) [IF EXISTS] database_name [RESTRICT\|CASCADE]; ``` - If `IF EXISTS` is not specified, the default behavior is to issue a warning message if `database_name` does't exist - `RESTRICT` is the default behavior. This PR is to provide a native parsing support for `DROP DATABASE`. #### How was this patch tested? Added a test case `DDLCommandSuite` Author: gatorsmile <gatorsmile@gmail.com> Closes #11962 from gatorsmile/parseDropDatabase.
*	[MINOR] Fix newly added java-lint errors	Dongjoon Hyun	2016-03-26	1	-12/+16
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR fixes some newly added java-lint errors(unused-imports, line-lengsth). ## How was this patch tested? Pass the Jenkins tests. Author: Dongjoon Hyun <dongjoon@apache.org> Closes #11968 from dongjoon-hyun/SPARK-14167.
*	[SPARK-14109][SQL] Fix HDFSMetadataLog to fallback from FileContext to ↵	Tathagata Das	2016-03-25	3	-8/+127
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	FileSystem API ## What changes were proposed in this pull request? HDFSMetadataLog uses newer FileContext API to achieve atomic renaming. However, FileContext implementations may not exist for many scheme for which there may be FileSystem implementations. In those cases, rather than failing completely, we should fallback to the FileSystem based implementation, and log warning that there may be file consistency issues in case the log directory is concurrently modified. In addition I have also added more tests to increase the code coverage. ## How was this patch tested? Unit test. Tested on cluster with custom file system. Author: Tathagata Das <tathagata.das1565@gmail.com> Closes #11925 from tdas/SPARK-14109.
*	[SPARK-14073][STREAMING][TEST-MAVEN] Move flume back to Spark	Shixiong Zhu	2016-03-25	1	-8/+25
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR moves flume back to Spark as per the discussion in the dev mail-list. ## How was this patch tested? Existing Jenkins tests. Author: Shixiong Zhu <shixiong@databricks.com> Closes #11895 from zsxwing/move-flume-back.
*	[SQL][HOTFIX] Fix flakiness in StateStoreRDDSuite	Tathagata Das	2016-03-25	2	-4/+8
\| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? StateStoreCoordinator.reportActiveInstance is async, so subsequence state checks must be in eventually. ## How was this patch tested? Jenkins tests Author: Tathagata Das <tathagata.das1565@gmail.com> Closes #11924 from tdas/state-store-flaky-fix.
*	[SPARK-14061][SQL] implement CreateMap	Wenchen Fan	2016-03-25	2	-8/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? As we have `CreateArray` and `CreateStruct`, we should also have `CreateMap`. This PR adds the `CreateMap` expression, and the DataFrame API, and python API. ## How was this patch tested? various new tests. Author: Wenchen Fan <wenchen@databricks.com> Closes #11879 from cloud-fan/create_map.
*	[SPARK-14014][SQL] Integrate session catalog (attempt #2)	Andrew Or	2016-03-24	5	-20/+36
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This reopens #11836, which was merged but promptly reverted because it introduced flaky Hive tests. ## How was this patch tested? See `CatalogTestCases`, `SessionCatalogSuite` and `HiveContextSuite`. Author: Andrew Or <andrew@databricks.com> Closes #11938 from andrewor14/session-catalog-again.
*	[SPARK-14145][SQL] Remove the untyped version of Dataset.groupByKey	Reynold Xin	2016-03-24	3	-73/+1
\| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? Dataset has two variants of groupByKey, one for untyped and the other for typed. It actually doesn't make as much sense to have an untyped API here, since apps that want to use untyped APIs should just use the groupBy "DataFrame" API. ## How was this patch tested? This patch removes a method, and removes the associated tests. Author: Reynold Xin <rxin@databricks.com> Closes #11949 from rxin/SPARK-14145.
*	[SPARK-14142][SQL] Replace internal use of unionAll with union	Reynold Xin	2016-03-24	8	-14/+14
\| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? unionAll has been deprecated in SPARK-14088. ## How was this patch tested? Should be covered by all existing tests. Author: Reynold Xin <rxin@databricks.com> Closes #11946 from rxin/SPARK-14142.
*	[SPARK-13957][SQL] Support Group By Ordinal in SQL	gatorsmile	2016-03-25	1	-7/+85
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	#### What changes were proposed in this pull request? This PR is to support group by position in SQL. For example, when users input the following query ```SQL select c1 as a, c2, c3, sum() from tbl group by 1, 3, c4 ``` The ordinals are recognized as the positions in the select list. Thus, `Analyzer` converts it to ```SQL select c1, c2, c3, sum() from tbl group by c1, c3, c4 ``` This is controlled by the config option `spark.sql.groupByOrdinal`. - When true, the ordinal numbers in group by clauses are treated as the position in the select list. - When false, the ordinal numbers are ignored. - Only convert integer literals (not foldable expressions). If found foldable expressions, ignore them. - When the positions specified in the group by clauses correspond to the aggregate functions in select list, output an exception message. - star is not allowed to use in the select list when users specify ordinals in group by Note: This PR is taken from https://github.com/apache/spark/pull/10731. When merging this PR, please give the credit to zhichao-li Also cc all the people who are involved in the previous discussion: rxin cloud-fan marmbrus yhuai hvanhovell adrian-wang chenghao-intel tejasapatil #### How was this patch tested? Added a few test cases for both positive and negative test cases. Author: gatorsmile <gatorsmile@gmail.com> Author: xiaoli <lixiao1983@gmail.com> Author: Xiao Li <xiaoli@Xiaos-MacBook-Pro.local> Closes #11846 from gatorsmile/groupByOrdinal.
*	Revert "[SPARK-14014][SQL] Replace existing catalog with SessionCatalog"	Andrew Or	2016-03-23	5	-36/+20
\| \| \| \|	This reverts commit 5dfc01976bb0d72489620b4f32cc12d620bb6260.
*	[SPARK-14085][SQL] Star Expansion for Hash	gatorsmile	2016-03-24	2	-0/+27
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	#### What changes were proposed in this pull request? This PR is to support star expansion in hash. For example, ```SQL val structDf = testData2.select("a", "b").as("record") structDf.select(hash($"") ``` In addition, it refactors the codes for the rule `ResolveStar` and fixes a regression for star expansion in group by when using SQL API. For example, ```SQL SELECT FROM testData2 group by a, b ``` cc cloud-fan Now, the code for star resolution is much cleaner. The coverage is better. Could you check if this refactoring is good? Thanks! #### How was this patch tested? Added a few test cases to cover it. Author: gatorsmile <gatorsmile@gmail.com> Closes #11904 from gatorsmile/starResolution.
*	[SPARK-14014][SQL] Replace existing catalog with SessionCatalog	Andrew Or	2016-03-23	5	-20/+36
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? `SessionCatalog`, introduced in #11750, is a catalog that keeps track of temporary functions and tables, and delegates metastore operations to `ExternalCatalog`. This functionality overlaps a lot with the existing `analysis.Catalog`. As of this commit, `SessionCatalog` and `ExternalCatalog` will no longer be dead code. There are still things that need to be done after this patch, namely: - SPARK-14013: Properly implement temporary functions in `SessionCatalog` - SPARK-13879: Decide which DDL/DML commands to support natively in Spark - SPARK-?????: Implement the ones we do want to support through `SessionCatalog`. - SPARK-?????: Merge SQL/HiveContext ## How was this patch tested? This is largely a refactoring task so there are no new tests introduced. The particularly relevant tests are `SessionCatalogSuite` and `ExternalCatalogSuite`. Author: Andrew Or <andrew@databricks.com> Author: Yin Huai <yhuai@databricks.com> Closes #11836 from andrewor14/use-session-catalog.
*	[SPARK-14078] Streaming Parquet Based FileSink	Michael Armbrust	2016-03-23	3	-0/+180
\| \| \| \| \| \| \| \| \| \|	This PR adds a new `Sink` implementation that writes out Parquet files. In order to correctly handle partial failures while maintaining exactly once semantics, the files for each batch are written out to a unique directory and then atomically appended to a metadata log. When a parquet based `DataSource` is initialized for reading, we first check for this log directory and use it instead of file listing when present. Unit tests are added, as well as a stress test that checks the answer after non-deterministic injected failures. Author: Michael Armbrust <michael@databricks.com> Closes #11897 from marmbrus/fileSink.
*	[SPARK-13809][SQL] State store for streaming aggregations	Tathagata Das	2016-03-23	3	-0/+877
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? In this PR, I am implementing a new abstraction for management of streaming state data - State Store. It is a key-value store for persisting running aggregates for aggregate operations in streaming dataframes. The motivation and design is discussed here. https://docs.google.com/document/d/1-ncawFx8JS5Zyfq1HAEGBx56RDet9wfVp_hDM8ZL254/edit# ## How was this patch tested? - [x] Unit tests - [x] Cluster tests Coverage from unit tests <img width="952" alt="screen shot 2016-03-21 at 3 09 40 pm" src="https://cloud.githubusercontent.com/assets/663212/13935872/fdc8ba86-ef76-11e5-93e8-9fa310472c7b.png"> ## TODO - [x] Fix updates() iterator to avoid duplicate updates for same key - [x] Use Coordinator in ContinuousQueryManager - [x] Plugging in hadoop conf and other confs - [x] Unit tests - [x] StateStore object lifecycle and methods - [x] StateStoreCoordinator communication and logic - [x] StateStoreRDD fault-tolerance - [x] StateStoreRDD preferred location using StateStoreCoordinator - [ ] Cluster tests - [ ] Whether preferred locations are set correctly - [ ] Whether recovery works correctly with distributed storage - [x] Basic performance tests - [x] Docs Author: Tathagata Das <tathagata.das1565@gmail.com> Closes #11645 from tdas/state-store.
*	[SPARK-14075] Refactor MemoryStore to be testable independent of BlockManager	Josh Rosen	2016-03-23	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \| \|	This patch refactors the `MemoryStore` so that it can be tested without needing to construct / mock an entire `BlockManager`. - The block manager's serialization- and compression-related methods have been moved from `BlockManager` to `SerializerManager`. - `BlockInfoManager `is now passed directly to classes that need it, rather than being passed via the `BlockManager`. - The `MemoryStore` now calls `dropFromMemory` via a new `BlockEvictionHandler` interface rather than directly calling the `BlockManager`. This change helps to enforce a narrow interface between the `MemoryStore` and `BlockManager` functionality and makes this interface easier to mock in tests. - Several of the block unrolling tests have been moved from `BlockManagerSuite` into a new `MemoryStoreSuite`. Author: Josh Rosen <joshrosen@databricks.com> Closes #11899 from JoshRosen/reduce-memorystore-blockmanager-coupling.
*	[SPARK-13817][SQL][MINOR] Renames Dataset.newDataFrame to Dataset.ofRows	Cheng Lian	2016-03-24	4	-4/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This PR does the renaming as suggested by marmbrus in [this comment][1]. ## How was this patch tested? Existing tests. [1]: https://github.com/apache/spark/commit/6d37e1eb90054cdb6323b75fb202f78ece604b15#commitcomment-16654694 Author: Cheng Lian <lian@databricks.com> Closes #11889 from liancheng/spark-13817-follow-up.
*	[HOTFIX][SQL] Don't stop ContinuousQuery in quietly	Shixiong Zhu	2016-03-23	2	-25/+12
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? Try to fix a flaky hang ## How was this patch tested? Existing Jenkins test Author: Shixiong Zhu <shixiong@databricks.com> Closes #11909 from zsxwing/hotfix2.
*	[SPARK-14088][SQL] Some Dataset API touch-up	Reynold Xin	2016-03-22	3	-4/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? 1. Deprecated unionAll. It is pretty confusing to have both "union" and "unionAll" when the two do the same thing in Spark but are different in SQL. 2. Rename reduce in KeyValueGroupedDataset to reduceGroups so it is more consistent with rest of the functions in KeyValueGroupedDataset. Also makes it more obvious what "reduce" and "reduceGroups" mean. Previously it was confusing because it could be reducing a Dataset, or just reducing groups. 3. Added a "name" function, which is more natural to name columns than "as" for non-SQL users. 4. Remove "subtract" function since it is just an alias for "except". ## How was this patch tested? All changes should be covered by existing tests. Also added couple test cases to cover "name". Author: Reynold Xin <rxin@databricks.com> Closes #11908 from rxin/SPARK-14088.
*	[MINOR][SQL][DOCS] Update `sql/README.md` and remove some unused imports in ↵	Dongjoon Hyun	2016-03-22	7	-8/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	`sql` module. ## What changes were proposed in this pull request? This PR updates `sql/README.md` according to the latest console output and removes some unused imports in `sql` module. This is done by manually, so there is no guarantee to remove all unused imports. ## How was this patch tested? Manual. Author: Dongjoon Hyun <dongjoon@apache.org> Closes #11907 from dongjoon-hyun/update_sql_module.
*	[SPARK-13401][SQL][TESTS] Fix SQL test warnings.	Yong Tang	2016-03-22	6	-0/+6
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? This fix tries to fix several SQL test warnings under the sql/core/src/test directory. The fixed warnings includes "[unchecked]", "[rawtypes]", and "[varargs]". ## How was this patch tested? All existing tests passed. Author: Yong Tang <yong.tang.github@outlook.com> Closes #11857 from yongtang/SPARK-13401.
*	[HOTFIX][SQL] Add a timeout for 'cq.stop'	Shixiong Zhu	2016-03-22	1	-1/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	## What changes were proposed in this pull request? Fix an issue that DataFrameReaderWriterSuite may hang forever. ## How was this patch tested? Existing tests. Author: Shixiong Zhu <shixiong@databricks.com> Closes #11902 from zsxwing/hotfix.