[SPARK-4019] [SPARK-3740] Fix MapStatus compression bug that could lead to empty results or Snappy errors

This commit fixes a bug in MapStatus that could cause jobs to wrongly return empty results if those jobs contained stages with more than 2000 partitions where most of those partitions were empty. For jobs with > 2000 partitions, MapStatus uses HighlyCompressedMapStatus, which only stores the average size of blocks. If the average block size is zero, then this will cause all blocks to be reported as empty, causing BlockFetcherIterator to mistakenly skip them. For example, this would return an empty result: sc.makeRDD(0 until 10, 1000).repartition(2001).collect() This can also lead to deserialization errors (e.g. Snappy decoding errors) for jobs with > 2000 partitions where the average block size is non-zero but there is at least one empty block. In this case, the BlockFetcher attempts to fetch empty blocks and fails when trying to deserialize them. The root problem here is that MapStatus has a (previously undocumented) correctness property that was violated by HighlyCompressedMapStatus: If a block is non-empty, then getSizeForBlock must be non-zero. I fixed this by modifying HighlyCompressedMapStatus to store the average size of _non-empty_ blocks and to use a compressed bitmap to track which blocks are empty. I also removed a test which was broken as originally written: it attempted to check that HighlyCompressedMapStatus's size estimation error was < 10%, but this was broken because HighlyCompressedMapStatus is only used for map statuses with > 2000 partitions, but the test only created 50. Author: Josh Rosen <joshrosen@databricks.com> Closes #2866 from JoshRosen/spark-4019 and squashes the following commits: fc8b490 [Josh Rosen] Roll back hashset change, which didn't improve performance. 5faa0a4 [Josh Rosen] Incorporate review feedback c8b8cae [Josh Rosen] Two performance fixes: 3b892dd [Josh Rosen] Address Reynold's review comments ba2e71c [Josh Rosen] Add missing newline 609407d [Josh Rosen] Use Roaring Bitmap to track non-empty blocks. c23897a [Josh Rosen] Use sets when comparing collect() results 91276a3 [Josh Rosen] [SPARK-4019] Fix MapStatus compression bug that could lead to empty results.
author: Josh Rosen <joshrosen@databricks.com> 2014-10-23 16:39:32 -0700
committer: Patrick Wendell <pwendell@gmail.com> 2014-10-23 16:39:32 -0700
commit: 83b7a1c6503adce1826fc537b4db47e534da5cae (patch)
tree: 1e0f4b3c78c13db94d24b0c9b011a6148b5d3433 /pom.xml
parent: 222fa47f0dfd6c53aac513655a519521d9396e72 (diff)
download: spark-83b7a1c6503adce1826fc537b4db47e534da5cae.tar.gz
spark-83b7a1c6503adce1826fc537b4db47e534da5cae.tar.bz2
spark-83b7a1c6503adce1826fc537b4db47e534da5cae.zip
1 files changed, 5 insertions, 0 deletions
diff --git a/pom.xml b/pom.xml
index 288bbf1114..a7e71f9ca5 100644
--- a/pom.xml
+++ b/pom.xml
@@ -429,6 +429,11 @@
         </exclusions>
       </dependency>
       <dependency>
+        <groupId>org.roaringbitmap</groupId>
+        <artifactId>RoaringBitmap</artifactId>
+        <version>0.4.1</version>
+      </dependency>
+      <dependency>
         <groupId>commons-net</groupId>
         <artifactId>commons-net</artifactId>
         <version>2.2</version>
author	Josh Rosen <joshrosen@databricks.com>	2014-10-23 16:39:32 -0700
committer	Patrick Wendell <pwendell@gmail.com>	2014-10-23 16:39:32 -0700
commit	83b7a1c6503adce1826fc537b4db47e534da5cae (patch)
tree	1e0f4b3c78c13db94d24b0c9b011a6148b5d3433 /pom.xml
parent	222fa47f0dfd6c53aac513655a519521d9396e72 (diff)
download	spark-83b7a1c6503adce1826fc537b4db47e534da5cae.tar.gz spark-83b7a1c6503adce1826fc537b4db47e534da5cae.tar.bz2 spark-83b7a1c6503adce1826fc537b4db47e534da5cae.zip