summaryrefslogtreecommitdiff
path: root/site/docs/0.6.1/running-on-yarn.html
diff options
context:
space:
mode:
Diffstat (limited to 'site/docs/0.6.1/running-on-yarn.html')
-rw-r--r--site/docs/0.6.1/running-on-yarn.html205
1 files changed, 205 insertions, 0 deletions
diff --git a/site/docs/0.6.1/running-on-yarn.html b/site/docs/0.6.1/running-on-yarn.html
new file mode 100644
index 000000000..67c70a0c8
--- /dev/null
+++ b/site/docs/0.6.1/running-on-yarn.html
@@ -0,0 +1,205 @@
+<!DOCTYPE html>
+<!--[if lt IE 7]> <html class="no-js lt-ie9 lt-ie8 lt-ie7"> <![endif]-->
+<!--[if IE 7]> <html class="no-js lt-ie9 lt-ie8"> <![endif]-->
+<!--[if IE 8]> <html class="no-js lt-ie9"> <![endif]-->
+<!--[if gt IE 8]><!--> <html class="no-js"> <!--<![endif]-->
+ <head>
+ <meta charset="utf-8">
+ <meta http-equiv="X-UA-Compatible" content="IE=edge,chrome=1">
+ <title>Launching Spark on YARN - Spark 0.6.1 Documentation</title>
+ <meta name="description" content="">
+
+ <link rel="stylesheet" href="css/bootstrap.min.css">
+ <style>
+ body {
+ padding-top: 60px;
+ padding-bottom: 40px;
+ }
+ </style>
+ <meta name="viewport" content="width=device-width">
+ <link rel="stylesheet" href="css/bootstrap-responsive.min.css">
+ <link rel="stylesheet" href="css/main.css">
+
+ <script src="js/vendor/modernizr-2.6.1-respond-1.1.0.min.js"></script>
+
+ <link rel="stylesheet" href="css/pygments-default.css">
+ <script type="text/javascript">
+ var _gaq = _gaq || [];
+ _gaq.push(['_setAccount', 'UA-32518208-1']);
+ _gaq.push(['_trackPageview']);
+
+ (function() {
+ var ga = document.createElement('script'); ga.type = 'text/javascript'; ga.async = true;
+ ga.src = ('https:' == document.location.protocol ? 'https://ssl' : 'http://www') + '.google-analytics.com/ga.js';
+ var s = document.getElementsByTagName('script')[0]; s.parentNode.insertBefore(ga, s);
+ })();
+ </script>
+ </head>
+ <body>
+ <!--[if lt IE 7]>
+ <p class="chromeframe">You are using an outdated browser. <a href="http://browsehappy.com/">Upgrade your browser today</a> or <a href="http://www.google.com/chromeframe/?redirect=true">install Google Chrome Frame</a> to better experience this site.</p>
+ <![endif]-->
+
+ <!-- This code is taken from http://twitter.github.com/bootstrap/examples/hero.html -->
+
+ <div class="navbar navbar-fixed-top" id="topbar">
+ <div class="navbar-inner">
+ <div class="container">
+ <div class="brand"><a href="index.html">
+ <img src="img/spark-logo-77x50px-hd.png" /></a><span class="version">0.6.1</span>
+ </div>
+ <ul class="nav">
+ <!--TODO(andyk): Add class="active" attribute to li some how.-->
+ <li><a href="index.html">Overview</a></li>
+
+ <li class="dropdown">
+ <a href="#" class="dropdown-toggle" data-toggle="dropdown">Programming Guides<b class="caret"></b></a>
+ <ul class="dropdown-menu">
+ <li><a href="quick-start.html">Quick Start</a></li>
+ <li><a href="scala-programming-guide.html">Scala</a></li>
+ <li><a href="java-programming-guide.html">Java</a></li>
+ </ul>
+ </li>
+
+ <li><a href="api/core/index.html">API (Scaladoc)</a></li>
+
+ <li class="dropdown">
+ <a href="#" class="dropdown-toggle" data-toggle="dropdown">Deploying<b class="caret"></b></a>
+ <ul class="dropdown-menu">
+ <li><a href="ec2-scripts.html">Amazon EC2</a></li>
+ <li><a href="spark-standalone.html">Standalone Mode</a></li>
+ <li><a href="running-on-mesos.html">Mesos</a></li>
+ <li><a href="running-on-yarn.html">YARN</a></li>
+ </ul>
+ </li>
+
+ <li class="dropdown">
+ <a href="api.html" class="dropdown-toggle" data-toggle="dropdown">More<b class="caret"></b></a>
+ <ul class="dropdown-menu">
+ <li><a href="configuration.html">Configuration</a></li>
+ <li><a href="tuning.html">Tuning Guide</a></li>
+ <li><a href="bagel-programming-guide.html">Bagel (Pregel on Spark)</a></li>
+ <li><a href="contributing-to-spark.html">Contributing to Spark</a></li>
+ </ul>
+ </li>
+ </ul>
+ <!--<p class="navbar-text pull-right"><span class="version-text">v0.6.1</span></p>-->
+ </div>
+ </div>
+ </div>
+
+ <div class="container" id="content">
+ <h1 class="title">Launching Spark on YARN</h1>
+
+ <p>Experimental support for running over a <a href="http://hadoop.apache.org/docs/r2.0.2-alpha/hadoop-yarn/hadoop-yarn-site/YARN.html">YARN (Hadoop
+NextGen)</a>
+cluster was added to Spark in version 0.6.0. Because YARN depends on version
+2.0 of the Hadoop libraries, this currently requires checking out a separate
+branch of Spark, called <code>yarn</code>, which you can do as follows:</p>
+
+<pre><code>git clone git://github.com/mesos/spark
+cd spark
+git checkout -b yarn --track origin/yarn
+</code></pre>
+
+<h1 id="preparations">Preparations</h1>
+
+<ul>
+ <li>In order to distribute Spark within the cluster, it must be packaged into a single JAR file. This can be done by running <code>sbt/sbt assembly</code></li>
+ <li>Your application code must be packaged into a separate JAR file.</li>
+</ul>
+
+<p>If you want to test out the YARN deployment mode, you can use the current Spark examples. A <code>spark-examples_2.9.2-0.6.1</code> file can be generated by running <code>sbt/sbt package</code>. NOTE: since the documentation you&rsquo;re reading is for Spark version 0.6.1, we are assuming here that you have downloaded Spark 0.6.1 or checked it out of source control. If you are using a different version of Spark, the version numbers in the jar generated by the sbt package command will obviously be different.</p>
+
+<h1 id="launching-spark-on-yarn">Launching Spark on YARN</h1>
+
+<p>The command to launch the YARN Client is as follows:</p>
+
+<pre><code>SPARK_JAR=&lt;SPARK_YAR_FILE&gt; ./run spark.deploy.yarn.Client \
+ --jar &lt;YOUR_APP_JAR_FILE&gt; \
+ --class &lt;APP_MAIN_CLASS&gt; \
+ --args &lt;APP_MAIN_ARGUMENTS&gt; \
+ --num-workers &lt;NUMBER_OF_WORKER_MACHINES&gt; \
+ --worker-memory &lt;MEMORY_PER_WORKER&gt; \
+ --worker-cores &lt;CORES_PER_WORKER&gt;
+</code></pre>
+
+<p>For example:</p>
+
+<pre><code>SPARK_JAR=./core/target/spark-core-assembly-0.6.1.jar ./run spark.deploy.yarn.Client \
+ --jar examples/target/scala-2.9.2/spark-examples_2.9.2-0.6.1.jar \
+ --class spark.examples.SparkPi \
+ --args standalone \
+ --num-workers 3 \
+ --worker-memory 2g \
+ --worker-cores 2
+</code></pre>
+
+<p>The above starts a YARN Client programs which periodically polls the Application Master for status updates and displays them in the console. The client will exit once your application has finished running.</p>
+
+<h1 id="important-notes">Important Notes</h1>
+
+<ul>
+ <li>When your application instantiates a Spark context it must use a special &ldquo;standalone&rdquo; master url. This starts the scheduler without forcing it to connect to a cluster. A good way to handle this is to pass &ldquo;standalone&rdquo; as an argument to your program, as shown in the example above.</li>
+ <li>YARN does not support requesting container resources based on the number of cores. Thus the numbers of cores given via command line arguments cannot be guaranteed.</li>
+</ul>
+
+ <!-- Main hero unit for a primary marketing message or call to action -->
+ <!--<div class="hero-unit">
+ <h1>Hello, world!</h1>
+ <p>This is a template for a simple marketing or informational website. It includes a large callout called the hero unit and three supporting pieces of content. Use it as a starting point to create something more unique.</p>
+ <p><a class="btn btn-primary btn-large">Learn more &raquo;</a></p>
+ </div>-->
+
+ <!-- Example row of columns -->
+ <!--<div class="row">
+ <div class="span4">
+ <h2>Heading</h2>
+ <p>Donec id elit non mi porta gravida at eget metus. Fusce dapibus, tellus ac cursus commodo, tortor mauris condimentum nibh, ut fermentum massa justo sit amet risus. Etiam porta sem malesuada magna mollis euismod. Donec sed odio dui. </p>
+ <p><a class="btn" href="#">View details &raquo;</a></p>
+ </div>
+ <div class="span4">
+ <h2>Heading</h2>
+ <p>Donec id elit non mi porta gravida at eget metus. Fusce dapibus, tellus ac cursus commodo, tortor mauris condimentum nibh, ut fermentum massa justo sit amet risus. Etiam porta sem malesuada magna mollis euismod. Donec sed odio dui. </p>
+ <p><a class="btn" href="#">View details &raquo;</a></p>
+ </div>
+ <div class="span4">
+ <h2>Heading</h2>
+ <p>Donec sed odio dui. Cras justo odio, dapibus ac facilisis in, egestas eget quam. Vestibulum id ligula porta felis euismod semper. Fusce dapibus, tellus ac cursus commodo, tortor mauris condimentum nibh, ut fermentum massa justo sit amet risus.</p>
+ <p><a class="btn" href="#">View details &raquo;</a></p>
+ </div>
+ </div>
+
+ <hr>-->
+
+ <!--<footer>
+ <p></p>
+ </footer>-->
+
+ </div> <!-- /container -->
+
+ <script src="js/vendor/jquery-1.8.0.min.js"></script>
+ <script src="js/vendor/bootstrap.min.js"></script>
+ <script src="js/main.js"></script>
+
+ <!-- A script to fix internal hash links because we have an overlapping top bar.
+ Based on https://github.com/twitter/bootstrap/issues/193#issuecomment-2281510 -->
+ <script>
+ $(function() {
+ function maybeScrollToHash() {
+ if (window.location.hash && $(window.location.hash).length) {
+ var newTop = $(window.location.hash).offset().top - $('#topbar').height() - 5;
+ $(window).scrollTop(newTop);
+ }
+ }
+ $(window).bind('hashchange', function() {
+ maybeScrollToHash();
+ });
+ // Scroll now too in case we had opened the page on a hash, but wait 1 ms because some browsers
+ // will try to do *their* initial scroll after running the onReady handler.
+ setTimeout(function() { maybeScrollToHash(); }, 1)
+ })
+ </script>
+
+ </body>
+</html>