{"id":17066611,"url":"https://github.com/curran/setuphadoop","last_synced_at":"2025-10-05T06:52:20.635Z","repository":{"id":27987434,"uuid":"31481381","full_name":"curran/setupHadoop","owner":"curran","description":"Shell scripts and instructions for setting up Hadoop on a cluster.","archived":false,"fork":false,"pushed_at":"2015-03-11T17:19:50.000Z","size":157,"stargazers_count":1,"open_issues_count":0,"forks_count":8,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-05-07T03:43:12.313Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Shell","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/curran.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2015-03-01T00:04:37.000Z","updated_at":"2021-10-23T20:13:54.000Z","dependencies_parsed_at":"2022-08-26T11:20:22.615Z","dependency_job_id":null,"html_url":"https://github.com/curran/setupHadoop","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/curran%2FsetupHadoop","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/curran%2FsetupHadoop/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/curran%2FsetupHadoop/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/curran%2FsetupHadoop/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/curran","download_url":"https://codeload.github.com/curran/setupHadoop/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":252810272,"owners_count":21807759,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-10-14T11:07:28.493Z","updated_at":"2025-10-05T06:52:20.537Z","avatar_url":"https://github.com/curran.png","language":"Shell","funding_links":[],"categories":[],"sub_categories":[],"readme":"# setupHadoop\nShell scripts and instructions for setting up Hadoop on a cluster. Intended for use with fresh instances of Ubuntu Server 14.04.1 LTS on Amazon EC2. The instructions below show how to set up a 2 node cluster.\n\n## Creating Instances\nFirst, create two virtual machines using the Amazon Web Interface.\n\n * Click through Services -\u003e EC2 -\u003e Instances -\u003e Launch Instance.\n * Select \"Ubuntu Server 14.04 LTS (HVM), SSD Volume Type\".\n * Click \"Review and Launch\", then click \"Launch\".\n * On the dialog \"Select an existing key pair...\", select \"Create a new pair\"\n * Enter a name for the key pair (I used \"cloudTest\")\n * Click \"Download Key Pair\", which will download \"cloudTest.pem\"\n * Click \"Launch Instance\"\n * Click \"View Instances\" to get back to the instances page.\n * Repeat this process a second time, this time using the existing key pair \"cloudTest\", to create a second instance.\n\nOnce you have created two instances, you can name them by clicking on the empty \"Name\" field. For example, you can name them \"Master\" and \"Slave\", to help keep track of which is which. After doing this, you should see a listing like this:\n\n![Instances](http://curran.github.io/images/setupHadoop/instances.png)\n\n## Connecting to an Instance\n\nIn a terminal, go to the directory where \"cloudTest.pem\" is.\n\n`cd ~/Downloads`\n\nMake sure the key file is not visible to others via ssh.\n\n`chmod 400 cloudTest.pem`\n\nIf you don't do this, then you'll see this error later:\n\n```\n@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\n@         WARNING: UNPROTECTED PRIVATE KEY FILE!          @\n@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\n```\n\nFind the public IP of your instance in the AWS Web Interface.\n\nConnect to the instance via SSH using the following command.\n\n`ssh -i cloudTest.pem ubuntu@\u003cyour IP here\u003e`\n\nFor example,\n\n`ssh -i cloudTest.pem ubuntu@54.67.81.195`\n\nType \"yes\" at the prompt `Are you sure you want to continue connecting (yes/no)? yes`\n\nOnce logged in, you can check what your Ubuntu version is by running\n\n`lsb_release -a`\n\n## Set Up Hadoop\n\nCopy and paste this entire script into your console after logging into an instance. This should be done for all machines in the cluster.\n\n```\nsudo apt-get update;\\\nsudo apt-get install -y git default-jdk;\\\ncurl -O http://mirror.cogentco.com/pub/apache/hadoop/common/hadoop-2.6.0/hadoop-2.6.0.tar.gz;\\\ntar xfz hadoop-2.6.0.tar.gz;\\\nsudo mv hadoop-2.6.0 /usr/local/hadoop;\\\nrm hadoop-2.6.0.tar.gz;\\\nssh-keygen -t rsa -P '' -f ~/.ssh/id_rsa;\\\ncat ~/.ssh/id_rsa.pub \u003e\u003e ~/.ssh/authorized_keys;\\\necho export JAVA_HOME=/usr/lib/jvm/java-7-openjdk-amd64 \u003e\u003e ~/.bashrc;\\\necho export HADOOP_CONF_DIR=/usr/local/hadoop/etc/hadoop \u003e\u003e ~/.bashrc;\\\necho export YARN_CONF_DIR=/usr/local/hadoop/etc/hadoop \u003e\u003e ~/.bashrc;\\\necho export PATH=\\$PATH:/usr/local/hadoop/bin \u003e\u003e ~/.bashrc;\\\necho export PATH=\\$PATH:/usr/local/hadoop/sbin \u003e\u003e ~/.bashrc;\\\nsource ~/.bashrc\n```\n\nFor the master, copy these config files into Hadoop directory.\n\n```\ngit clone https://github.com/curran/setupHadoop.git; \\\ncd setupHadoop; \\\ncp -r master/* $HADOOP_CONF_DIR\n```\n\nWe will want to access the Web Interfaces for HDSF and Yarn, which are blocked by default with AWS. Allow all traffic into the master node by scrolling to the right in the AWS instance listing page, clicking the link in the \"Security Groups\" column -\u003e \"Inbound\" tab -\u003e \"Edit\" button -\u003e \"Add Rule\" button -\u003e change \"Custom TCP Rule\" to \"All TCP\" -\u003e \"Save\" button\n\n# Setting up a Cluster\n\nThe approach to setting up a cluster is to first set up many machines as independent single-node Hadoop clusters, then reconfigure them such that one is a master and the others are slaves. Then you need to \"start the cluster\" by starting HDFS and YARN from the _master node only_. The master node will connect to slave nodes via SSH and start the appropriate processes on each, namely DataNode for HDFS and ResourceManager for YARN.\n\nSince the master node needs to communicate to slaves over SSH, we need to add the public key of the master machine to the list of allowed hosts in the slave machine(s). To do this:\n\n * Log into the master node over SSH\n * Execute `cat ~/.ssh/id_rsa.pub`\n * Copy the output to the clipboard\n * Log into the slave node over SSH\n * Edit the file `~/.ssh/authorized_keys`\n * Paste the contents of the clipboard into a new line of the file\n * Save and close the file\n\n(There very well may be a nicer way of doing this, please send a pull request if there is!)\n\n## Getting HDFS to Work\n\nChoose a single machine to be the master node for HDFS, which will run the NameNode daemon. All other machines will be slaves, and will run the DataNode daemon.\n\nTo set up a slave machine, do the following:\n\nEdit the file `/usr/local/hadoop/etc/hadoop/core-site.xml`. Change the `fs.defaultFS` value to use the IP of the master node (found using `ifconfig` ran on the master node). This IP is also listed in the Amazon Web Interface, called \"Private IP\".\n\nThe file should look something like this:\n```\n\u003cconfiguration\u003e\n  \u003cproperty\u003e\n    \u003cname\u003efs.default.name\u003c/name\u003e\n    \u003cvalue\u003ehdfs://52.11.95.33:9000\u003c/value\u003e\n  \u003c/property\u003e\n\u003c/configuration\u003e\n```\n\nThe following commands are defined in `/usr/local/hadoop/bin` and `/usr/local/hadoop/sbin`. They are available as commands to execute because these paths were added to `$PATH` by `setupHadoop.sh`.\n\nFormat the file system.\n\n`hdfs namenode -format` WARNING - if you want to set up a cluster, make sure all configurations are set before executing this. If this is executed with config for a single machine, then it seems to break the state of the system and you cannot get the full cluster working without starting from scratch.\n\n`start-dfs.sh` Start HDFS. This will launch the NameNode and DataNode Hadoop daemons on the master AND slaves (via SSH).\n\n`stop-dfs.sh` Stop HDFS on master and slaves.\n\nYou can always check to see which daemons are running on a given node by executing `jps`. The output will look something like this:\n\n```\n$ jps\n11296 NameNode\n11453 DataNode\n11768 Jps\n```\n\nNow you should see the following page on port `50070` of your master node:\n\n![workingHDFS](http://curran.github.io/images/setupHadoop/workingHDFS.png)\nA working HDFS cluster with 2 DataNodes. (on port `50070`)\n\n# Getting YARN to work\n\nTo get YARN working, edit the file `/usr/local/hadoop/etc/hadoop/yarn-site.xml` to set the IP of the Resource Manager (on both master and slave machines). In my case this is the same as the HDFS name node.\n```\n\u003cconfiguration\u003e\n  \u003cproperty\u003e\n    \u003cname\u003eyarn.resourcemanager.hostname\u003c/name\u003e\n    \u003cvalue\u003e52.11.95.33\u003c/value\u003e\n  \u003c/property\u003e\n\u003c/configuration\u003e\n```\n\nNow run\n\n`start-yarn.sh` to start,\n`stop-yarn.sh` to stop.\n\nAfter starting YARN, you should see the following page on port `8088` of your master node:\n\n![workingYARN](http://curran.github.io/images/setupHadoop/workingYARN.png)\nA working YARN cluster with 2 NodeManagers. (on port `8088`)\n\n### Reformatting HDFS\n\nIf your HDFS somehow gets corrupted, you can reformat everything like this:\n\n```\nstop-yarn.sh\nstop-dfs.sh\nrm -r -f /tmp/hadoop-ubuntu/* # do this on all machines\nhdfs namenode -format # do this on master\nstart-dfs.sh\nstart-yarn.sh\n```\n\n### Notes\n\nWhen trying to run a Spark shell in YARN with the following command\n\n`./bin/spark-shell --master yarn-client`\n\nThe YARN application initializes (I can see it in the YARN Web UI), but I get the following error after about 10 seconds:\n\n```\n15/03/04 00:32:35 WARN remote.ReliableDeliverySupervisor: Association with remote system [akka.tcp://sparkYarnAM@ip-172-31-4-232.us-west-1.compute.internal:57241] has failed, address is now gated for [5000] ms. Reason is: [Disassociated].\n```\n\nThis seems to be [related to RAM capacity](http://stackoverflow.com/questions/28671171/spark-shell-cannot-connect-to-yarn). I am using AWS Micro instances that have only 1GB of RAM. Using the following command will show you memory usage every second on a given machine:\n\n`watch -n 1 free -m`\n\nThe free memory was falling to around 60MB when the YARN connection gets \"dissociated\".\n\nTo get the Spark Shell to work on YARN on my Mac laptop, I experienced the same error as described in the blog post [YARN Job Problem: Application application_** failed 1 times due to AM Container for XX exited with exitCode: 127](https://cloudcelebrity.wordpress.com/2014/01/31/yarn-job-problem-application-application_-failed-1-times-due-to-am-container-for-xx-exited-with-exitcode-127/) and applied his solution:\n\n`sudo ln -s /usr/bin/java /bin/java`\n\nWhat finally worked in the Spark Shell in YARN-client mode:\n\n```\nval data = sc.textFile(\"hdfs://localhost:9000/data/adult/data.csv\")\ndata.first()\n```\n\nNote the port 9000 in the hdfs URL. If no port is specified, the system assumes post 8020 (as listed in the [default HDFS ports](https://ambari.apache.org/1.2.3/installing-hadoop-using-ambari/content/reference_chap2_1.html)), which is not the default used by HDFS. The default is 9000.\n\nDraws from\n\n * https://www.digitalocean.com/community/tutorials/how-to-install-hadoop-on-ubuntu-13-10\n * http://hadoop.apache.org/docs/current/hadoop-project-dist/hadoop-common/ClusterSetup.html\n * http://hadoop.apache.org/docs/current/hadoop-project-dist/hadoop-common/SingleCluster.html\n * http://www.alexjf.net/blog/distributed-systems/hadoop-yarn-installation-definitive-guide/\n * https://help.ubuntu.com/community/CheckingYourUbuntuVersion\n * http://www.michael-noll.com/tutorials/running-hadoop-on-ubuntu-linux-multi-node-cluster/\n * https://www.youtube.com/watch?v=3rb111Z9TVI\n\nBy Curran Kelleher March 2015\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcurran%2Fsetuphadoop","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcurran%2Fsetuphadoop","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcurran%2Fsetuphadoop/lists"}