{"id":17685855,"url":"https://github.com/sepandhaghighi/hadoop","last_synced_at":"2026-05-02T19:37:00.584Z","repository":{"id":66134952,"uuid":"97870841","full_name":"sepandhaghighi/hadoop","owner":"sepandhaghighi","description":"Anagram Python Script In Hadoop","archived":false,"fork":false,"pushed_at":"2017-07-23T09:06:57.000Z","size":46840,"stargazers_count":3,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-03-30T20:29:54.975Z","etag":null,"topics":["anagram","anagram-solver","linux","localhost","map-reduce","python3","script"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/sepandhaghighi.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2017-07-20T19:24:01.000Z","updated_at":"2023-05-10T12:39:43.000Z","dependencies_parsed_at":"2023-03-13T20:30:46.796Z","dependency_job_id":null,"html_url":"https://github.com/sepandhaghighi/hadoop","commit_stats":{"total_commits":36,"total_committers":1,"mean_commits":36.0,"dds":0.0,"last_synced_commit":"28a6bb05c77af0e4f22bb7cb8b386a715164a415"},"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/sepandhaghighi/hadoop","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sepandhaghighi%2Fhadoop","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sepandhaghighi%2Fhadoop/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sepandhaghighi%2Fhadoop/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sepandhaghighi%2Fhadoop/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/sepandhaghighi","download_url":"https://codeload.github.com/sepandhaghighi/hadoop/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sepandhaghighi%2Fhadoop/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":32547651,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-02T19:18:06.202Z","status":"ssl_error","status_checked_at":"2026-05-02T19:16:21.335Z","response_time":132,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["anagram","anagram-solver","linux","localhost","map-reduce","python3","script"],"created_at":"2024-10-24T10:29:19.645Z","updated_at":"2026-05-02T19:37:00.567Z","avatar_url":"https://github.com/sepandhaghighi.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\u003cimg src=\"images/anagram.png\"\u003e\n\u003c/div\u003e\t\t\t\t\n\n----------\n\t\t\t\n\n# Anagram Sorter (Python + Hadoop)\n### Sharif University Of Technology\n### Green Computing Final Project Spring 2017\n\u003chr/\u003e\n\u003cdiv align=\"center\"\u003e\n\u003ca href=\"https://scrutinizer-ci.com/g/sepandhaghighi/hadoop/\"\u003e\u003cimg src=\"https://scrutinizer-ci.com/g/sepandhaghighi/hadoop/badges/quality-score.png?b=master\"\u003e\u003c/a\u003e\n\u003ca href=\"https://www.codacy.com/app/sepand-haghighi/hadoop?utm_source=github.com\u0026amp;utm_medium=referral\u0026amp;utm_content=sepandhaghighi/hadoop\u0026amp;utm_campaign=Badge_Grade\"\u003e\u003cimg src=\"https://api.codacy.com/project/badge/Grade/5ffa02becfca4f20b3ff036769942244\"/\u003e\u003c/a\u003e\n\u003ca href=\"https://scrutinizer-ci.com/g/sepandhaghighi/hadoop/\"\u003e\u003cimg src=\"https://scrutinizer-ci.com/g/sepandhaghighi/hadoop/badges/build.png?b=master\"\u003e\u003c/a\u003e\n\u003ca href=\"https://www.python.org/ftp/python/2.7.13/python-2.7.13.msi\"\u003e\u003cimg src=\"https://img.shields.io/badge/python-2.7%2B-blue.svg\"\u003e\u003c/a\u003e\n\u003c/div\u003e\t\t\t\t\t\n\n## Requirements\n\n\n- [Python 2.7+](https://www.python.org/ftp/python/2.7.13/python-2.7.13.msi \"Python 2.7+\")\n- [Hadoop Framework](http://hadoop.apache.org/ \"Hadoop\")\n- [JDK 7+](www.oracle.com/technetwork/java/javase/downloads/index.html \"JDK 7+\")\n- SSH\n- Git\t\n\n## Install Hadoop\t\t\t\t\t\t\t\n\n\n\n1. Install Java\n\t- ```sudo apt-get update```\n\t- ```sudo apt-get install default-jdk```\n2. \tAdd User\n\t- ```sudo addgroup hadoop```\n\t- ```sudo adduser --ingroup hadoop username```\n\t- ```sudo adduser username sudo```\n3. Install And Config SSH\n\t- ```sudo apt-get install ssh```\n\t- ```su username```\n\t- ``` ssh-keygen -t rsa -P \"\"```\n\t- ``` cat $HOME/.ssh/id_rsa.pub \u003e\u003e $HOME/.ssh/authorized_keys```\n4. Install Hadoop\n\t- [http://hadoop.apache.org/releases.html](http://hadoop.apache.org/releases.html \"Hadoop Download Page\")\n\t- Extract File (without Error !!)\n\t- ```sudo mv * /usr/local/hadoop```\n\t- ```sudo chmod -R 777 /usr/local/hadoop```\n5. Setup Configuration Files\n\t- ```update-alternatives --config java```\n\t- ```nano ~/.bashrc``` and append this items\n\t\t- ```export JAVA_HOME=/usr/lib/jvm/java-7-openjdk-amd64```\n\t\t- ```export HADOOP_INSTALL=/usr/local/hadoop```\n\t\t- ```export PATH=$PATH:$HADOOP_INSTALL/bin```\n\t\t- ```export PATH=$PATH:$HADOOP_INSTALL/sbin```\n\t\t- ```export HADOOP_MAPRED_HOME=$HADOOP_INSTALL```\n\t\t- ```export HADOOP_COMMON_HOME=$HADOOP_INSTALL```\n\t\t- ```export HADOOP_HDFS_HOME=$HADOOP_INSTALL```\n\t\t- ```export YARN_HOME=$HADOOP_INSTALL```\n\t\t- ```export HADOOP_COMMON_LIB_NATIVE_DIR=$HADOOP_INSTALL/lib/native```\n\t\t- ```export HADOOP_OPTS=\"-Djava.library.path=$HADOOP_INSTALL/lib\" ``` and save file\n\t- ```source ~/.bashrc```\n\t- ``` nano /usr/local/hadoop/etc/hadoop/hadoop-env.sh``` append this line\n\t\t- ```export JAVA_HOME=/usr/lib/jvm/java-7-openjdk-amd64```\n\t- ``` sudo mkdir -p /app/hadoop/tmp```\n\t- ``` sudo chown username:hadoop /app/hadoop/tmp```\n\t- ``` nano /usr/local/hadoop/etc/hadoop/core-site.xml``` enter the following in between the \u003cconfiguration\u003e\u003c/configuration\u003e tag\n\t\t- ``` \u003cconfiguration\u003e\n \t\t\t\t\u003cproperty\u003e\n  \t\t\t\t\u003cname\u003ehadoop.tmp.dir\u003c/name\u003e\n  \t\t\t\t\u003cvalue\u003e/app/hadoop/tmp\u003c/value\u003e\n  \t\t\t\t\u003cdescription\u003eA base for other temporary directories.\u003c/description\u003e\n \t\t\t\t\u003c/property\u003e\n\n \t\t\t\t\u003cproperty\u003e\n  \t\t\t\t\u003cname\u003efs.default.name\u003c/name\u003e\n  \t\t\t\t\u003cvalue\u003ehdfs://localhost:54310\u003c/value\u003e\n  \t\t\t\t\u003cdescription\u003eThe name of the default file system.  A URI whose\n  \t\t\t\tscheme and authority determine the FileSystem implementation.  The\n  \t\t\t\turi's scheme determines the config property (fs.SCHEME.impl) naming\n  \t\t\t\tthe FileSystem implementation class.  The uri's authority is used to\n  \t\t\t\tdetermine the host, port, etc. for a filesystem.\u003c/description\u003e\n \t\t\t\t\u003c/property\u003e\n\t\t\t\t\u003c/configuration\u003e ```\n\t- ```cp /usr/local/hadoop/etc/hadoop/mapred-site.xml.template /usr/local/hadoop/etc/hadoop/mapred-site.xml```\n\t- ``` nano /usr/local/hadoop/etc/hadoop/mapred-site.xml``` enter the following in between the \u003cconfiguration\u003e\u003c/configuration\u003e tag\n\t\t- ```\u003cconfiguration\u003e\n \t\t\t\u003cproperty\u003e\n  \t\t\t\u003cname\u003emapred.job.tracker\u003c/name\u003e\n  \t\t\t\u003cvalue\u003elocalhost:54311\u003c/value\u003e\n  \t\t\t\u003cdescription\u003eThe host and port that the MapReduce job tracker runs\n  \t\t\tat.  If \"local\", then jobs are run in-process as a single map\n  \t\t\tand reduce task.\n  \t\t\t\u003c/description\u003e\n \t\t\t\u003c/property\u003e\n\t\t\t\u003c/configuration\u003e ```\n\t- ```sudo mkdir -p /usr/local/hadoop_store/hdfs/namenode```\n\t- ```sudo mkdir -p /usr/local/hadoop_store/hdfs/datanode```\n\t- ```sudo chown -R username:hadoop /usr/local/hadoop_store```\n\t- ```nano /usr/local/hadoop/etc/hadoop/hdfs-site.xml``` enter the following in between the \u003cconfiguration\u003e\u003c/configuration\u003e tag\n\t\t- ```\u003cconfiguration\u003e\n \t\t\t\u003cproperty\u003e\n  \t\t\t\u003cname\u003edfs.replication\u003c/name\u003e\n  \t\t\t\u003cvalue\u003e1\u003c/value\u003e\n  \t\t\t\u003cdescription\u003eDefault block replication.\n  \t\t\tThe actual number of replications can be specified when the file is created.\n  \t\t\tThe default is used if replication is not specified in create time.\n  \t\t\t\u003c/description\u003e\n \t\t\t\u003c/property\u003e\n \t\t\t\u003cproperty\u003e\n   \t\t\t\u003cname\u003edfs.namenode.name.dir\u003c/name\u003e\n   \t\t\t\u003cvalue\u003efile:/usr/local/hadoop_store/hdfs/namenode\u003c/value\u003e\n \t\t\t\u003c/property\u003e\n \t\t\t\u003cproperty\u003e\n   \t\t\t\u003cname\u003edfs.datanode.data.dir\u003c/name\u003e\n   \t\t\t\u003cvalue\u003efile:/usr/local/hadoop_store/hdfs/datanode\u003c/value\u003e\n \t\t\t\u003c/property\u003e\n\t\t\t\u003c/configuration\u003e```\n6. Start Hadoop\n\t- ```hadoop namenode -format```\n\t- ```cd /usr/local/hadoop/sbin```\n\t- ```start-all.sh```\n\t- In Connection Refused Problems Check [https://wiki.apache.org/hadoop/ConnectionRefused](https://wiki.apache.org/hadoop/ConnectionRefused \"Connection Refused\")\n\n\n\n## Clone Repo\t\t\t\n- ```cd /home/username```\n- ```git clone https://github.com/sepandhaghighi/hadoop```\n- ``` chmod -R 777 /home/username/hadoop```\n\n## HDFS Commands\n\t\n1. Add input to filesystem --\u003e \t``` hadoop fs -put inputfile inputfile```\n2. Read output file --\u003e ``` hadoop fs -cat /output/part-00000``` \n3. Copy output file --\u003e ``` hadoop fs -get /output/part-00000 /home/username/hadoop/output.txt```\n4. Remove output folder --\u003e ``` hadoop fs -rmr /output/```  \t\t\n\n\n## Run Map/Reduce\n``` hadoop jar /usr/local/hadoop/share/hadoop/tools/lib/hadoop-streaming-2.8.0.jar -input input.txt -mapper /home/hduser/hadoop/mapper.py -reducer /home/hduser/hadoop/reducer.py -output /output```\n\n\n## Samples \u0026 Screenshots\n\n[Download Run Full Video](http://www.shaghighi.ir/hadoop/Full.mkv \"Download\") 43MB ( This video recorded by [simplescreenrecorder](http://www.maartenbaert.be/simplescreenrecorder/ \"simplescreenrecorder\") )\t\t\t\n\nInput.txt and output.txt in data folder\n\n\u003cdiv align=\"center\"\u003e\n\u003cimg src=\"images/input.jpg\"\u003e\n\u003cp\u003eInput file\u003c/p\u003e\n\u003cimg src=\"images/output.jpg\"\u003e\n\u003cp\u003eOutput file\u003c/p\u003e\n\u003cimg src=\"images/screen1.png\"\u003e\n\u003cp\u003eScreenshot 1\u003c/p\u003e\n\u003cimg src=\"images/screen2.png\"\u003e\n\u003cp\u003eScreenshot 2\u003c/p\u003e\n\u003cimg src=\"images/screen3.png\"\u003e\n\u003cp\u003eScreenshot 3\u003c/p\u003e\n\u003c/div\u003e\n\n\n\t\t\t\t\n\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsepandhaghighi%2Fhadoop","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsepandhaghighi%2Fhadoop","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsepandhaghighi%2Fhadoop/lists"}