{"id":38710810,"url":"https://github.com/heidsoft/heidsoft-bigdata","last_synced_at":"2026-01-17T11:00:01.003Z","repository":{"id":147696931,"uuid":"11424293","full_name":"heidsoft/heidsoft-bigdata","owner":"heidsoft","description":"大数据研究与应用","archived":false,"fork":false,"pushed_at":"2014-10-07T14:04:25.000Z","size":244,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"master","last_synced_at":"2024-04-15T00:43:07.501Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/heidsoft.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2013-07-15T13:52:30.000Z","updated_at":"2017-01-05T04:54:29.000Z","dependencies_parsed_at":null,"dependency_job_id":"f7ae29b8-0941-43da-8390-4cac96346f57","html_url":"https://github.com/heidsoft/heidsoft-bigdata","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/heidsoft/heidsoft-bigdata","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/heidsoft%2Fheidsoft-bigdata","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/heidsoft%2Fheidsoft-bigdata/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/heidsoft%2Fheidsoft-bigdata/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/heidsoft%2Fheidsoft-bigdata/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/heidsoft","download_url":"https://codeload.github.com/heidsoft/heidsoft-bigdata/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/heidsoft%2Fheidsoft-bigdata/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28506593,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-17T10:25:30.148Z","status":"ssl_error","status_checked_at":"2026-01-17T10:25:29.718Z","response_time":85,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2026-01-17T11:00:00.831Z","updated_at":"2026-01-17T11:00:00.980Z","avatar_url":"https://github.com/heidsoft.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"#Cloudera’s VM\n\tThese virtual machines make it easy to get started with CDH (\n\tCloudera’s 100% open source Hadoop platform that includes \n\tImpala, Search, Spark, and more) and Cloudera Manager. \n\tThey come complete with everything you need to learn Hadoop.\n\n#Important\n\tThese are 64-bit VMs. They requires a 64-bit host OS and a virtualization product that can support a 64-bit guest OS.\n\tTo use a VMware VM, you must use a player compatible with WorkStation 8.x or higher: \n\tPlayer 4.x or higher, ESXi 5.x or higher, or Fusion 4.x or higher. \n\tOlder versions of WorkStation can be used to create a new VM using the same virtual disk (VMDK file),\n\tbut some features in VMware Tools won't be available.\n\tThe VM and file size vary according to the CDH version as follows:\n\n\tCDH and Cloudera Manager Version\tVM Size\tFile Size\n\tCDH 5 and Cloudera Manager 5\t8 GB\t3 GB\n\n#CDH 5\t\n\tTo learn more about CDH 5, see the CDH 5 documentation.\n\tFor the latest important information about new features, incompatible changes, and known issues in CDH 5, see the CDH 5 Release Notes.\n\tFor information on the versions of the components in the latest release of CDH 5, and links to each project's changes files and release notes, see the packaging section of CDH Version and Packaging Information.\n\tCloudera Manager 5\n\tAs part of the boot process, the VM automatically launches Cloudera Manager and configures Flume, HBase, HDFS, Hive, Hue, Impala, Key-Value Store Indexer, Oozie, Solr, Sqoop, Spark, YARN, and ZooKeeper services. Only the HDFS, Hive, Hue, YARN, and ZooKeeper services are started automatically. The remaining services are not automatically started to conserve RAM but can be launched manually.\n\tYou can start or reconfigure any installed services using Cloudera Manager, which is automatically displayed when the VM starts. Standard CDH 5 command-line tools are available for use as well.\n\tFor more information about Cloudera Manager including installation, configuration, and usage instructions, see the Cloudera Manager 5 documentation.\n\n#Accounts\n\tOnce you launch the VM, you are automatically logged in as the cloudera user. The account details are:\n\tusername: cloudera\n\tpassword: cloudera\n\tThe cloudera account has sudo privileges in the VM. The root account password is cloudera.\n\tHue and Cloudera Manager use the same credentials.\n\n##QuickStart VMware Image\n\tTo launch the VMware image, you will either need VMware Player for Windows and Linux, or VMware Fusion for Mac. Note that VMware Fusion only works on Intel architectures, so older Macs with PowerPC processors cannot run the QuickStart VM.\n\n##QuickStart VirtualBox Image\n\tSome users have reported problems running CentOS 6.2 in VirtualBox. If a kernel panic occurs while the VirtualBox VM is booting, you can try working around this problem by opening the Settings \u003e System \u003e Motherboard tab, and selecting ICH9 instead of PIIX3 for the chip set. If you have not already done so, you must also enable I/O APIC on the same tab.\n\n##QuickStart KVM Image\n\tThe KVM image provides a raw disk image that can be used by many hypervisors. Configure machines that use this image with sufficient RAM. See Cloudera QuickStart VM for the VM size requirements.\n\t\n#组件说明\n##Flume\n\tFlume 从几乎所有来源收集数据并将这些数据聚合到永久性存储（如 HDFS）中。\n\t\n##HBase\n\tApache HBase 提供对大型数据集的随机、实时的读/写访问权限（需要 HDFS 和 ZooKeeper）。\n\t\n##HDFS\n\tApache Hadoop 分布式文件系统 (HDFS) 是 Hadoop 应用程序使用的主要存储系统。HDFS 创建多个数据块副本并将它们分布在整个群集的计算主机上，以启用可靠且极其快速的计算功能。\n\t\n##Hive\n\tHive 是一种数据仓库系统，提供名为 HiveQL 的 SQL 类语言。\n\t\t\n##Hue\n\tHue 是与包括 Apache Hadoop 的 Cloudera Distribution 一起配合使用的图形用户界面（需要 HDFS、MapReduce 和 Hive）。\n\t\n##Impala\n\tImpala 为存储在 HDFS 和 HBase 中的数据提供了一个实时 SQL 查询接口。Impala 需要 Hive 服务，并与 Hue 共享 Hive Metastore。\n\t\n##Key-Value Store Indexer\n\t键/值 Store Indexer 侦听 HBase 中所含表内的数据变化，并使用 Solr 为其创建索引。\n\t\n##MapReduce\n\tApache Hadoop MapReduce 支持对整个群集中的大型数据集进行分布式计算（需要 HDFS）。建议改用 YARN（包括 MapReduce 2）。包括 MapReduce 用于向后兼容。\n\t\n##Oozie\n\tOozie 是群集中管理数据处理作业的工作流协调服务。\n\t\n##Solr\n\tSolr 是一个分布式服务，用于编制存储在 HDFS 中的数据的索引并搜索这些数据。\n\t\n##Spark\n\tApache Spark is an open source cluster computing system\n\t\n##Sqoop\n\tSqoop 是一个设计用于在 Apache Hadoop 和结构化数据存储（如关系数据库）之间高效地传输大批量数据的工具。Cloudera Manager 支持的版本为 Sqoop 2。\n\t\n##Sqoop 1 Client\n\tConfiguration and connector management for Sqoop 1.\n\t\n##YARN (MR2 Included)\n\tApache Hadoop MapReduce 2.0 (MRv2) 或 YARN 是支持 MapReduce 应用程序的数据计算框架（需要 HDFS）。\n\t\n##ZooKeeper\n\tApache ZooKeeper 是用于维护和同步配置数据的集中服务。\t\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fheidsoft%2Fheidsoft-bigdata","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fheidsoft%2Fheidsoft-bigdata","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fheidsoft%2Fheidsoft-bigdata/lists"}