{"id":19152491,"url":"https://github.com/gesiscss/promoss","last_synced_at":"2026-06-15T18:32:36.520Z","repository":{"id":72057979,"uuid":"99912000","full_name":"gesiscss/promoss","owner":"gesiscss","description":"PROMOSS official","archived":false,"fork":false,"pushed_at":"2017-08-24T12:15:46.000Z","size":19990,"stargazers_count":1,"open_issues_count":0,"forks_count":2,"subscribers_count":14,"default_branch":"master","last_synced_at":"2025-02-22T21:15:02.390Z","etag":null,"topics":["clustering","lda","topic-modeling"],"latest_commit_sha":null,"homepage":null,"language":"Java","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/gesiscss.png","metadata":{"files":{"readme":"readme.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2017-08-10T10:36:43.000Z","updated_at":"2021-03-10T07:34:07.000Z","dependencies_parsed_at":"2023-07-23T00:45:44.367Z","dependency_job_id":null,"html_url":"https://github.com/gesiscss/promoss","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/gesiscss/promoss","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/gesiscss%2Fpromoss","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/gesiscss%2Fpromoss/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/gesiscss%2Fpromoss/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/gesiscss%2Fpromoss/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/gesiscss","download_url":"https://codeload.github.com/gesiscss/promoss/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/gesiscss%2Fpromoss/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34376122,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-15T02:00:07.085Z","response_time":63,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["clustering","lda","topic-modeling"],"created_at":"2024-11-09T08:18:04.408Z","updated_at":"2026-06-15T18:32:36.500Z","avatar_url":"https://github.com/gesiscss.png","language":"Java","funding_links":[],"categories":[],"sub_categories":[],"readme":"\n# Promoss Topic Modelling Toolbox\n(C) Copyright 2016, Christoph Carl Kling\n\nPromoss makes use of multiple free software packages -- thanks to the authors of:\n\nKnoceans by Gregor Heinrich Gregor Heinrich (gregor :: arbylon : net) published under GNU GPL.\n\nTartarus Snowball stemmer by Martin Porter and Richard Boulton published under BSD License (see http://www.opensource.org/licenses/bsd-license.html ), with Copyright (c) 2001, Dr Martin Porter, and (for the Java developments) Copyright (c) 2002, Richard Boulton. \n\nQuickhull3D Copyright by John E. Lloyd, 2004. \n\nApache Xerces Java and NekoHTML are released under Apache License 2.0.\n\nMcCallum, Andrew Kachites.  \"MALLET: A Machine Learning for Language Toolkit.\" http://mallet.cs.umass.edu. 2002, released under the Common Public License. http://www.opensource.org/licenses/cpl1.0.php\n\nPromoss is free software; you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation; either version 3 of the License, or (at your option) any later version.\n\nPromoss is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.\n\nYou should have received a copy of the GNU General Public License along with this program; if not, write to the Free Software Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307 USA\n\n## Support\nPlease contact me if you need help running the code: promoss (ät) c-kling.de\n\n---\n\n## First steps\n\n### Building the jar file\nYou can build the promoss.jar using Ant. Go to the directory of the extracted promoss.tar.gz file (in which the build.xml is located) and enter the command:\n```\nant; ant build-jar\n```\n(The ant build might yield errors for classes under development which can be ignored.)\n\n### Demo files\nIf you would like to have demo files to play around with, just write a mail to promoss@c-kling.de\n\n---\n\n## Latent Dirichlet Allocation (LDA)\nCollapsed stochastic variational inference for LDA with an asymmetric document-topic prior.\n\n### Example command line usage\n```\njava -Xmx11000M -jar promoss.jar -directory demo/ml_demo/ -method \"LDA\" -MIN_DICT_WORDS 0 -T 5\n```\n\n### Input files\nThe most simple way to feed your documents into the topic model is via the corpus.txt file, which can include raw documents (each line corresponds to a document). From this corpus.txt, a wordsets file with the processed documents in SVMlight format is created, called wordsets. You can also directly give the wordsets file and a words.txt dictionary, where the line number (starting with 0) corresponds to the word ID in the SVMlight file.\n\n#### corpus.txt\nEach line corresponds to a document. Words of documents are separated by spaces. (However, one can also input raw text and set the -processed parameter to false in order to use a library-specific code for splitting words.)\n\nExample corpus.txt:\n```\nexist distribut origin softwar distributor agre gpl\ngpl establish term distribut origin softwar even goe unmodifi word distribut gpl softwar one agre \ndynam link constitut make deriv work allow dynam link long rule follow code make deriv work rule\ngpl also deal deriv work link creat deriv work gpl affect gpl defin scope copyright law gpl section \n```\n\n#### words.txt \nThis optional file gives the vocabulary, one word per row. The line numbers correspond to the later indices in the topic-word matrix.\n\n### Output files\nAfter each 10 runs, important parameters are stored in the *output_LDA/* subfolder, with the number of runs as folder name. The *topktopic_words* file contains the top words of each topic (the number of returned top words can be set via the -topk parameter). The *nkt* file contains the word counts for each topic: Each line corresponds to a topic, and each column to a word (starting with index 0), corrsponding to the line numbers in *words.txt* file located in the main directory. The *doc_topic* file contains the topic probabilities for each document, rows correspond to documents (same ordering as given), columns to topics.\n\n### Mandatory parameter\n* directory \t\tString. Gives the directory of the texts.txt file.\n\n### Optional parameters:\n* T\t\t\tInteger. Number of topics. Default: 100\n* RUNS\t\t\tInteger. Number of iterations the sampler will run. Default: 200\n* SAVE_STEP\t\tInteger. Number of iterations after which the learned paramters are saved. Default: 10\n* TRAINING_SHARE\t\tDouble. Gives the share of documents which are used for training (0 to 1). Default: 1\n* BATCHSIZE\t\tInteger. Batch size for topic estimation. Default: 128\n* BURNIN\t\t\tInteger. Number of iterations till the topics are updated. Default: 0\n* INIT_RAND\t\tDouble. Topic-word counts are initiatlised as INIT_RAND * RANDOM(). Default: 0\n* MIN_DICT_WORDS\t\tInteger. If the words.txt file is missing, words.txt is created by using words which occur at least MIN_DICT_WORDS times in the corpus. Default: 100\n* save_prefix\t\tString. If given, this String is appended to all output files.\n* alpha\t\t\tDouble. Initial value of alpha_0. Default: 1\n* rhokappa\t\tDouble. Initial value of kappa, a parameter for the learning rate of topics. Default: 0.5\n* rhotau\t\t\tInteger. Initial value of tau, a parameter for the learning rate of topics. Default: 64\n* rhos\t\t\tInteger. Initial value of s, a parameter for the learning rate of topics. Default: 1\n* rhokappa_document\tDouble. Initial value of kappa, a parameter for the learning rate of the document-topic distribution. Default: kappa\n* rhotau_document\tInteger. Initial value of tau, a parameter for the learning rate of the document-topic distribution. Default: tau\n* rhos_document\t\tInteger. Initial value of tau, a parameter for the learning rate of the document-topic distribution. Default: rhos\n* processed\t\tBoolean. Tells if the text is already processed, or if words should be split with complex regular expressions. Otherwise split by spaces. Default: true.\n* stemming\t\tBoolean. Activates word stemming in case no words.txt/wordsets file is given. Default: false\n* stopwords\t\tBoolean. Activates stopword removal in case no words.txt/wordsets file is given. Default: false\n* language\t\tString. Currently \"en\" and \"de\" are available languages for stemming. Default: \"en\"\n* store_empty\t\tBoolean. Determines if empty documents should be omitted in the final document-topic matrix or if the topic distribution should be predicted using the context. Default: True\n* topk\t\t\tInteger. Set the number of top words returned in the topktopics file of the output. Default: 100\n\n---\n\n## Hierarchical Multi-Dirichlet Process Topic Model (Promoss)\nAn efficient topic model which uses arbitrary document metadata!\n\nFor a description of the model, I refer to Chapter 4 of my dissertation: \n\n[Christoph Carl Kling. Probabilistic Models for Context in Social Media - Novel Approaches and Inference Schemes. 2016](https://kola.opus.hbz-nrw.de/frontdoor/deliver/index/docId/1397/file/DissertationChristophKling.pdf)\n\n### Example command line usage\n```\njava -Xmx11000M -jar promoss.jar -directory demo/ml_demo/ -meta_params \"T(L1000,W1000,D10,Y100,M20);N\" -MIN_DICT_WORDS 1000\n```\n\nThis will sample topics from a demo dataset of 1000 messages of the linux kernel mailing list. Messages are already stemmed and stopwords were removed. There are 1000 clusters for the first four contexts (which are the timeline and the yearly, weekly and daily cycle). Many clusters are empty, because the original dataset contained \u003e3m documents. This is just for testing if the algorithm runs, a demo dataset  with nicer results is in preparation.\n\n### Input file format\nThere is two standard input formats for the data.\nThe first is based on raw, unclustered metadata stored in meta.txt and the corpus, stored in corpus.txt\nThe second is based on already clustered data (texts.txt) with given document groups defined in groups.txt (groups of documents share the same parent clusters in the Promoss).\nWhen running the code, the script first looks for a texts.txt and a groups.txt. If any of those documents is missing, the script looks for the corpus.txt and meta.txt, from which it generates the texts.txt and corpus.txt.\nFinally, a groups file with the groups of the documents and a wordsets file with the processed documents in SVMlight format are created.\n\n#### Variant 1\n##### corpus.txt \nEach line corresponds to a document. Words of documents are separated by spaces. (However, one can also input raw text and set the -processed parameter to false in order to use a library-specific code for splitting words.)\n\nExample corpus.txt:\n```\n  exist distribut origin softwar distributor agre gpl\n  gpl establish term distribut origin softwar even goe unmodifi word distribut gpl softwar one agre \n  dynam link constitut make deriv work allow dynam link long rule follow code make deriv work rule\n  gpl also deal deriv work link creat deriv work gpl affect gpl defin scope copyright law gpl section \n```\n\n\n##### meta.txt \nHere we give the metadata values separated by semicolons. Possible metadata are geographical coordinates (latitude and longitude separated by comma), UNIX timestamps (in seconds), nominal values (e.g. category names, numbers) or oordinal variables (stored numbers which correspond to the ordering). The metadata types have to be specified via the -meta_params parameter (see below for a description).\n\nExample meta.txt:\n```\n33.150051,-114.365448;1139316299;1\n34.150051,-118.365448;1139316058;2\n43.59772,-116.235705;1139261931;3\n14.559243,120.982732;1139256458;2\n```\n\n#### Variant 2\n##### texts.txt \nEach line corresponds to a document. First, the context group IDs (for each context one) are given, separated by commas. The context group in context 0 is given first, then the context group in context 1 and so on. Then follows a space and the words of the documents separated by spaces. \n\nExample texts.txt:\n```\n254,531,790,157,0  exist distribut origin softwar distributor agre gpl\n254,528,789,157,0  gpl establish term distribut origin softwar even goe unmodifi word distribut gpl softwar one agre \n254,901,700,157,0  dynam link constitut make deriv work allow dynam link long rule follow code make deriv work rule\n254,838,691,157,0  gpl also deal deriv work link creat deriv work gpl affect gpl defin scope copyright law gpl section \n```\n\n##### groups.txt \nEach line gives the parent context clusters of a context group. Data are separated by spaces. The first column gives the context id, the second column gives the group ID of the context group, and then the IDs of the context clusters from which the documents of that context group draw their topics are given.\n\nExample groups.txt\n```\n0 0 0 1\n0 1 0 1 2\n0 2 1 2 3\n0 3 2 3 4\n0 4 3 4 5\n0 5 4 5 6\n0 6 5 6 7\n0 7 6 7 8\n0 8 7 8 9\n0 9 8 9 10\n0 10 9 10 11\n[...]\n0 254 123 23 53\n```\n\nThe first line reads: For context 0, documents which are assigned to context group 0 draw their topics from context cluster 0 and context cluster 1.\nThe last line reads: For context 0, documents which are assigned to context group 254 draw their topics from context cluster 123, 23 and 53.\nIf no groups.txt is given, all context groups will be linked to a context cluster with the same ID, which means that all context clusters are independent.\n\n#### words.txt \nThis optional file gives the vocabulary, one word per row. The line numbers correspond to the later indices in the topic-word matrix.\n\n### Output files\nCluster descriptions (e.g. means of the geographical clusters, bins of timestamps etc.) are saved in the *cluster_desc/* folder.\nAfter each 10 runs, important parameters are stored in the *output_HMDP/* subfolder, with the number of runs as folder name. The *clusters_X* file contains the topic loadings of each cluster of the *X*th metadata. The *topktopic_words* file contains the top words of each topic (the number of returned top words can be set via the -topk parameter).\nThe *nkt* file contains the word counts for each topic: Each line corresponds to a topic, and each column to a word (starting with index 0), corrsponding to the line numbers in *words.txt* file located in the main directory. The *doc_topic* file contains the topic probabilities for each document, rows correspond to documents (same ordering as given), columns to topics.\n\n### Mandatory parameter\n* directory \t\tString. Gives the directory of the texts.txt and groups.txt file.\n\n### Mandatory Parameters when Using corpus.txt and meta.txt (Input Variant 1)\n* meta_params\t\tString. Specifies the metadata types and gives the desired clustering. Types of metadata are given separated by semicolons (and correspond to the number of different metadata in the meta.txt file. Possible datatypes are:\n * G\tGeographical coordinates. The number of desired clusters is specified in brackets, i.e. G(1000) will cluster the documents into 1000 clusters based on the geographical coordinates. (Technical detail: we use EM to fit a mixture of fisher distributions.)\n * T\tUNIX timestamps (in seconds). The number of clusters (based on binning) is given in brackets, and there can be multiple clusterings based on a binning on the timeline or temporal cycles. This is indicated by a letter followed by the number of desired clusters:\n * L\tBinning based on the timeline. Example: L1000 gives 1000 bins.\n * Y\tBinning based on the yearly cycle. Example: L1000 gives 1000 bins.\n * M\tBinning based on the monthly cycle. Example: L1000 gives 1000 bins.\n * W\tBinning based on the weekly cycle. Example: L1000 gives 1000 bins.\n * D\tBinning based on the daily  cycle. Example: L1000 gives 1000 bins.\n* O\tOrdinal values (numbers)\n* N\tNominal values (text strings)\n\t\t\t\n\nExample usage in the -meta_params parameter: \n```\n-meta_params \"G(1000);T(L1000,Y100,M10,W20,D10);O\"\n```\n\nThis command can be used for the meta.txt given above. It would create 1000 geographical clusters based on the latitude and longitude. Then it would parse each UNIX timestamp to create 1000 clusters on the timeline, 100 clusters on the yearly, 10 clusters on the monthly, 20 clusters on the weekly and 10 clusters on the daily cycle (based on simple binning). Then the third metadata variable would be interpreted as an ordinal variable, meaning that each different value is an own cluster which is smoothed with the previous and next cluster (if existent).\n\n#### Rule of thumb for clustering\nClusters should not be too small, because the observed documents in a cluster should be sufficient to learn a cluster-specific topic prior.\nOn the other hand, too few clusters prevent the model from capturing differences in topic frequencies in the context space.\nOne rule of thumb for the number of clusters C in a corpus with M documents and (an expected number of) T topics is: C = M/T. I.e. if we have 1.000.000 documents and expect about 100 topics, it is reasonable to pick 10.000 clusters. This approximation is very simplistic, I recommend to use e.g. Dirichlet process-based methods such as infinite Gaussian mixture models for cluster detection before running the model. \n\n### Optional parameters\nThe parameters are sorted, most common parameters are on top:\n* T\t\t\tInteger. Number of truncated topics. Default: 100\n* RUNS\t\t\tInteger. Number of iterations the sampler will run. Default: 200\n* processed\t\tBoolean. Tells if the text is already processed, or if words should be split with complex regular expressions. Otherwise split by spaces. Default: true.\n* stemming\t\tBoolean. Activates word stemming in case no words.txt/wordsets file is given. Default: false\n* stopwords\t\tBoolean. Activates stopword removal in case no words.txt/wordsets file is given. Default: false\n* language\t\tString. Currently \"en\" and \"de\" are available languages for stemming. Default: \"en\"\n* store_empty\t\tBoolean. Determines if empty documents should be omitted in the final document-topic matrix or if the topic distribution should be predicted using the context. Default: True\n* TRAINING_SHARE\t\tDouble. Gives the share of documents which are used for training (0 to 1). Default: 1\n* topk\t\t\tInteger. Set the number of top words returned in the topktopics file of the output. Default: 100\n* gamma\t\t\tDouble. Initial scaling parameter of the top-level Dirichlet process. Default: 1\n* learn_gamma\t\tBoolean. Should gamma be learned during inference? Default: True\n* SAVE_STEP\t\tInteger. Number of iterations after which the learned paramters are saved. Default: 10\n* BATCHSIZE\t\tInteger. Batch size for topic estimation. Default: 128\n* BATCHSIZE_GROUPS\tInteger. Batch size for group-specific parameter estimation. Default: BATCHSIZE\n* BURNIN\t\t\tInteger. Number of iterations till the topics are updated. Default: 0\n* BURNIN_DOCUMENTS\tInteger. Gives the number of sampling iterations where the group-specific parameters are not updated yet. Default: 0\n* INIT_RAND\t\tDouble. Topic-word counts are initiatlised as INIT_RAND * RANDOM(). Default: 0\n* SAMPLE_ALPHA\t\tInteger. Every SAMPLE_ALPHAth document is used to estimate alpha_1. Default: 1\n* BATCHSIZE_ALPHA\tInteger. How many observations do we take before updating alpha_1. Default: 1000\n* MIN_DICT_WORDS\t\tInteger. If the words.txt file is missing, words.txt is created by using words which occur at least MIN_DICT_WORDS times in the corpus. Default: 100\n* save_prefix\t\tString. If given, this String is appended to all output files.\n* alpha_0\t\tDouble. Initial value of alpha_0. Default: 1\n* alpha_1\t\tDouble. Initial value of alpha_1. Default: 1\n* epsilon\t\tComma-separated double. Dirichlet prior over the weights of contexts. Comma-separated double values, with dimensionality equal to the number of contexts.\n* delta_fix \t\tIf set, delta is fixed and set to this value. Otherwise delta is learned during inference.\n* rhokappa\t\tDouble. Initial value of kappa, a parameter for the learning rate of topics. Default: 0.5\n* rhotau\t\t\tInteger. Initial value of tau, a parameter for the learning rate of topics. Default: 64\n* rhos\t\t\tInteger. Initial value of s, a parameter for the learning rate of topics. Default: 1\n* rhokappa_document\tDouble. Initial value of kappa, a parameter for the learning rate of the document-topic distribution. Default: kappa\n* rhotau_document\tInteger. Initial value of tau, a parameter for the learning rate of the document-topic distribution. Default: tau\n* rhos_document\t\tInteger. Initial value of tau, a parameter for the learning rate of the document-topic distribution. Default: rhos\n* rhokappa_group\t\tDouble. Initial value of kappa, a parameter for the learning rate of the group-topic distribution. Default: kappa\n* rhotau_group\t\tInteger. Initial value of tau, a parameter for the learning rate of the group-topic distribution. Default: tau\n* rhos_group\t\tInteger. Initial value of tau, a parameter for the learning rate of the group-topic distribution. Default: rhos\n\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgesiscss%2Fpromoss","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fgesiscss%2Fpromoss","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgesiscss%2Fpromoss/lists"}