{"id":18320188,"url":"https://github.com/cnag-biomedical-informatics/convert-pheno","last_synced_at":"2025-04-05T22:31:52.092Z","repository":{"id":44156307,"uuid":"501218398","full_name":"CNAG-Biomedical-Informatics/convert-pheno","owner":"CNAG-Biomedical-Informatics","description":"A software toolkit for the interconversion of standard data models for phenotypic data","archived":false,"fork":false,"pushed_at":"2024-05-28T11:32:39.000Z","size":83738,"stargazers_count":9,"open_issues_count":1,"forks_count":1,"subscribers_count":2,"default_branch":"main","last_synced_at":"2024-05-29T03:03:28.359Z","etag":null,"topics":["beacon","beacon-v2","bff","biomedical-informatics","cdisc","cdisc-odm","cnag","convert","convert-pheno","csv","ehr","ehr-data","ehr-phenotyping","health-data","omop","omop-cdm","phenopacket","phenopackets-v2","pxf","redcap"],"latest_commit_sha":null,"homepage":"","language":"Perl","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"artistic-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/CNAG-Biomedical-Informatics.png","metadata":{"files":{"readme":"README.md","changelog":"Changes","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":"docs/supported-formats.md","governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-06-08T11:18:29.000Z","updated_at":"2024-05-30T18:32:00.837Z","dependencies_parsed_at":"2024-05-28T14:26:15.005Z","dependency_job_id":null,"html_url":"https://github.com/CNAG-Biomedical-Informatics/convert-pheno","commit_stats":null,"previous_names":[],"tags_count":10,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/CNAG-Biomedical-Informatics%2Fconvert-pheno","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/CNAG-Biomedical-Informatics%2Fconvert-pheno/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/CNAG-Biomedical-Informatics%2Fconvert-pheno/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/CNAG-Biomedical-Informatics%2Fconvert-pheno/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/CNAG-Biomedical-Informatics","download_url":"https://codeload.github.com/CNAG-Biomedical-Informatics/convert-pheno/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247411237,"owners_count":20934650,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["beacon","beacon-v2","bff","biomedical-informatics","cdisc","cdisc-odm","cnag","convert","convert-pheno","csv","ehr","ehr-data","ehr-phenotyping","health-data","omop","omop-cdm","phenopacket","phenopackets-v2","pxf","redcap"],"created_at":"2024-11-05T18:15:28.245Z","updated_at":"2025-04-05T22:31:47.078Z","avatar_url":"https://github.com/CNAG-Biomedical-Informatics.png","language":"Perl","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"left\"\u003e\n  \u003ca href=\"https://github.com/cnag-biomedical-informatics/convert-pheno\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/cnag-biomedical-informatics/convert-pheno/main/docs/img/CP-logo.png\" width=\"220\" alt=\"Convert-Pheno\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/cnag-biomedical-informatics/convert-pheno\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/cnag-biomedical-informatics/convert-pheno/main/docs/img/CP-text.png\" width=\"500\" alt=\"Convert-Pheno\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\u003cp align=\"center\"\u003e\n    \u003cem\u003eA software toolkit for the interconversion of standard data models for phenotypic data\u003c/em\u003e\n\u003c/p\u003e\n\n[![Build and Test](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/build-and-test.yml/badge.svg)](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/build-and-test.yml)\n[![Coverage Status](https://coveralls.io/repos/github/CNAG-Biomedical-Informatics/convert-pheno/badge.svg?branch=main)](https://coveralls.io/github/CNAG-Biomedical-Informatics/convert-pheno?branch=main)\n[![CPAN Publish](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/cpan-publish.yml/badge.svg)](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/cpan-publish.yml)\n[![Kwalitee Score](https://cpants.cpanauthors.org/dist/Convert-Pheno.svg)](https://cpants.cpanauthors.org/dist/Convert-Pheno)\n![version](https://img.shields.io/badge/version-0.24_beta-orange)\n[![Docker Build](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/docker-build.yml/badge.svg)](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/docker-build.yml)\n[![Docker Pulls](https://badgen.net/docker/pulls/manuelrueda/convert-pheno?icon=docker\u0026label=pulls)](https://hub.docker.com/r/manuelrueda/convert-pheno/)\n[![Docker Image Size](https://badgen.net/docker/size/manuelrueda/convert-pheno?icon=docker\u0026label=image%20size)](https://hub.docker.com/r/manuelrueda/convert-pheno/)\n[![Documentation Status](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/documentation.yml/badge.svg)](https://github.com/cnag-biomedical-informatics/convert-pheno/actions/workflows/documentation.yml)\n[![License: Artistic-2.0](https://img.shields.io/badge/License-Artistic%202.0-0298c3.svg)](https://opensource.org/licenses/Artistic-2.0)\n[![Google Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1T6F3bLwfZyiYKD6fl1CIxs9vG068RHQ6?usp=sharing)\n\n**Documentation**: \u003ca href=\"https://cnag-biomedical-informatics.github.io/convert-pheno\" target=\"_blank\"\u003ehttps://cnag-biomedical-informatics.github.io/convert-pheno\u003c/a\u003e\n\n**Google Colab tutorial**: \u003ca href=\"https://colab.research.google.com/drive/1T6F3bLwfZyiYKD6fl1CIxs9vG068RHQ6?usp=sharing\" target=\"_blank\"\u003ehttps://colab.research.google.com/drive/1T6F3bLwfZyiYKD6fl1CIxs9vG068RHQ6?usp=sharing\u003c/a\u003e\n\n**CLI Source Code**: \u003ca href=\"https://github.com/cnag-biomedical-informatics/convert-pheno\" target=\"_blank\"\u003ehttps://github.com/cnag-biomedical-informatics/convert-pheno\u003c/a\u003e\n\n**CPAN Distribution**: \u003ca href=\"https://metacpan.org/pod/Convert::Pheno\" target=\"_blank\"\u003ehttps://metacpan.org/pod/Convert::Pheno\u003c/a\u003e\n\n**Docker Hub Image**: \u003ca href=\"https://hub.docker.com/r/manuelrueda/convert-pheno/tags\" target=\"_blank\"\u003ehttps://hub.docker.com/r/manuelrueda/convert-pheno/tags\u003c/a\u003e\n\n**Web App UI**: \u003ca href=\"https://convert-pheno.cnag.cat\" target=\"_blank\"\u003ehttps://convert-pheno.cnag.cat\u003c/a\u003e\n\n# NAME\n\nconvert-pheno - A script to interconvert common data models for phenotypic data\n\n# SYNOPSIS\n\n    convert-pheno [-i input-type] \u003cinfile\u003e [-o output-type] \u003coutfile\u003e [-options]\n\n        Arguments:                       \n          (input-type): \n                -ibff                    Beacon v2 Models ('individuals' JSON|YAML) file\n                -iomop                   OMOP-CDM CSV files or PostgreSQL dump\n                -ipxf                    Phenopacket v2 (JSON|YAML) file\n                -iredcap (experimental)  REDCap (raw data) export CSV file\n                -icdisc  (experimental)  CDISC-ODM v1 XML file\n                -icsv    (experimental)  Raw data CSV\n\n                (Wish-list)\n                #-iopenehr               openEHR\n                #-ifhir                  HL7/FHIR\n\n          (output-type):\n                -obff                    Beacon v2 Models ('individuals' JSON|YAML) file\n                -opxf                    Phenopacket v2 (JSON|YAML) file\n\n                (Wish-list)\n                #-oomop                  OMOP-CDM PostgreSQL dump\n\n                Compatible with -i(bff|pxf):\n                -ocsv                    Flatten data to CSV\n                -ojsonf                  Flatten data to 1D-JSON (or 1D-YAML if suffix is .yml|.yaml)\n                -ojsonld (experimental)  JSON-LD (interoperable w/ RDF ecosystem; YAML-LD if suffix is .ymlld|.yamlld)\n\n        Options:\n          -exposures-file \u003cfile\u003e         CSV file with a list of 'concept_id' considered to be exposures (with -iomop)\n          -mapping-file \u003cfile\u003e           Fields mapping YAML (or JSON) file\n          -max-lines-sql \u003cnumber\u003e        Maximum lines read per table from SQL dump [500]\n          -min-text-similarity-score \u003cscore\u003e Minimum score for cosine similarity (or Sorensen-Dice coefficient) [0.8] (to be used with --search mixed)\n          -ohdsi-db                      Use Athena-OHDSI database (~2.2GB) with -iomop\n          -omop-tables \u003ctables\u003e          OMOP-CDM tables to be processed. Tables \u003cCONCEPT\u003e and \u003cPERSON\u003e are always included.\n          -out-dir \u003cdirectory\u003e           Output (existing) directory\n          -O                             Overwrite output file\n          -path-to-ohdsi-db \u003cdirectory\u003e  Directory for the file \u003cohdsi.db\u003e\n          -phl|print-hidden-labels       Print original values (before DB mapping) of text fields \u003c_labels\u003e\n          -rcd|redcap-dictionary \u003cfile\u003e  REDCap data dictionary CSV file\n          -schema-file \u003cfile\u003e            Alternative JSON Schema for mapping file\n          -search \u003ctype\u003e                 Type of search [\u003eexact|mixed]\n          -svs|self-validate-schema      Perform a self-validation of the JSON schema that defines mapping (requires IO::Socket::SSL)\n          -sep|separator \u003cchar\u003e          Delimiter character for CSV files [;] e.g., --sep $'\\t'\n          -stream                        Enable incremental processing with -iomop and -obff [\u003eno-stream|stream]\n          -sql2csv                       Print SQL TABLES (only valid with -iomop). Mutually exclusive with --stream\n          -test                          Does not print time-changing-events (useful for file-based cmp)\n          -text-similarity-method \u003cmethod\u003e The method used to compare values to DB [\u003ecosine|dice]\n          -u|username \u003cusername\u003e         Set the username\n\n        Generic Options:\n          -debug \u003clevel\u003e                 Print debugging level (from 1 to 5, being 5 max)\n          -help                          Brief help message\n          -log                           Save log file (JSON). If no argument is given then the log is named [convert-pheno-log.json]\n          -man                           Full documentation\n          -no-color                      Don't print colors to STDOUT [\u003ecolor|no-color]\n          -v|verbose                     Verbosity on\n          -V|version                     Print Version\n\n# DESCRIPTION\n\n`convert-pheno` is a command-line front-end to the CPAN's module [Convert::Pheno](https://metacpan.org/pod/Convert%3A%3APheno).\n\n# SUMMARY\n\nA script that uses [Convert::Pheno](https://metacpan.org/pod/Convert%3A%3APheno) to interconvert common data models for phenotypic data\n\n# INSTALLATION\n\nIf you plan to only use the CLI, we recommend installing it via CPAN. See details below.\n\n## Non containerized\n\nThe script runs on command-line Linux and it has been tested on Debian/RedHat/MacOS based distributions (only showing commands for Debian's). Perl 5 is installed by default on Linux, \nbut we will install a few CPAN modules with `cpanminus`.\n\n### Method 1: From CPAN\n\nFirst install system level dependencies:\n\n    sudo apt-get install cpanminus libbz2-dev zlib1g-dev libperl-dev libssl-dev\n\nNow you have two choose between one of the 3 options below:\n\n**Option 1:** System-level installation:\n\n    cpanm --notest --sudo Convert::Pheno\n    convert-pheno -h\n\n**Option 2:** Install Convert-Pheno and the dependencies at `~/perl5`\n\n    cpanm --local-lib=~/perl5 local::lib \u0026\u0026 eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib)\n    cpanm --notest Convert::Pheno\n    convert-pheno --help\n\nTo ensure Perl recognizes your local modules every time you start a new terminal, you should type:\n\n    echo 'eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib)' \u003e\u003e ~/.bashrc\n\n**Option 3:** Install Convert-Pheno and the dependencies in a \"virtual environment\" (at `local/`) . We'll be using the module `Carton` for that:\n\n    mkdir local\n    cpanm --notest --local-lib=local/ Carton\n    echo \"requires 'Convert::Pheno';\" \u003e cpanfile\n    export PATH=$PATH:local/bin; export PERL5LIB=$(pwd)/local/lib/perl5:$PERL5LIB\n    carton install\n    carton exec -- convert-pheno -help\n\n### Method 2: From CPAN in a Conda environment\n\nPlease follow [these instructions](https://cnag-biomedical-informatics.github.io/convert-pheno/download-and-installation/#__tabbed_1_2).\n\n### Method 3: From Github\n\n    git clone https://github.com/cnag-biomedical-informatics/convert-pheno.git\n    cd convert-pheno\n\nInstall system level dependencies:\n\n    sudo apt-get install cpanminus libbz2-dev zlib1g-dev libperl-dev libssl-dev\n\nNow you have two choose between one of the 3 options below:\n\n**Option 1:** Install dependencies (they're harmless to your system) as `sudo`:\n\n    cpanm --notest --sudo --installdeps .\n    bin/convert-pheno --help            \n\n**Option 2:** Install the dependencies at `~/perl5`:\n\n    cpanm --local-lib=~/perl5 local::lib \u0026\u0026 eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib)\n    cpanm --notest --installdeps .\n    bin/convert-pheno --help\n\nTo ensure Perl recognizes your local modules every time you start a new terminal, you should type:\n\n    echo 'eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib)' \u003e\u003e ~/.bashrc\n\n**Option 3:** Install the dependencies in a \"virtual environment\" (at `local/`) . We'll be using the module `Carton` for that:\n\n    mkdir local\n    cpanm --notest --local-lib=local/ Carton\n    export PATH=$PATH:local/bin; export PERL5LIB=$(pwd)/local/lib/perl5:$PERL5LIB\n    carton install\n    carton exec -- bin/convert-pheno -help\n\n## Containerized\n\n### Method 4: From Docker Hub\n\nDownload a docker image (latest version - amd64|x86-64) from [Docker Hub](https://hub.docker.com/r/manuelrueda/convert-pheno) by executing:\n\n    docker pull manuelrueda/convert-pheno:latest\n    docker image tag manuelrueda/convert-pheno:latest cnag/convert-pheno:latest\n\nSee additional instructions below.\n\n### Method 5: With Dockerfile\n\nPlease download the `Dockerfile` from the repo:\n\n    wget https://raw.githubusercontent.com/cnag-biomedical-informatics/convert-pheno/main/Dockerfile\n\nAnd then run:\n\n    docker buildx build -t cnag/convert-pheno:latest .\n\n### Additional instructions for Methods 4 and 5\n\nTo run the container (detached) execute:\n\n    docker run -tid -e USERNAME=root --name convert-pheno cnag/convert-pheno:latest\n\nTo enter:\n\n    docker exec -ti convert-pheno bash\n\nThe command-line executable can be found at:\n\n    /usr/share/convert-pheno/bin/convert-pheno\n\nThe default container user is `root` but you can also run the container as `$UID=1000` (`dockeruser`). \n\n     docker run --user 1000 -tid --name convert-pheno cnag/convert-pheno:latest\n    \n\nAlternatively, you can use `make` to perform all the previous steps:\n\n    wget https://raw.githubusercontent.com/cnag-biomedical-informatics/convert-pheno/main/Dockerfile\n    wget https://raw.githubusercontent.com/cnag-biomedical-informatics/convert-pheno/main/makefile.docker\n    make -f makefile.docker install\n    make -f makefile.docker run\n    make -f makefile.docker enter\n\n### Mounting volumes\n\nDocker containers are fully isolated. If you need the mount a volume to the container please use the following syntax (`-v host:container`). \nFind an example below (note that you need to change the paths to match yours):\n\n    docker run -tid --volume /media/mrueda/4TBT/data:/data --name convert-pheno-mount cnag/convert-pheno:latest\n\nThen I will do something like this:\n\n    # First I create an alias to simplify invocation (from the host)\n    alias convert-pheno='docker exec -ti convert-pheno-mount /usr/share/convert-pheno/bin/convert-pheno'\n\n    # Now I use the alias to run the command (note that I use the flag --out-dir to specify the output directory)\n    convert-pheno -ibff /data/individuals.json -opxf pxf.json --out-dir /data\n\n### System requirements\n\n    * Ideally a Debian-based distribution (Ubuntu or Mint), but any other (e.g., CentOs, OpenSuse, MacOS) should do as well.\n      (It should also work on macOS and Windows Server, but we are only providing information for Linux here)\n    * Perl 5 (\u003e= 5.26 core; installed by default in most Linux distributions). Check the version with \"perl -v\".\n    * \u003e= 4GB of RAM\n    * 1 core\n    * At least 16GB HDD\n\n# HOW TO RUN CONVERT-PHENO\n\nFor executing convert-pheno you will need:\n\n- Input file(s):\n\n    A text file in one of the accepted formats. With `--iomop` I/O files can be gzipped.\n\n- Optional: \n\n    Athena-OHDSI database\n\n    The database file is available at this [link](https://drive.google.com/drive/folders/1-5Ywf-hhwb8bX1sRNV2Tf3EjH4TCaC8P?usp=sharing) (~2.2GB). The database may be needed when using `-iomop`.\n\n    Regardless if you're using the containerized or non-containerized version, the download procedure is the same. For CLI users, Google makes it difficult to use `wget`, `curl` or `aria2c` so we will use a `Python` module instead:\n\n        $ pip install gdown\n\n    And then run the following script\n\n        import gdown\n\n        url = 'https://drive.google.com/uc?export=download\u0026id=1-Ls1nmgxp-iW-8LkRIuNNdNytXa8kgNw'\n        output = './ohdsi.db'\n        gdown.download(url, output, quiet=False)\n\n    Once downloaded, you have two options:\n\n    a) Move the file `ohdsi.db` inside the `share/db/` directory.\n\n    or\n\n    b) Use the option `--path-to-ohdsi-db`\n\n**Examples:**\n\n    $ bin/convert-pheno -ipxf phenopackets.json -obff individuals.json\n\n    $ $path/convert-pheno -ibff individuals.json -opxf phenopackets.yaml --out-dir my_out_dir \n\n    $ $path/convert-pheno -iredcap redcap.csv -opxf phenopackets.json --redcap-dictionary redcap_dict.csv --mapping-file mapping_file.yaml\n\n    $ $path/convert-pheno -iomop dump.sql -obff individuals.json\n\n    $ $path/convert-pheno -iomop dump.sql.gz -obff individuals.json.gz --stream -omop-tables measurement -verbose\n\n    $ $path/convert-pheno -cdisc cdisc_odm.xml -obff individuals.json --rcd redcap_dict.csv --mapping-file mapping_file.yaml --search mixed --min-text-similarity-score 0.6\n\n    $ $path/convert-pheno -iomop *csv -obff individuals.json -sep ','\n\n    $ carton exec -- $path/convert-pheno -ibff individuals.json -opxf phenopackets.json # If using Carton\n\n## COMMON ERRORS AND SOLUTIONS\n\n    * Error message: CSV_XS ERROR: 2023 - EIQ - QUO character not allowed @ rec 1 pos 21 field 1\n      Solution: Make sure you use the right character separator for your data with --sep \u003cchar\u003e. \n                The script tries to guess it from the file extension, but sometimes extension and actual separator do not match. \n                When using REDCap as input, make sure that \u003c--iredcap\u003e and \u003c--rcd\u003e files use the same separator field.\n                The defauly value for the separator is ';'. \n      Example for tab separator in CLI.\n       --sep  $'\\t' \n\n    * Error message: Foo\n      Solution: Bar\n\n# CITATION\n\nThe author requests that any published work that utilizes `Convert-Pheno` includes a cite to the the following reference:\n\nRueda, M et al., (2024). Convert-Pheno: A software toolkit for the interconversion of standard data models for phenotypic data. Journal of Biomedical Informatics. [DOI](https://doi.org/10.1016/j.jbi.2023.104558)\n\n# AUTHOR \n\nWritten by Manuel Rueda, PhD. Info about CNAG can be found at [https://www.cnag.eu](https://www.cnag.eu).\n\n# COPYRIGHT AND LICENSE\n\nCopyright (C) 2022-2024, Manuel Rueda - CNAG.\n\nThis program is free software, you can redistribute it and/or modify it under the terms of the [Artistic License version 2.0](https://metacpan.org/pod/perlartistic).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcnag-biomedical-informatics%2Fconvert-pheno","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcnag-biomedical-informatics%2Fconvert-pheno","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcnag-biomedical-informatics%2Fconvert-pheno/lists"}