{"id":48413668,"url":"https://github.com/pangaea-data-publisher/fuji","last_synced_at":"2026-04-06T06:35:11.542Z","repository":{"id":43694787,"uuid":"259678183","full_name":"pangaea-data-publisher/fuji","owner":"pangaea-data-publisher","description":"FAIRsFAIR Research Data Object Assessment Service","archived":false,"fork":false,"pushed_at":"2026-01-19T12:09:51.000Z","size":10289,"stargazers_count":64,"open_issues_count":27,"forks_count":46,"subscribers_count":8,"default_branch":"master","last_synced_at":"2026-01-29T16:23:10.424Z","etag":null,"topics":["fairdata","hacktoberfest"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/pangaea-data-publisher.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":"AUTHORS","dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2020-04-28T15:33:40.000Z","updated_at":"2026-01-19T12:09:56.000Z","dependencies_parsed_at":"2024-02-28T09:36:08.902Z","dependency_job_id":"dd480435-7144-49f2-a181-019beeea34ef","html_url":"https://github.com/pangaea-data-publisher/fuji","commit_stats":null,"previous_names":[],"tags_count":16,"template":false,"template_full_name":null,"purl":"pkg:github/pangaea-data-publisher/fuji","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pangaea-data-publisher%2Ffuji","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pangaea-data-publisher%2Ffuji/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pangaea-data-publisher%2Ffuji/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pangaea-data-publisher%2Ffuji/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/pangaea-data-publisher","download_url":"https://codeload.github.com/pangaea-data-publisher/fuji/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pangaea-data-publisher%2Ffuji/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31463015,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-05T21:22:52.476Z","status":"online","status_checked_at":"2026-04-06T02:00:07.287Z","response_time":112,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["fairdata","hacktoberfest"],"created_at":"2026-04-06T06:35:09.149Z","updated_at":"2026-04-06T06:35:11.530Z","avatar_url":"https://github.com/pangaea-data-publisher.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# F-UJI (FAIRsFAIR Research Data Object Assessment Service)\nDevelopers: [Robert Huber](mailto:rhuber@marum.de), [Anusuriya Devaraju](mailto:anusuriya.devaraju@googlemail.com)\n\nThanks to [Heinz-Alexander Fuetterer](https://github.com/afuetterer) for his contributions and his help in cleaning up the code.\n\n| __CI__ | [![CI](https://github.com/pangaea-data-publisher/fuji/actions/workflows/ci.yml/badge.svg)](https://github.com/pangaea-data-publisher/fuji/actions/workflows/ci.yml) [![Coverage](https://coveralls.io/repos/github/pangaea-data-publisher/fuji/badge.svg?branch=master)](https://coveralls.io/github/pangaea-data-publisher/fuji?branch=master) |\n| :--- | :--- |\n| __CD__ | [![Publish Docker image](https://github.com/pangaea-data-publisher/fuji/actions/workflows/publish-docker.yml/badge.svg)](https://github.com/pangaea-data-publisher/fuji/actions/workflows/publish-docker.yml) |\n| __Package__ | [![Python Version](https://img.shields.io/badge/python-3.11-blue)](https://www.python.org/) |\n| __Meta__    | [![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.11084909.svg)](https://doi.org/10.5281/zenodo.11084909) [![GitHub License](https://img.shields.io/github/license/pangaea-data-publisher/fuji.svg)](https://github.com/pangaea-data-publisher/fuji/blob/master/LICENSE) [![pre-commit](https://img.shields.io/badge/pre--commit-enabled-brightgreen?logo=pre-commit\u0026logoColor=white)](https://github.com/pre-commit/pre-commit) [![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff) |\n\n## Overview\n\nF-UJI is a web service to programmatically assess FAIRness of research data objects based on [metrics](https://doi.org/10.5281/zenodo.3775793) developed by the [FAIRsFAIR](https://www.fairsfair.eu/) project.\nThe service will be applied to demonstrate the evaluation of objects in repositories selected for in-depth collaboration with the project.\n\nThe '__F__' stands for FAIR (of course) and '__UJI__' means 'Test' in Malay. So __F-UJI__ is a FAIR testing tool.\n\n**Cite as**\n\nDevaraju, A. and Huber, R. (2021). An automated solution for measuring the progress toward FAIR research data. Patterns, vol 2(11), https://doi.org/10.1016/j.patter.2021.100370\n\n### Clients and User Interface\n\nA web demo using F-UJI is available at \u003chttps://www.f-uji.net\u003e.\n\nAn R client package that was generated from the F-UJI OpenAPI definition is available from \u003chttps://github.com/NFDI4Chem/rfuji\u003e.\n\nAn open source web client for F-UJI is available at \u003chttps://github.com/MaastrichtU-IDS/fairificator\u003e.\n\n### Output Formats\n\nAs as REST service the default output of F-UJI is JSON. Since version 3.5.0 F-UJI additionally offers output of results in various RDF formats following the DQV specifications. Content negotiation can be used to retrieve a RDF DQV result output.\nSupported RDF serialisations are: application/rdf+xml, application/x-turtle, text/turtle, application/ld+json, text/n3\n\n## Assessment Scope, Constraint and Limitation\nThe service is **in development** and its assessment depends on several factors.\n- In the FAIR ecosystem, FAIR assessment must go beyond the object itself. FAIR enabling services and repositories are vital to ensure that research data objects remain FAIR over time. Importantly, machine-readable services (e.g., registries) and documents (e.g., policies) are required to enable automated tests.\n- In addition to repository and services requirements, automated testing depends on clear machine assessable criteria. Some aspects (rich, plurality, accurate, relevant) specified in FAIR principles still require human mediation and interpretation.\n- The tests must focus on generally applicable data/metadata characteristics until domain/community-driven criteria have been agreed (e.g., appropriate schemas and required elements for usage/access control, etc.). For example, for some metrics (i.e., on I and R principles), the automated tests we proposed only inspect the ‘surface’ of criteria to be evaluated. Therefore, tests are designed in consideration of generic cross-domain metadata standards such as Dublin Core, DCAT, DataCite, schema.org, etc.\n- FAIR assessment is performed based on aggregated metadata; this includes metadata embedded in the data (landing) page, metadata retrieved from a PID provider (e.g., DataCite content negotiation) and other services (e.g., re3data).\n\n![alt text](https://github.com/pangaea-data-publisher/fuji/blob/master/fuji_server/static/main.png?raw=true)\n\n## Requirements\n[Python](https://www.python.org/downloads/) `3.11`\n\n### Google Dataset Search\nSince FsF metric 0.8, The Google Corpus is no longer used, **the following steps are therefore not required anymore**.\n* Download the latest Dataset Search corpus file from: \u003chttps://www.kaggle.com/googleai/dataset-search-metadata-for-datasets\u003e\n* Open file `fuji_server/helper/create_google_cache_db.py` and set variable 'google_file_location' according to the file location of the corpus file\n* Run `create_google_cache_db.py` which creates a SQLite database in the data directory. From root directory run `python3 -m fuji_server.helper.create_google_cache_db`.\n\nThe service was generated by the [swagger-codegen](https://github.com/swagger-api/swagger-codegen) project. By using the\n[OpenAPI-Spec](https://github.com/swagger-api/swagger-core/wiki) from a remote server, you can easily generate a server stub.\nThe service uses the [Connexion](https://github.com/spec-first/connexion) library on top of Flask.\n\n## Usage\nBefore running the service, please set user details in the configuration file, see config/users.py.\n\nTo install F-UJI, you may execute the following Python-based or docker-based installation commands from the root directory:\n\n### Python module-based installation\n\nFrom the fuji source folder run:\n```bash\npython -m pip install .\n```\nThe F-UJI server can now be started with:\n```bash\npython -m fuji_server -c fuji_server/config/server.ini\n```\n\nThe OpenAPI user interface is then available at \u003chttp://localhost:1071/fuji/api/v1/ui/\u003e.\n\n### Docker-based installation\n\n```bash\ndocker run -d -p 1071:1071 ghcr.io/pangaea-data-publisher/fuji\n```\n\nTo access the OpenAPI user interface, open the URL below in the browser:\n\u003chttp://localhost:1071/fuji/api/v1/ui/\u003e\n\nYour OpenAPI definition lives here:\n\n\u003chttp://localhost:1071/fuji/api/v1/openapi.json\u003e\n\nYou can provide a different server config file this way:\n\n```bash\ndocker run -d -p 1071:1071 -v server.ini:/usr/src/app/fuji_server/config/server.ini ghcr.io/pangaea-data-publisher/fuji\n```\n\nYou can also build the docker image from the source code:\n\n```bash\ndocker build -t \u003ctag_name\u003e .\ndocker run -d -p 1071:1071 \u003ctag_name\u003e\n```\n\n### Notes\n\nTo avoid Tika startup warning message, set environment variable `TIKA_LOG_PATH`. For more information, see [https://github.com/chrismattmann/tika-python](https://github.com/chrismattmann/tika-python)\n\nIf you receive the exception `urllib2.URLError: \u003curlopen error [SSL: CERTIFICATE_VERIFY_FAILED]` on macOS, run the install command shipped with Python:\n`./Install\\ Certificates.command`.\n\nF-UJI is using [basic authentication](https://en.wikipedia.org/wiki/Basic_access_authentication), so username and password have to be provided for each REST call which can be configured in `fuji_server/config/users.py`.\n\n#### GitHub API\n\nF-UJI can optionally use the GitHub API to evaluate software repositories hosted on GitHub.\nUnauthorised requests to the GitHub API are subject to a very low rate limit however, so it's recommended to authenticate using a personal access token.\n\nTo create an access token, log into your GitHub account and navigate to \u003chttps://github.com/settings/tokens\u003e, either by clicking on the link or through Settings -\u003e Developer Settings -\u003e Personal access tokens -\u003e Tokens (classic). Next, click \"Generate new token\" and select \"Generate new token (classic)\" from the drop-down menu.\n\nWrite the purpose of the token into the \"Note\" field (for example, *F-UJI deployment*) and set a suitable expiration date. Leave all the checkboxes underneath *unchecked*.\n\n\u003e Note: When the token expires, you will receive an e-mail asking you to renew it if you still need it. The e-mail will provide a link to do so, and you will only need to change the token in the f-uji configuration as described below to continue using it. Setting no expiration date for a token is thus not recommended.\n\nWhen you click \"Generate new token\" at the bottom of the page, the new token will be displayed. Make a note of it now.\n\nTo use F-UJI with a single access token, open [`fuji_server/config/github.ini`](./fuji_server/config/github.ini) locally and set `token` to the token you just created. When F-UJI receives an evaluation request that uses the GitHub API, it will run this request authenticated as your account.\n\nIf you still run into rate limiting issues, you can use multiple GitHub API tokens.\nThese need to be generated by different GitHub accounts, as the rate limit applies to the user, not the token.\nF-UJI will automatically switch to another token if the rate limit is near.\nTo do so, create a local file in [`fuji_server/data/`](./fuji_server/data/), called e.g. `github_api_tokens.txt`. Put all API tokens in that file, one token on each line. Then, open [`fuji_server/config/github.ini`](./fuji_server/config/github.ini) locally and set `token_file` to the absolute path to your local API token file.\n\n\u003e Note: If you push a change containing a GitHub API token, GitHub will usually recognise this and invalidate the token immediately. You will need to regenerate the token. Please take care not to publish your API tokens anywhere. Even though they have very limited scope if you leave all the checkboxes unchecked during creation, they can allow someone else to run a request in your name.\n\n## Development\n\nFirst, make sure to read the [contribution guidelines](./CONTRIBUTING.md).\nThey include instructions on how to set up your environment with `pre-commit` and how to run the tests.\n\nThe repository includes a [simple web client](./simpleclient/) suitable for interacting with the API during development.\nOne way to run it would be with a LEMP stack (Linux, Nginx, MySQL, PHP), which is described in the following.\n\nFirst, install the necessary packages:\n\n```bash\nsudo apt-get update\nsudo apt-get install nginx\nsudo ufw allow 'Nginx HTTP'\nsudo service mysql start  # expects that mysql is already installed, if not run sudo apt install mysql-server\nsudo service nginx start\nsudo apt install php8.1-fpm php-mysql\nsudo apt install php8.1-curl\nsudo phpenmod curl\n```\n\nNext, configure the service by running `sudo vim /etc/nginx/sites-available/fuji-dev` and paste:\n\n```php\nserver {\n    listen 9000;\n    server_name fuji-dev;\n    root /var/www/fuji-dev;\n\n    index index.php;\n\n    location / {\n        try_files $uri $uri/ =404;\n    }\n\n    location ~ \\.php$ {\n        include snippets/fastcgi-php.conf;\n        fastcgi_pass unix:/var/run/php/php8.1-fpm.sock;\n        fastcgi_read_timeout 3600s;\n     }\n\n    location ~ /\\.ht {\n        deny all;\n    }\n}\n```\n\nLink `simpleclient/index.php` and `simpleclient/icons/` to `/var/www/fuji-dev` by running `sudo ln \u003cpath_to_fuji\u003e/fuji/simpleclient/* /var/www/fuji-dev/`. You might need to adjust the file permissions to allow non-root writes.\n\nNext,\n```bash\nsudo ln -s /etc/nginx/sites-available/fuji-dev /etc/nginx/sites-enabled/\nsudo nginx -t\nsudo service nginx reload\nsudo service php8.1-fpm start\n```\n\nThe web client should now be available at \u003chttp://localhost:9000/\u003e. Make sure to adjust the username and password in [`simpleclient/index.php`](./simpleclient/index.php).\n\nAfter a restart, it may be necessary to start the services again:\n\n```bash\nsudo service php8.1-fpm start\nsudo service nginx start\npython -m fuji_server -c fuji_server/config/server.ini\n```\n\n### Component interaction (walkthrough)\n\nThis walkthrough can guide you through the comprehensive codebase.\n\nA good starting point is [`fair_object_controller/assess_by_id`](fuji_server/controllers/fair_object_controller.py#36).\nHere, we create a [`FAIRCheck`](fuji_server/controllers/fair_check.py) object called `ft`.\nThis reads the metrics YAML file during initialisation and will provide all the `check` methods.\n\nNext, several harvesting methods are called, first [`harvest_all_metadata`](fuji_server/controllers/fair_check.py#329), followed by [`harvest_re3_data`](fuji_server/controllers/fair_check.py#345) (Datacite) and [`harvest_github`](fuji_server/controllers/fair_check.py#366) and finally [`harvest_all_data`](fuji_server/controllers/fair_check.py#359).\nThe harvesters are implemented separately in [`harvester/`](./fuji_server/harvester/), and each of them collects different kinds of data.\nThis is regardless of the defined metrics, the harvesters always run.\n- The metadata harvester looks through HTML markup following schema.org, Dublincore etc., through signposting/typed links.\nIdeally, it can find things like author information or license names that way.\n- The data harvester is only run if the metadata harvester finds an `object_content_identifier` pointing at content files.\nThen, the data harvester runs over the files and checks things like the file format.\n- The Github harvester connects with the GitHub API to retrieve metadata and data from software repositories.\nIt relies on an access token being defined in [`config/github.cfg`](./fujji_server/config/github.cfg).\n\nAfter harvesting, all evaluators are called.\nEach specific evaluator, e.g. [`FAIREvaluatorLicense`](fuji_server/evaluators/fair_evaluator_license.py), is associated with a specific FsF and/or FAIR4RS metric.\nBefore the evaluator runs any checks on the harvested data, it asserts that its associated metric is listed in the metrics YAML file.\nOnly if it is, the evaluator runs through and computes a local score.\n\nIn the end, all scores are aggregated into F, A, I, R scores.\n\n### Adding support for new metrics\n\nStart by adding a new metrics YAML file in [`yaml/`](./fuji_server/yaml).\nIts name has to match the following regular expression: `(metrics_v)?([0-9]+\\.[0-9]+)(_[a-z]+)?(\\.yaml)`,\nand the content should be structured similarly to the existing metric files.\n\nMetric names are tested for validity using regular expressions throughout the code.\nIf your metric names do not match those, not all components of the tool will execute as expected, so make sure to adjust the expressions.\nRegular expression groups are also used for mapping to F, A, I, R categories for scoring, and debug messages are only displayed if they are associated with a valid metric.\n\nEvaluators are mapped to metrics in their `__init__` methods, so adjust existing evaluators to associate with your metric as well or define new evaluators if needed.\nThe multiple test methods within an evaluator also check whether their specific test is defined.\n[`FAIREvaluatorLicense`](fuji_server/evaluators/fair_evaluator_license.py) is an example of an evaluator corresponding to metrics from different sources.\n\nFor each metric, the maturity is determined as the maximum of the maturity associated with each passed test.\nThis means that if a test indicating maturity 3 is passed and one indicating maturity 2 is not passed, the metric will still be shown to be fulfilled with maturity 3.\n\n### Community specific metrics\n\nSome, not all, metrics can be configured using the following guidelines:\n[Metrics configuration guide](https://github.com/pangaea-data-publisher/fuji/blob/master/metrics_configuration.md)\n\n### Updates to the API\n\nMaking changes to the API requires re-generating parts of the code using Swagger.\nFirst, edit [`fuji_server/yaml/openapi.yaml`](fuji_server/yaml/openapi.yaml).\nThen, use the [Swagger Editor](https://editor.swagger.io/) to generate a python-flask server.\nThe zipped files should be automatically downloaded.\nUnzip it.\n\nNext:\n1. Place the files in `swagger_server/models` into `fuji_server/models`, except `swagger_server/models/__init__.py`.\n2. Rename all occurrences of `swagger_server` to `fuji_server`.\n3. Add the content of `swagger_server/models/__init__.py` into `fuji_server/__init__.py`.\n\nUnfortunately, the Swagger Editor doesn't always produce code that is compliant with PEP standards.\nRun `pre-commit run` (or try to commit) and fix any errors that cannot be automatically fixed.\n\n## License\nThis project is licensed under the MIT License; for more details, see the [LICENSE](https://github.com/pangaea-data-publisher/fuji/blob/master/LICENSE) file.\n\n\n## Acknowledgements\n\nF-UJI is a result of the [FAIRsFAIR](https://www.fairsfair.eu/) “Fostering FAIR Data Practices In Europe” project which received funding from the European Union’s Horizon 2020 project call H2020-INFRAEOSC-2018-2020 (grant agreement 831558).\n\nThe project was also supported through our contributors by the [Helmholtz Metadata Collaboration (HMC)](https://www.helmholtz-metadaten.de/en), an incubator-platform of the Helmholtz Association within the framework of the Information and Data Science strategic initiative.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpangaea-data-publisher%2Ffuji","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fpangaea-data-publisher%2Ffuji","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpangaea-data-publisher%2Ffuji/lists"}