{"id":18818404,"url":"https://github.com/smithsonian/smax-postgres","last_synced_at":"2025-09-11T05:14:24.839Z","repository":{"id":258387975,"uuid":"804754480","full_name":"Smithsonian/smax-postgres","owner":"Smithsonian","description":"Record SMA-X history in PostgreSQL / TimescaleDB","archived":false,"fork":false,"pushed_at":"2025-02-17T07:45:45.000Z","size":689,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":3,"default_branch":"main","last_synced_at":"2025-02-17T08:32:33.456Z","etag":null,"topics":["historical-data","real-time-systems","structured-data","time-series"],"latest_commit_sha":null,"homepage":"https://smithsonian.github.io/smax-postgres/","language":"C","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"unlicense","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Smithsonian.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null},"funding":{"github":"attipaci","patreon":null,"open_collective":null,"ko_fi":null,"tidelift":null,"community_bridge":null,"liberapay":null,"issuehunt":null,"lfx_crowdfunding":null,"polar":null,"buy_me_a_coffee":null,"thanks_dev":null,"custom":null}},"created_at":"2024-05-23T07:50:17.000Z","updated_at":"2025-02-17T07:45:48.000Z","dependencies_parsed_at":"2024-10-18T19:19:19.294Z","dependency_job_id":"7e3fd81e-5622-4d48-a51d-300afcba8766","html_url":"https://github.com/Smithsonian/smax-postgres","commit_stats":null,"previous_names":["smithsonian/smax-postgres"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Smithsonian%2Fsmax-postgres","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Smithsonian%2Fsmax-postgres/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Smithsonian%2Fsmax-postgres/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Smithsonian%2Fsmax-postgres/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Smithsonian","download_url":"https://codeload.github.com/Smithsonian/smax-postgres/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":239753733,"owners_count":19691162,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["historical-data","real-time-systems","structured-data","time-series"],"created_at":"2024-11-08T00:16:38.417Z","updated_at":"2025-09-11T05:14:24.827Z","avatar_url":"https://github.com/Smithsonian.png","language":"C","funding_links":["https://github.com/sponsors/attipaci"],"categories":[],"sub_categories":[],"readme":"![Build Status](https://github.com/Smithsonian/smax-postgres/actions/workflows/build.yml/badge.svg)\n![Static Analysis](https://github.com/Smithsonian/smax-postgres/actions/workflows/analyze.yml/badge.svg)\n\u003ca href=\"https://smithsonian.github.io/smax-postgres/apidoc/html/files.html\"\u003e\n ![API documentation](https://github.com/Smithsonian/smax-postgres/actions/workflows/dox.yml/badge.svg)\n\u003c/a\u003e\n\u003ca href=\"https://smithsonian.github.io/smax-postgres/index.html\"\u003e\n ![Project page](https://github.com/Smithsonian/smax-postgres/actions/workflows/pages/pages-build-deployment/badge.svg)\n\u003c/a\u003e\n\n\u003cpicture\u003e\n  \u003csource srcset=\"resources/CfA-logo-dark.png\" alt=\"CfA logo\" media=\"(prefers-color-scheme: dark)\"/\u003e\n  \u003csource srcset=\"resources/CfA-logo.png\" alt=\"CfA logo\" media=\"(prefers-color-scheme: light)\"/\u003e\n  \u003cimg src=\"resources/CfA-logo.png\" alt=\"CfA logo\" width=\"400\" height=\"67\" align=\"right\"/\u003e\n\u003c/picture\u003e\n\u003cbr clear=\"all\"\u003e\n\n\n# smax-postgres\n\nRecord [SMA-X](https://docs.google.com/document/d/1eYbWDClKkV7JnJxv4MxuNBNV47dFXuUWu7C4Ve_YTf0/edit?usp=sharing) \nhistory in PostgreSQL / TimescaleDB. It is free to use, in any way you like, without licensing restrictions.\n\n - [API documentation](https://smithsonian.github.io/smax-postgres/apidoc/html/files.html)\n - [Project pages](https://smithsonian.github.io/smax-postgres) on github.io\n\nAuthor: Attila Kovacs\n\nLast Updated: 18 September 2024\n\n----------------------------------------------------------------------------------------------------------------------\n\n## Table of Contents\n\n - [Introduction](#smaxpg-introduction)\n - [Prerequisites](#smaxpg-prerequisites)\n - [Building `smax-postgres`](#building-smaxpg)\n - [Installation](#smaxpg-installation)\n - [Database organization (for clients)](#database-organization)\n - [Configuration reference](#configuration-reference)\n\n----------------------------------------------------------------------------------------------------------------------\n\n\u003ca name=\"smaxpg-introduction\"\u003e\u003c/a\u003e\n## Introduction\n\n`smax-postgres` is a daemon application, which can collect data from an \n[SMA information eXchange (SMA-X)](https://docs.google.com/document/d/1eYbWDClKkV7JnJxv4MxuNBNV47dFXuUWu7C4Ve_YTf0/edit?usp=sharing) \nrealtime database and insert these into a __PostgreSQL__ database to create a time-series historical record for all or \nselected SMA-X variables. The program is highly customizable and supports both regular updates for changing variables \nas well as regular snapshots of all selected SMA-X variables.\n\n\n\n\u003ca name=\"smaxpg-prerequisites\"\u003e\u003c/a\u003e\n## Prerequisites\n\nThe `smax-postgres` application has build and runtime dependencies on the following software:\n\n - __PostgreSQL__ installation and development files (`libpq.so` and `lipq.fe.h`).\n - [Smithsonian/smax-clib](https://github.com/Smithsonian/smax-clib)\n - [Smithsonian/redisx](https://github.com/Smithsonian/redisx)\n - [Smithsonian/xchange](https://github.com/Smithsonian/xchange)\n - __Popt__ development libraries (`libpopt-dev`in Debian, or `popt-devel` in RPM distros)\n - (_optional_) __TimescaleDB__ extensions.\n - (_optional_) __systemd__ development files (`libsystemd.so` and `sd-daemon.h`).\n\nAdditionally, to configure your SMA-X server, you will need the \n[Smithsonian/smax-server](https://github.com/Smithsonian/smax-server) repo also.\n\n----------------------------------------------------------------------------------------------------------------------\n\n\u003ca name=\"building-smaxpg\"\u003e\u003c/a\u003e\n## Building `smax-postgres`\n\nYou can configure the build, either by editing `config.mk` or else by defining the relevant environment variables \nprior to invoking `make`. The following build variables can be configured:\n\n - `PGVER`: Major version of PostgreSQL to use (default: 16). Needed e.g. for systemd integration.\n\n - `PGDIR`: Root directory of a specific PostgreSQL installation to build against (not set by default). If not set\n   we'll build against the default PostgreSQL available on your system.\n  \n - `SYSTEMD`: Sets whether to compile with SystemD integration (needs `libsystemd.so` and `sd-daemon.h`). Default\n   is 1 (enabled).\n   \n - `CC`: The C compiler to use (default: `gcc`).\n\n - `CPPFLAGS`: C preprocessor flags, such as externally defined compiler constants.\n \n - `CFLAGS`: Flags to pass onto the C compiler (default: `-g -Os -Wall`). Note, `-Iinclude` will be added \n   automatically.\n   \n - `CSTANDARD`: Optionally, specify the C standard to compile for, e.g. `c99` to compile for the C99 standard. If\n   defined then `-std=$(CSTANDARD)` is added to `CFLAGS` automatically.\n   \n - `WEXTRA`: If set to 1, `-Wextra` is added to `CFLAGS` automatically.\n   \n - `FORTIFY`: If set it will set the `_FORTIFY_SOURCE` macro to the specified value (`gcc` supports values 1 \n   through 3). It affords varying levels of extra compile time / runtime checks.\n   \n - `LDFLAGS`: Extra linker flags (default: _not set_). Note, `-lm -pthread -lsmax -lredisx -lxchange -lpq -lpopt` will \n   be added automatically.\n\n - `CHECKEXTRA`: Extra options to pass to `cppcheck` for the `make check` target\n \n - `DOXYGEN`: Specify the `doxygen` executable to use for generating documentation. If not set (default), `make` will\n   use `doxygen` in your `PATH` (if any). You can also set it to `none` to disable document generation and the\n   checking for a usable `doxygen` version entirely.\n \n - `XCHANGE`: If the [Smithsonian/xchange](https://github.com/Smithsonian/xchange) library is not installed on your\n   system (e.g. under `/usr`) set `XCHANGE` to where the distribution can be found. The build will expect to find \n   `xchange.h` under `$(XCHANGE)/include` and `libxchange.so` / `libxchange.a` under `$(XCHANGE)/lib` or else in the \n   default `LD_LIBRARY_PATH`.\n \n - `REDISX`: If the [Smithsonian/redisx](https://github.com/Smithsonian/redisx) library is not installed on your\n   system (e.g. under `/usr`) set `REDISX` to where the distribution can be found. The build will expect to find \n   `redisx.h` under `$(REDISX)/include` and `libredisx.so` / `libredisx.a` under `$(REDISX)/lib` or else in the \n   default `LD_LIBRARY_PATH`.\n   \n - `SMAXLIB`: If the [Smithsonian/smax-clib](https://github.com/Smithsonian/smax-clib) library is not installed on \n   your system (e.g. under `/usr`) set `SMAXLIB` to where the distribution can be found. The build will expect to find \n   `smax.h` under `$(SMAXLIB)/include` and `libsmax.so` / `libsmax.a` under `$(SMAXLIB)/lib` or else in the default \n   `LD_LIBRARY_PATH`.\n \nAfter configuring, you can simply run `make`, which will build `bin/smax-postgres`, and user documentation. You may \nalso build other `make` target(s). (You can use `make help` to get a summary of the available `make` targets). \n\nNow you may compile `smax-postgres`:\n\n```bash\n  $ make\n```\n\nAfter building the library you can install the above components to the desired locations on your system. For a \nsystem-wide install you may simply run:\n\n```bash\n  $ sudo make install\n```\n\nOr, to install in some other locations, you may set a prefix and/or `DESTDIR`. For example, to install under `/opt` \ninstead, you can:\n\n```bash\n  $ sudo make prefix=\"/opt\" install\n```\n\nOr, to stage the installation (to `/usr`) under a 'build root':\n\n```bash\n  $ make DESTDIR=\"/tmp/stage\" install\n```\n\n----------------------------------------------------------------------------------------------------------------------\n\n\u003ca name=\"smaxpg-installation\"\u003e\u003c/a\u003e\n## Installation\n\nPrior to installation, you should check that the PostgreSQL service name is correct in `smax-postgres.service`, and\nedit it as necessary for your system configuration. You may also edit the `cfg/smax-postgres.cfg` now, or after it\nis installed.\n\nProvided the build was successful, you can install the executables, configuration files, and optionally the SystemD \nunit files via:\n\n```bash\n  $ sudo make install\n```\n\n(When installing at the SMA, you may want `make install-sma` instead, to install with SMA-specific configuration).\nIn case of SystemD integration you should also reload the SystemD daemon so `smax-postgres.service` can be enabled \nand managed as desired:\n\n```bash\n  $ sudo systemd daemon-reload\n```\n\nAfter that you can start the service as:\n\n```bash\n  $ sudo systemd start smax-postgres\n```\n\n### Staging / advanced installation\n\nBy default, `make install` will install the `smax-postgress` executable to `/usr/bin`, configuration under `/etc/`,\nSystemD service unit under `/etc/systemd/system`, and documentation under `/usr/share/doc/smax-postgres/`. Instead of\n`/usr`, you may want to install into another destination, such as `/opt/` or `/usr/local`. You can do that by setting\nthe `DESTDIR` environment prior to `make install`, e.g.:\n\n```bash\n  $ export DESTDIR=\"/opt\"\n```\n\nAdditionally, you can also stage the installation under a different root, by setting the `PREFIX` environment variable,\ne.g.:\n\n```bash\n  $ export PREFIX=\"~/rmpbuild/BUILD/smax-postgres\"\n```\n\n### Standard error/output with SystemD integration\n\nIn case of SystemD integration, errors will get logged to the journal, and can be investigated by `journalctl`. E.g. to\nsee the errors in the last 3 hours, you may:\n\n```bash\n  $ journalctl -u smax-postgres --since \"3 hours ago\"\n```\n\nStandard output is logged to the file `/var/log/smax-postgres.out` for the current session. Restarting `smax-postgres` will \nstart a new file. (Because of buffering, there may be a long lag before you'll see stuff appear in this log file.)\n\n\n\n### Initial setup of the SQL database\n\nPrior to using `smax-postgres`, you will need to configure users and access privileges (roles) for the database, and may \nwant to create the database instance manually. You will need to create a user designated for the `smax-postgres` program, \nand specify its credentials in the `smax-postgres` configuration file. This user will not require `CREATEDB` permission, \nbut it will need permission to create tables in the existing database and to insert data or to search the tables\n(i.e. read/write privileges). Additionally, you may also create the designated database instance assigned to whatever \nuser to own. If you create the database manually, do not forget to set the name of the designated database in the \n`smax-postgres` configuration file.\n\nNormally `smax-postgres` will assume that the database to use has been fully set up, including a 'titles' table \n(containing 2 columns: _text_ variable IDs, and auto-incremented integer _serial_ numbers). However, `smax-postgres` \ncan create the database and set up the required 'titles' table as needed (including indexing, and TimescaleDB \nextension as appropriate), when launched  with the `-b` (or `--bootstrap`) option.\n\n```bash\n  $ smax-postgres -c /usr/local/etc/smax-postgres/myconfig.cfg -b\n```\n\nWill log into the existing (new) database using the credentials specified in the configuration file \n`/usr/local/etc/smax-postgres/myconfig.cfg`, then configures that database (e.g. set up the TimescaleSB extension as \nappropriate), and creates the 'titles' table and its index.\n\nYou may also let the bootstrapping process create the database itself, in which case you may have to provide the\npassword for the 'postgres' admin account, or the credentials for another account with `CREATEDB` privileges with \nthe necessary privileges to create databases. E.g.:\n\n```bash\n  $ smax-postgres -c /usr/local/etc/myconfig.cfg -b -p \"S3cur1ty!\"\n```\n\nwill attempt create the database as the 'postgres' admin with password 'S3curity!', before proceeding to configure the \nnewly created database as the user designated for the `smax-postgres` program. (Alternatively, you may use the `-a` and \n`-p` options together to create the database with another privileged user). The newly created database will be \nautomatically assigned to the designated `smax-postgres` user as its owner).\n\nOnce the database is configured, you will not need the `-b` option again (but it also will not wreck the previous\ninitialization if accidentally used again after the initial setup).\n\n----------------------------------------------------------------------------------------------------------------------\n\n\n\u003ca name=\"database-organization\"\u003e\u003c/a\u003e\n## Database organization (for clients)\n\nEach SMA-X variable has its own time-series data table in the SQL DB. These tables are named `var_\u003ctid\u003e`, where \n`\u003ctid\u003e` is a 6-digit serial number, e.g. `var_000001` for the first variable added to the SQL database. The variable \nname to `\u003ctid\u003e` pairings are listed in the `titles` table, which contains just two columns: (1) the full SMA-X \nvariable id (_text_), and (2) the corresponding `\u003ctid\u003e` (integer _serial_) used in the SQL database to store the time \nseries of the variable. (Thus, to figure out what table stored data for a given SMA-X variable you will need to search \nthe 'titles' table for the `\u003ctid\u003e` first.)\n\nThe time series data in the `var_\u003ctid\u003e` tables contains (at least) 2 + `n` columns for an SMA-X variable, which has \n`n` array elements. The first column is the UTC timestamp at which the data was pulled from the SMA-X database. The \nsecond is an integer 'age' (in seconds) that informs how much before the pull was the last time that variable was \nupdated in the SMA-X database. After that, the remaining columns list the array elements stored in the SMA-X variable. \n(Thus scalar entries will have just one additional column, labeled as 'c0').\n\nBecause the SMA-X variables may have dynamic types and array dimensions, the SQL tables may automatically expand \nto an enclosing type (for example if a variable changes from `int16` to `int32` or to a `float32`), and columns will \nbe added as necessary to store an expanded set of array elements. When SMA-X data 'shrinks',  containing fewer \nelements than the existing SQL record, the SQL entry will be padded with `NULL` values as necessary.\n\nIn addition to the time-series data stored in the SQL database, it also stores versioned metadata in tables with\nmatching `\u003ctid\u003e` values. For example, the metadata for the time-series `var_000001` is stored in the table \n`var_000001_meta`. Metadata tables contained serial-numbered versions of infrequently changing metadata, such\nas array dimensions and shapes (scalar values are stored with `ndim = 0` and `shape = NULL`), associated physical \nunits, and downsampling factors. Each metadata entry is timestamped also to indicate when a change (if any) \noccurred in these characteristics.\n\nThus, to query an SMA-X variable `system:subsystem:property`, you first want to find the 'tid' of the variable in \n'titles':\n\n```sql\n  SELECT tid FROM titles WHERE name = 'system.subsystem.property';\n```\n\nSay the query returns the tid `192`, then the time-series for that variable will be stored in the table named\n`var_000192`, while metadata versions are stored in the table named `var_000192_meta`.\n\n----------------------------------------------------------------------------------------------------------------------\n\n\u003ca name=\"configuration-reference\"\u003e\u003c/a\u003e\n## Configuration Reference\n\nSee `cfg/example.cfg` as an example configuration file. Based on it may create your own configuration file, which \nyou can then load via the `-c` option to `smax-postgres` at startup. If using SystemD integration, you may want to update \n`/etc/systemd/system/smax-postgres.service` to load the configuration file from the location of your choice when the\nservice is started via `systemd`.\n\n### Database configuration options\n\n#### `smax_server \u003chost\u003e`\n\nHost name or IP address of the SMA-X server (default 'smax').\n\n#### `sql_auth \u003cpassword\u003e`\n\nPassword for authenticating user on the SQL server (no default).\n\n#### `sql_db` \u003cdb-name\u003e`\n\nSQL database name to use (default is 'smax_db'). \n\n#### `sql_server \u003chost\u003e`\n\nHost name or IP address of the SMA-X server (default 'localhost').\n\n#### `sql_user \u003cuser-name\u003e`\n\nSQL user name to use (default is 'smax_db'). \n\n#### `use_hypertables \u003c1|0\u003e`\n\nDetermines whether to create hypertables via the TimescaleDB extension. The value 1 enables, 0 disables the used of \nhypertables (the default is to not use hypertables). TimescaleDB hypertables allow for faster access of time series \ndata by organizing large datasets into smaller blocks of data, which can be handled more efficiently.\n\n### Update frequency options\n\n#### interval specification\n\nTimescales may be specified by a numeric value (integer or decimal) followed immediately by a unit designator, such as \n'1d' for one day. Alternatively, the value 'none' can be used to disable a timescale-specific option. The following \ntimescale units are understood:\n\n |  unit     | description |\n | --------- | ----------- |\n |   `s`     | second(s)   |\n |   `m`     | minute(s)   |\n |   `h`     | hour(s)     |\n |   `d`     | day(s)      |\n |   `w`     | week(s)     |\n |   `y`     | year(s)     |\n\n\n#### `snapshot_interval \u003cinterval\u003e`\n\nSpecifies the interval at which all designated variables are pushed into the SQL database, regardless whether they \nhave changed or not since the last time they were pushed (default '1h'). See the section further above on interval \nspecifications.\n\n#### `update_interval \u003cinterval\u003e`\n\nSpecifies the regular interval at which to push changing variables into the database (default: '1m'). See the section \nfurtherabove on interval specifications. Variables that have no changed since the last update will be excluded from \nthe regular updates until it is time for the next full snapshot. The snapshot interval is controlled separately (see \nabove).\n\n\n### Variable-specific options\n\n#### glob patterns\n\nSMA-X variables can be specified individually or via \n[glob patterns](https://man7.org/linux/man-pages/man7/glob.7.html), similarly to how these are used in UNIX shells. \n\n\n#### `always \u003cpattern\u003e`\n\nSpecifies a variable or a glob pattern of variables, which are to be logged into the SQL database regardless of all \nother directives, which may otherwise limit if and when they are to be pushed. Reserve using this option only for the \nmost critical cases, when the other configuration options do not provide the desired level of assurances for some \nabsolutely critical data points.\n\n#### `exclude \u003cpattern\u003e`\n\nSpecifies a variable or a glob pattern of variables that are to be excluded from logging to the SQL database. \n`exclude` and `include` directives take effect in the order they were specified, so for a given variable only the last \n`include` or `exclude` statement, which pertains to it, will decide whether or not to log that given variable. \nVariables that are configured with an `always` directive will be logged to the SQL database regardless of any \nexclusions that may have been specified, either before or after. By default all metadata variables (ones whose name \nbegin with `\u003c`) and all temporary variables (whose names begin with an underscore `_`) are excluded from logging, \nunless they are explicitly re-included.\n\n#### `include \u003cpattern\u003e`\n\nSpecifies a variable or a glob pattern of variables that are to be included for logging to the SQL database. \n`exclude` and `include` directives take effect in the order they were specified, so for a given variable only the last \n`include` or `exclude` statement, which pertains to it, will decide whether or not to log that given variable. \n\n\n#### `max_age \u003cinterval\u003e`\n\nSets a maximum age for variables to push to the SQL database (default: '90d'). Variables that have not been updated in \nSMA-X for longer than the specified interval will not be pushed to the SQL database. Variables that are configured \nwith an `always` directive will be logged to the SQL database regardless of their age.\n\n#### `max_size \u003cbytes\u003e`\n\nSets a maximum byte size for variables to push to the SQL database (default: '1024'). Variables that have larger \nbinary representations (after downsampling via the `sampling` directive, if any) will not be logged to the database to \navoid bloating. However, variables explicitly configured via an `always` directive will be logged to the SQL database \nregardless of their storage requirements.\n\n#### `sample \u003cn\u003e \u003cpattern\u003e`\n\nLog sparse samples of data for a variable or a glob pattern of variables. In some cases you may store large arrays in \nthe SMA-X database, logging of which may bloat the time series history stored in the SQL database. However, you may \nwant to still get a preview of what that data was, by storing every n'th sample of the original only. For example, for \nan array with 1000 elemenrs, you may want to store say 20 samples. Setting `\u003cn\u003e` to 50 will achieve that, by storing \nevery 50th element in the SQL database only. (Still, the SQL database will store the original dimensionality of the \ndownsampled variables, and note the downsampling factor used also as metadata).\n\n-----------------------------------------------------------------------------\nCopyright (C) 2025 Attila Kovács\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsmithsonian%2Fsmax-postgres","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsmithsonian%2Fsmax-postgres","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsmithsonian%2Fsmax-postgres/lists"}