{"id":13929081,"url":"https://github.com/juji-io/datalevin","last_synced_at":"2025-05-13T19:10:06.669Z","repository":{"id":37501964,"uuid":"270858421","full_name":"juji-io/datalevin","owner":"juji-io","description":"A simple, fast and versatile Datalog database","archived":false,"fork":false,"pushed_at":"2025-04-22T06:11:06.000Z","size":48646,"stargazers_count":1243,"open_issues_count":38,"forks_count":71,"subscribers_count":25,"default_branch":"master","last_synced_at":"2025-04-22T07:38:12.741Z","etag":null,"topics":["client-server-database","embedded-database","fulltext-search","key-value-store","vector-database"],"latest_commit_sha":null,"homepage":"https://github.com/juji-io/datalevin","language":"Clojure","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"epl-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/juji-io.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null},"funding":{"github":"huahaiy"}},"created_at":"2020-06-08T23:53:52.000Z","updated_at":"2025-04-22T06:11:10.000Z","dependencies_parsed_at":"2023-10-01T23:30:07.893Z","dependency_job_id":"adadd894-72d8-4dd6-8358-f942d3fe54ed","html_url":"https://github.com/juji-io/datalevin","commit_stats":{"total_commits":1933,"total_committers":76,"mean_commits":25.43421052631579,"dds":0.281427832384894,"last_synced_commit":"13f83abcd29ecb850a3694e774c7fd34191adc14"},"previous_names":[],"tags_count":219,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/juji-io%2Fdatalevin","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/juji-io%2Fdatalevin/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/juji-io%2Fdatalevin/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/juji-io%2Fdatalevin/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/juji-io","download_url":"https://codeload.github.com/juji-io/datalevin/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":251299817,"owners_count":21567315,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["client-server-database","embedded-database","fulltext-search","key-value-store","vector-database"],"created_at":"2024-08-07T18:02:06.204Z","updated_at":"2025-05-13T19:10:06.662Z","avatar_url":"https://github.com/juji-io.png","language":"Clojure","funding_links":["https://github.com/sponsors/huahaiy"],"categories":["Research \u0026 Data Analysis","others","Clojure","数据库","Database"],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\u003cimg src=\"logo.png\" alt=\"datalevin logo\"\nheight=\"140\"\u003e\u003c/img\u003e\u003c/p\u003e\n\u003ch1 align=\"center\"\u003eDatalevin\u003c/h1\u003e\n\u003cp align=\"center\"\u003e 🧘 Simple, fast and versatile Datalog database for everyone\n💽 \u003c/p\u003e\n\u003cp align=\"center\"\u003e\n\u003ca href=\"https://cljdoc.org/d/datalevin/datalevin\"\u003e\u003cimg\nsrc=\"https://cljdoc.org/badge/datalevin/datalevin\" alt=\"datalevin on\ncljdoc\"\u003e\u003c/img\u003e\u003c/a\u003e\n\u003ca href=\"https://clojars.org/datalevin\"\u003e\u003cimg\nsrc=\"https://img.shields.io/clojars/v/datalevin.svg?color=success\"\nalt=\"datalevin on clojars\"\u003e\u003c/img\u003e\u003c/a\u003e\n\u003ca\nhref=\"https://github.com/juji-io/datalevin/blob/master/doc/install.md#babashka-pod\"\u003e\u003cimg\nsrc=\"https://raw.githubusercontent.com/babashka/babashka/master/logo/badge.svg\"\nalt=\"bb compatible\"\u003e\u003c/img\u003e\u003c/a\u003e\n\u003c/p\u003e\n\u003cp align=\"center\"\u003e\n\u003ca href=\"https://github.com/juji-io/datalevin/actions\"\u003e\u003cimg\nsrc=\"https://github.com/juji-io/datalevin/actions/workflows/release.binaries.yml/badge.svg\"\nalt=\"datalevin linux/macos amd64 build status\"\u003e\u003c/img\u003e\u003c/a\u003e\n\u003ca href=\"https://ci.appveyor.com/project/huahaiy/datalevin\"\u003e\u003cimg\nsrc=\"https://ci.appveyor.com/api/projects/status/github/juji-io/datalevin?svg=true\"\nalt=\"datalevin windows build status\"\u003e\u003c/img\u003e\u003c/a\u003e\n\u003ca href=\"https://cirrus-ci.com/github/juji-io/datalevin\"\u003e\u003cimg\nsrc=\"https://api.cirrus-ci.com/github/juji-io/datalevin.svg\" alt=\"datalevin\narm64 build status\"\u003e\u003c/img\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\n\u003e I love Datalog, why hasn't everyone used this already?\n\n**Datalevin** (/ˈdadə ˈlevən/, \"levin\" means \"lightning\") is a simple durable\n[Datalog](https://en.wikipedia.org/wiki/Datalog) database. Here's what a Datalog\nquery looks like in Datalevin:\n\n```Clojure\n(d/q '[:find  ?name ?total\n       :in    $ ?year\n       :where [?sales :sales/year ?year]\n              [?sales :sales/total ?total]\n              [?sales :sales/customer ?customer]\n              [?customer :customers/name ?name]]\n      (d/db conn) 2024)\n```\n\n## :question: Why\n\nThe rationale is to have a simple, fast and open source Datalog query engine\nrunning on durable storage.\n\nIt is our observation that many developers prefer\nthe flavor of Datalog popularized by [Datomic®](https://www.datomic.com) over\nany flavor of SQL, once they get to use it. Perhaps it is because Datalog is\nmore declarative and composable than SQL, e.g. the automatic implicit joins seem\nto be its killer feature. In addition, the recursive rules feature of Datalog\nmakes it suitable for graph processing and deductive reasoning.\n\nThe feature set of Datomic® may not be a good fit for some use cases. One thing\nthat may [confuse some\nusers](https://vvvvalvalval.github.io/posts/2017-07-08-Datomic-this-is-not-the-history-youre-looking-for.html)\nis its [temporal\nfeatures](https://docs.datomic.com/cloud/whatis/data-model.html#time-model). To\nkeep things simple and familiar, Datalevin behaves the same way as most other\ndatabases: when data are deleted, they are gone. Datalevin also follows the\nwidely accepted principles of ACID, instead of introducing [unusual\nsemantics](https://jepsen.io/analyses/datomic-pro-1.0.7075).\n\nIn addition to support Datomic® flavor of Datalog query language, Datalevin has\na [novel cost-based query optimizer](doc/query.md) with a much better query\nperformance, which is [competitive](benchmarks/JOB-bench) with popular SQL RDBMS\nsuch as PostgreSQL.\n\nDatalevin provides robust ACID transaction features on the basis of\n[LMDB](https://en.wikipedia.org/wiki/Lightning_Memory-Mapped_Database), known\nfor its high read performance. With built-in support of asynchronous\ntransaction, Datalevin can also handle [write](benchmarks/write-bench) intensive\nworkload, as well as storing large documents.\n\nDatalevin supports [vector database](doc/vector.md) features by integrating an\nefficient SIMD accelerated vector indexing and search\n[library](https://github.com/unum-cloud/usearch).\n\nDatalevin has a [novel full-text search engine](doc/search.md) that has\n[competitive](benchmarks/search-bench) search performance.\n\nDatalevin can be used as a fast key-value store for\n[EDN](https://en.wikipedia.org/wiki/Extensible_Data_Notation) data. The native\nEDN data capability of Datalevin should be beneficial for Clojure programs.\n\nDatalevin can be used as a library, embedded in applications to manage state,\ne.g. used like SQLite; or it can run in a networked\n[client/server](https://github.com/juji-io/datalevin/blob/master/doc/server.md)\nmode (default port is 8898) with full-fledged role-based access control (RBAC)\non the server, e.g. used like PostgreSQL; or it can be used as a [babashka\npod](https://github.com/babashka/pod-registry/blob/master/examples/datalevin.clj)\nfor shell scripting.\n\nMore information can be found in these articles and presentation:\n\n* [Achieving High Throughput and Low Latency through Adaptive Asynchronous Transaction](https://yyhh.org/blog/2025/02/achieving-high-throughput-and-low-latency-through-adaptive-asynchronous-transaction/)\n* [Competing for the JOB with a Triplestore](https://yyhh.org/blog/2024/09/competing-for-the-job-with-a-triplestore/)\n* [If I had to Pick One: Datalevin](https://vimsical.notion.site/If-I-Had-To-Pick-One-Datalevin-be5c4b62cda342278a10a5e5cdc2206d)\n* [T-Wand: Beat Lucene in Less Than 600 Lines of Code](https://yyhh.org/blog/2021/11/t-wand-beat-lucene-in-less-than-600-lines-of-code/)\n* [2020 London Clojurians Meetup](https://youtu.be/-5SrIUK6k5g)\n\n## :truck: [Installation](doc/install.md)\n\nAs a Clojure library, Datalevin is simple to add as a dependency to your Clojure\nproject. There are also several other installation options. Please see details in\n[Installation Documentation](doc/install.md)\n\n## :birthday: Upgrade\n\nPlease read\n[Upgrade\nDocumentation](https://github.com/juji-io/datalevin/blob/master/doc/upgrade.md)\nfor information regarding upgrading your existing Datalevin database from older\nversions.\n\n## :tada: Usage\n\nDatalevin is aimed to be a versatile database.\n\n### Use as a Datalog store\n\nIn addition to [our API doc](https://cljdoc.org/d/datalevin/datalevin),\nDatalevin has almost the same Datalog API as\n[Datascript](https://github.com/tonsky/datascript), which in turn has almost the\nsame API as Datomic®, please consult the abundant tutorials, guides and learning\nsites available online to learn about the usage of Datomic® flavor of Datalog.\n\nHere is a simple code example using Datalevin:\n\n```clojure\n(require '[datalevin.core :as d])\n\n;; Define an optional schema.\n;; Note that pre-defined schema is optional, as Datalevin does schema-on-write.\n;; However, attributes requiring special handling need to be defined in schema,\n;; e.g. range query, many cardinality, uniqueness, reference type, etc.\n;; Similar to Datascript, Datalevin schemas differ from Datomic®:\n;; - The schema must be a map of maps, not a vector of maps.\n;; - It is not `transact`ed into the db but passed when acquiring connections.\n;; - Use `update-schema` to update the schema of an open connection to a DB.\n(def schema {:aka  {:db/cardinality :db.cardinality/many}\n             ;; :db/valueType is optional, if unspecified, the attribute will be\n             ;; treated as EDN blobs, and may not be optimal for range queries\n             :name {:db/valueType :db.type/string\n                    :db/unique    :db.unique/identity}})\n\n;; Create DB on disk and connect to it, assume write permission to create the dir\n(def conn (d/get-conn \"/tmp/datalevin/mydb\" schema))\n;; or if you have a Datalevin server running on myhost with default port 8898\n;; (def conn (d/get-conn \"dtlv://myname:mypasswd@myhost/mydb\" schema))\n\n;; Transact some data\n;; `:nation` is not defined in schema, so it will be treated as an EDN blob\n(d/transact! conn\n            [{:name \"Frege\", :db/id -1, :nation \"France\", :aka [\"foo\" \"fred\"]}\n             {:name \"Peirce\", :db/id -2, :nation \"france\"}\n             {:name \"De Morgan\", :db/id -3, :nation \"English\"}])\n\n;; Query the data\n(d/q '[:find ?nation\n       :in $ ?alias\n       :where\n       [?e :aka ?alias]\n       [?e :nation ?nation]]\n     (d/db conn)\n     \"fred\")\n;; =\u003e #{[\"France\"]}\n\n;; Retract the name attribute of an entity\n(d/transact! conn [[:db/retract 1 :name \"Frege\"]])\n\n;; Pull the entity, now the name is gone\n(d/q '[:find (pull ?e [*])\n       :in $ ?alias\n       :where\n       [?e :aka ?alias]]\n     (d/db conn)\n     \"fred\")\n;; =\u003e ([{:db/id 1, :aka [\"foo\" \"fred\"], :nation \"France\"}])\n\n;; Close DB connection\n(d/close conn)\n```\n\n### Use as a key-value store\n\nDatalevin packages the underlying LMDB database as a convenient key-value store\nfor EDN data.\n\n```clojure\n(require '[datalevin.core :as d])\n(import '[java.util Date])\n\n;; Open a key value DB on disk and get the DB handle\n(def db (d/open-kv \"/tmp/datalevin/mykvdb\"))\n;; or if you have a Datalevin server running on myhost with default port 8898\n;; (def db (d/open-kv \"dtlv://myname:mypasswd@myhost/mykvdb\" schema))\n\n;; Define some table (called \"dbi\", or sub-databases in LMDB) names\n(def misc-table \"misc-test-table\")\n(def date-table \"date-test-table\")\n\n;; Open the tables\n(d/open-dbi db misc-table)\n(d/open-dbi db date-table)\n\n;; Transact some data, a transaction can put data into multiple tables\n;; Optionally, data type can be specified to help with range query\n(d/transact-kv\n  db\n  [[:put misc-table :datalevin \"Hello, world!\"]\n   [:put misc-table 42 {:saying \"So Long, and thanks for all the fish\"\n                        :source \"The Hitchhiker's Guide to the Galaxy\"}]\n   [:put date-table #inst \"1991-12-25\" \"USSR broke apart\" :instant]\n   [:put date-table #inst \"1989-11-09\" \"The fall of the Berlin Wall\" :instant]])\n\n;; Get the value with the key\n(d/get-value db misc-table :datalevin)\n;; =\u003e \"Hello, world!\"\n(d/get-value db misc-table 42)\n;; =\u003e {:saying \"So Long, and thanks for all the fish\",\n;;     :source \"The Hitchhiker's Guide to the Galaxy\"}\n\n\n;; Range query, from unix epoch time to now\n(d/get-range db date-table [:closed (Date. 0) (Date.)] :instant)\n;; =\u003e [[#inst \"1989-11-09T00:00:00.000-00:00\" \"The fall of the Berlin Wall\"]\n;;     [#inst \"1991-12-25T00:00:00.000-00:00\" \"USSR broke apart\"]]\n\n;; This returns a PersistentVector - e.g. reads all data in JVM memory\n(d/get-range db misc-table [:all])\n;; =\u003e [[42 {:saying \"So Long, and thanks for all the fish\",\n;;          :source \"The Hitchhiker's Guide to the Galaxy\"}]\n;;     [:datalevin \"Hello, world!\"]]\n\n;; This allows you to iterate over all DB keys inside a transaction.\n;; You can perform writes inside the transaction.\n;; Avoid long-lived transactions. Read transactions prevent reuse of pages freed by newer write transactions, thus the database can grow quickly.\n;; Write transactions prevent other write transactions, since writes are serialized.\n(d/visit db misc-table\n            (fn [kv]\n               (let [k (d/read-buffer (d/k kv) :data)]\n                  (when (= k 42)\n                    (d/transact-kv db [[:put misc-table 42 \"Don't panic\"]]))))\n              [:all])\n\n(d/get-range db misc-table [:all])\n;; =\u003e [[42 \"Don't panic\"] [:datalevin \"Hello, world!\"]]\n\n;; Delete some data\n(d/transact-kv db [[:del misc-table 42]])\n\n;; Now it's gone\n(d/get-value db misc-table 42)\n;; =\u003e nil\n\n;; Close key value db\n(d/close-kv db)\n```\n## :green_book: Documentation\n\nPlease refer to the [API\ndocumentation](https://cljdoc.org/d/datalevin/datalevin) for more details.\n\n## :rocket: Status\n\nDatalevin is extensively tested with property-based testing. It is also used\nin production at [Juji](https://juji.io), among other companies.\n\nRunning the [benchmark suite adopted from\nDatascript](https://github.com/juji-io/datalevin/tree/master/benchmarks/datascript-bench),\nwhich includes several queries on 100K random datoms, on a 2016 Ubuntu Linux server with an Intel i7 3.6GHz CPU and a 1TB SSD drive, here is how it looks.\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"benchmarks/datascript-bench/Read.png\" alt=\"query benchmark\" height=\"300\"\u003e\u003c/img\u003e\n\u003c/p\u003e\n\nIn all benchmarked queries, Datalevin is the fastest among the three tested\nsystems, as Datalevin has a [cost based query optimizer](doc/query.md) while Datascript and\nDatomic do not. Datalevin also has a caching layer for index access. See\n[here](benchmarks/datascript-bench) for a detailed analysis of the results.\n\nWe also compared Datalevin and PostgreSQL in handling complex queries, using\n[Join Order Benchmark](benchmarks/JOB-bench). On a\nMacBook Pro, Apple M3 chip with 12 cores, 30 GB memory and 1TB SSD drive, here's\nthe average times:\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"benchmarks/JOB-bench/means.png\" alt=\"JOB benchmark averages\" height=\"300\"\u003e\u003c/img\u003e\n\u003c/p\u003e\n\nDatalevin is about 1.3X faster than PostgreSQL on average in running the complex\nqueries that involves many joins. The gain is mainly due to shorter query\nexecution time as Datalevin's query optimizer generates better plans. Details of\nthe analysis can be found in [this\narticle](https://yyhh.org/blog/2024/09/competing-for-the-job-with-a-triplestore/)\n\nFor durable transaction performance, we compared Datalevin with\nSQLite using [this write benchmark](benchmark/write-bench) on a 2016 Ubuntu Linux server with an Intel i7 3.6GHz CPU and a 1TB SSD drive.\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"benchmarks/write-bench/throughput-1.png\" alt=\"Throughput at 1\" height=\"300\"\u003e\u003c/img\u003e\n\u003c/p\u003e\n\nWhen transacting one entity (equivalently, one row in SQLite) at a time,\nDatalevin's default transaction function is over 5X faster than SQLite's\ndefault; while Datlevin's asynchronous transaction mode is over 20X faster than\nSQLite's WAL mode.\n\nOn the other hand, when transacting ever larger number of entities (rows) at a\ntime, SQLite gains on Datalevin and eventually surpasses it. For bulk loading\ndata, it is recommended to use `init-db` and `fill-db` functions, instead of\ndoing transactions in Datalevin. See [transaction](doc/transact.md) for more\ndiscussions.\n\n## :earth_americas: Roadmap\n\nThese are the tentative goals that we try to reach as soon as we can. We may\nadjust the priorities based on feedback.\n\n* 0.4.0 ~~Native image and native command line tool.~~ [Done 2021/02/27]\n* 0.5.0 ~~Networked server mode with role based access control.~~ [Done 2021/09/06]\n* 0.6.0 ~~As a search engine: full-text search across database.~~ [Done 2022/03/10]\n* 0.7.0 ~~Explicit transactions, lazy results loading, and results spill to disk\n  when memory is low.~~ [Done 2022/12/15]\n* 0.8.0 ~~Long ids; composite tuples; enhanced search engine ingestion speed.~~\n  [Done 2023/01/19]\n* 0.9.0 ~~New Datalog query engine with improved performance.~~ [Done 2024/03/09]\n* 0.10.0 ~~Async transaction; boolean search expression and phrase search; as a\n  vector database;~~ TTL for KV DB; extensible de/serialization for arbitrary data; auto upgrade migration; compressed data storage.\n* 1.0.0 New rule evaluation algorithm and incremental view maintenance.\n* 1.1.0 Transaction log storage and access API; read-only replicas for server.\n* 1.2.0 JSON API and library/client for popular languages.\n* 2.0.0 Automatic document indexing.\n* 3.0.0 Distributed mode.\n\n\n## :arrows_clockwise: Contact\n\nWe appreciate and welcome your contributions or suggestions. Please feel free to\nfile issues or pull requests.\n\nIf commercial support is needed for Datalevin, talk to us.\n\nYou can talk to us in the `#datalevin` channel on [Clojurians Slack](http://clojurians.net/).\n\n## License\n\nCopyright © 2020-2025 [Juji, Inc.](https://juji.io).\n\nLicensed under Eclipse Public License (see [LICENSE](LICENSE)).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjuji-io%2Fdatalevin","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjuji-io%2Fdatalevin","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjuji-io%2Fdatalevin/lists"}