{"id":40699853,"url":"https://github.com/dssg/ohio","last_synced_at":"2026-01-21T12:02:59.997Z","repository":{"id":37601658,"uuid":"171942894","full_name":"dssg/ohio","owner":"dssg","description":"Python I/O extras","archived":false,"fork":false,"pushed_at":"2022-12-26T20:53:01.000Z","size":128,"stargazers_count":18,"open_issues_count":7,"forks_count":0,"subscribers_count":11,"default_branch":"master","last_synced_at":"2025-09-25T16:49:13.918Z","etag":null,"topics":["programming-utility"],"latest_commit_sha":null,"homepage":"https://ohio.readthedocs.io/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/dssg.png","metadata":{"files":{"readme":"README.rst","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2019-02-21T20:47:49.000Z","updated_at":"2025-03-02T22:25:08.000Z","dependencies_parsed_at":"2023-01-31T01:30:41.565Z","dependency_job_id":null,"html_url":"https://github.com/dssg/ohio","commit_stats":null,"previous_names":[],"tags_count":9,"template":false,"template_full_name":null,"purl":"pkg:github/dssg/ohio","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dssg%2Fohio","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dssg%2Fohio/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dssg%2Fohio/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dssg%2Fohio/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/dssg","download_url":"https://codeload.github.com/dssg/ohio/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dssg%2Fohio/sbom","scorecard":{"id":358024,"data":{"date":"2025-08-11","repo":{"name":"github.com/dssg/ohio","commit":"c0d992287a44067b266c97dafadc6d1ffeb09096"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":3,"checks":[{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Code-Review","score":2,"reason":"Found 6/29 approved changesets -- score normalized to 2","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Token-Permissions","score":-1,"reason":"No tokens found","details":null,"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Dangerous-Workflow","score":-1,"reason":"no workflows found","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: downloadThenRun not pinned by hash: develop:73","Info:   0 out of   1 downloadThenRun dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Vulnerabilities","score":10,"reason":"0 existing vulnerabilities detected","details":null,"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}},{"name":"License","score":9,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Warn: project license file does not contain an FSF or OSI license."],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'master'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 7 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}}]},"last_synced_at":"2025-08-18T10:06:07.605Z","repository_id":37601658,"created_at":"2025-08-18T10:06:07.606Z","updated_at":"2025-08-18T10:06:07.606Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28632781,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-21T04:47:28.174Z","status":"ssl_error","status_checked_at":"2026-01-21T04:47:22.943Z","response_time":86,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["programming-utility"],"created_at":"2026-01-21T12:02:51.060Z","updated_at":"2026-01-21T12:02:59.989Z","avatar_url":"https://github.com/dssg.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\nOhio\n****\n\nOh! IO: The I/O tools that ``io`` doesn’t want you to have.\n\nOhio provides the missing links between Python’s built-in I/O\nprimitives, to help ensure the efficiency, clarity and elegance of\nyour code.\n\nFor higher-level examples of what Ohio can do for you, see\n`Extensions`_ and `Recipes`_.\n\n\nContents\n^^^^^^^^\n\n* `Ohio`_\n\n   * `Installation`_\n\n   * `Modules`_\n\n      * `csvio`_\n\n      * `iterio`_\n\n      * `pipeio`_\n\n      * `baseio`_\n\n      * `Extensions`_\n\n         * `Extensions for NumPy`_\n\n         * `Extensions for Pandas`_\n\n         * `Benchmarking`_\n\n      * `Recipes`_\n\n         * `dbjoin`_\n\n\nInstallation\n============\n\nOhio is a distributed library with support for Python v3. It is\navailable from `pypi.org \u003chttps://pypi.org/project/ohio/\u003e`_:\n\n::\n\n   $ pip install ohio\n\n\nModules\n=======\n\n\ncsvio\n-----\n\nFlexibly encode data to CSV format.\n\n**ohio.encode_csv(rows, *writer_args, writer=\u003cbuilt-in function\nwriter\u003e, write_header=False, **writer_kwargs)**\n\n   Encode the specified iterable of ``rows`` into CSV text.\n\n   Data is encoded to an in-memory ``str``, (rather than to the file\n   system), via an internally-managed ``io.StringIO``, (newly\n   constructed for every invocation of ``encode_csv``).\n\n   For example:\n\n   ::\n\n      \u003e\u003e\u003e data = [\n      ...     ('1/2/09 6:17', 'Product1', '1200', 'Mastercard', 'carolina'),\n      ...     ('1/2/09 4:53', 'Product1', '1200', 'Visa', 'Betina'),\n      ... ]\n\n      \u003e\u003e\u003e encoded_csv = encode_csv(data)\n\n      \u003e\u003e\u003e encoded_csv[:80]\n      '1/2/09 6:17,Product1,1200,Mastercard,carolina\\r\\n1/2/09 4:53,Product1,1200,Visa,Be'\n\n      \u003e\u003e\u003e encoded_csv.splitlines(keepends=True)\n      ['1/2/09 6:17,Product1,1200,Mastercard,carolina\\r\\n',\n       '1/2/09 4:53,Product1,1200,Visa,Betina\\r\\n']\n\n   By default, ``rows`` are encoded by built-in ``csv.writer``. You\n   may specify an alternate ``writer``, and provide construction\n   arguments:\n\n   ::\n\n      \u003e\u003e\u003e header = ('Transaction_date', 'Product', 'Price', 'Payment_Type', 'Name')\n\n      \u003e\u003e\u003e data = [\n      ...     {'Transaction_date': '1/2/09 6:17',\n      ...      'Product': 'Product1',\n      ...      'Price': '1200',\n      ...      'Payment_Type': 'Mastercard',\n      ...      'Name': 'carolina'},\n      ...     {'Transaction_date': '1/2/09 4:53',\n      ...      'Product': 'Product1',\n      ...      'Price': '1200',\n      ...      'Payment_Type': 'Visa',\n      ...      'Name': 'Betina'},\n      ... ]\n\n      \u003e\u003e\u003e encoded_csv = encode_csv(data, writer=csv.DictWriter, fieldnames=header)\n\n      \u003e\u003e\u003e encoded_csv.splitlines(keepends=True)\n      ['1/2/09 6:17,Product1,1200,Mastercard,carolina\\r\\n',\n       '1/2/09 4:53,Product1,1200,Visa,Betina\\r\\n']\n\n   And, for such writers featuring the method ``writeheader``, you may\n   instruct ``encode_csv`` to invoke this, prior to writing ``rows``:\n\n   ::\n\n      \u003e\u003e\u003e encoded_csv = encode_csv(\n      ...     data,\n      ...     writer=csv.DictWriter,\n      ...     fieldnames=header,\n      ...     write_header=True,\n      ... )\n\n      \u003e\u003e\u003e encoded_csv.splitlines(keepends=True)\n      ['Transaction_date,Product,Price,Payment_Type,Name\\r\\n',\n       '1/2/09 6:17,Product1,1200,Mastercard,carolina\\r\\n',\n       '1/2/09 4:53,Product1,1200,Visa,Betina\\r\\n']\n\n**class ohio.CsvTextIO(rows, *writer_args, write_header=False,\nchunk_size=10, **writer_kwargs)**\n\n   Readable file-like interface encoding specified data as CSV.\n\n   Rows of input data are only consumed and encoded as needed, as\n   ``CsvTextIO`` is read.\n\n   Rather than write to the file system, an internal ``io.StringIO``\n   buffer is used to store output temporarily, until it is read. (Also\n   unlike ``ohio.encode_csv``, this buffer is reused across read/write\n   cycles.)\n\n   For example, we might encode the following data as CSV:\n\n   ::\n\n      \u003e\u003e\u003e data = [\n      ...     ('1/2/09 6:17', 'Product1', '1200', 'Mastercard', 'carolina'),\n      ...     ('1/2/09 4:53', 'Product1', '1200', 'Visa', 'Betina'),\n      ... ]\n\n      \u003e\u003e\u003e csv_buffer = CsvTextIO(data)\n\n   Data may be encoded and retrieved via standard file object methods,\n   such as ``read``, ``readline`` and iteration:\n\n   ::\n\n      \u003e\u003e\u003e csv_buffer.read(15)\n      '1/2/09 6:17,Pro'\n\n      \u003e\u003e\u003e next(csv_buffer)\n      'duct1,1200,Mastercard,carolina\\r\\n'\n\n      \u003e\u003e\u003e list(csv_buffer)\n      ['1/2/09 4:53,Product1,1200,Visa,Betina\\r\\n']\n\n      \u003e\u003e\u003e csv_buffer.read()\n      ''\n\n   Note, in the above example, we first read 15 bytes of the encoded\n   CSV, then read the remainder of the line via iteration, (which\n   invokes ``readline``), and then collected the remaining CSV into a\n   list. Finally, we attempted to read the entirety still remaining –\n   which was nothing.\n\n**class ohio.CsvDictTextIO(rows, *writer_args, write_header=False,\nchunk_size=10, **writer_kwargs)**\n\n   ``CsvTextIO`` which accepts row data in the form of ``dict``.\n\n   Data is passed to ``csv.DictWriter``.\n\n   See also: ``ohio.CsvTextIO``.\n\n**ohio.iter_csv(rows, *writer_args, write_header=False,\n**writer_kwargs)**\n\n   Generate lines of encoded CSV from ``rows`` of data.\n\n   See: ``ohio.CsvWriterTextIO``.\n\n**ohio.iter_dict_csv(rows, *writer_args, write_header=False,\n**writer_kwargs)**\n\n   Generate lines of encoded CSV from ``rows`` of data.\n\n   See: ``ohio.CsvWriterTextIO``.\n\n**class ohio.CsvWriterTextIO(*writer_args, **writer_kwargs)**\n\n   csv.writer-compatible interface to iteratively encode CSV in\n   memory.\n\n   The writer instance may also be read, to retrieve written CSV, as\n   it is written.\n\n   Rather than write to the file system, an internal ``io.StringIO``\n   buffer is used to store output temporarily, until it is read.\n   (Unlike ``ohio.encode_csv``, this buffer is reused across\n   read/write cycles.)\n\n   Features class method ``iter_csv``: a generator to map an input\n   iterable of data ``rows`` to lines of encoded CSV text.\n   (``iter_csv`` differs from ``ohio.encode_csv`` in that it lazily\n   generates lines of CSV, rather than eagerly encoding the entire CSV\n   body.)\n\n   **Note**: If you don’t need to control *how* rows are written, but\n   do want an iterative and/or readable interface to encoded CSV,\n   consider also the more straight-forward ``ohio.CsvTextIO``.\n\n   For example, we may construct ``CsvWriterTextIO`` with the same\n   (optional) arguments as we would ``csv.writer``, (minus the file\n   descriptor):\n\n   ::\n\n      \u003e\u003e\u003e csv_buffer = CsvWriterTextIO(dialect='excel')\n\n   …and write to it, via either ``writerow`` or ``writerows``:\n\n   ::\n\n      \u003e\u003e\u003e csv_buffer.writerows([\n      ...     ('1/2/09 6:17', 'Product1', '1200', 'Mastercard', 'carolina'),\n      ...     ('1/2/09 4:53', 'Product1', '1200', 'Visa', 'Betina'),\n      ... ])\n\n   Written data is then available to be read, via standard file object\n   methods, such as ``read``, ``readline`` and iteration:\n\n   ::\n\n      \u003e\u003e\u003e csv_buffer.read(15)\n      '1/2/09 6:17,Pro'\n\n      \u003e\u003e\u003e list(csv_buffer)\n      ['duct1,1200,Mastercard,carolina\\r\\n',\n       '1/2/09 4:53,Product1,1200,Visa,Betina\\r\\n']\n\n   Note, in the above example, we first read 15 bytes of the encoded\n   CSV, and then collected the remaining CSV into a list, through\n   iteration, (which returns its lines, via ``readline``). However,\n   the first line was short by that first 15 bytes.\n\n   That is, reading CSV out of the ``CsvWriterTextIO`` empties that\n   content from its buffer:\n\n   ::\n\n      \u003e\u003e\u003e csv_buffer.read()\n      ''\n\n   We can repopulate our ``CsvWriterTextIO`` buffer by writing to it\n   again:\n\n   ::\n\n      \u003e\u003e\u003e csv_buffer.writerows([\n      ...     ('1/2/09 13:08', 'Product1', '1200', 'Mastercard', 'Federica e Andrea'),\n      ...     ('1/3/09 14:44', 'Product1', '1200', 'Visa', 'Gouya'),\n      ... ])\n\n      \u003e\u003e\u003e encoded_csv = csv_buffer.read()\n\n      \u003e\u003e\u003e encoded_csv[:80]\n      '1/2/09 13:08,Product1,1200,Mastercard,Federica e Andrea\\r\\n1/3/09 14:44,Product1,1'\n\n      \u003e\u003e\u003e encoded_csv.splitlines(keepends=True)\n      ['1/2/09 13:08,Product1,1200,Mastercard,Federica e Andrea\\r\\n',\n       '1/3/09 14:44,Product1,1200,Visa,Gouya\\r\\n']\n\n   Finally, class method ``iter_csv`` can do all this for us,\n   generating lines of encoded CSV as we request them:\n\n   ::\n\n      \u003e\u003e\u003e lines_csv = CsvWriterTextIO.iter_csv([\n      ...     ('Transaction_date', 'Product', 'Price', 'Payment_Type', 'Name'),\n      ...     ('1/2/09 6:17', 'Product1', '1200', 'Mastercard', 'carolina'),\n      ...     ('1/2/09 4:53', 'Product1', '1200', 'Visa', 'Betina'),\n      ...     ('1/2/09 13:08', 'Product1', '1200', 'Mastercard', 'Federica e Andrea'),\n      ...     ('1/3/09 14:44', 'Product1', '1200', 'Visa', 'Gouya'),\n      ... ])\n\n      \u003e\u003e\u003e next(lines_csv)\n      'Transaction_date,Product,Price,Payment_Type,Name\\r\\n'\n\n      \u003e\u003e\u003e next(lines_csv)\n      '1/2/09 6:17,Product1,1200,Mastercard,carolina\\r\\n'\n\n      \u003e\u003e\u003e list(lines_csv)\n      ['1/2/09 4:53,Product1,1200,Visa,Betina\\r\\n',\n       '1/2/09 13:08,Product1,1200,Mastercard,Federica e Andrea\\r\\n',\n       '1/3/09 14:44,Product1,1200,Visa,Gouya\\r\\n']\n\n**class ohio.CsvDictWriterTextIO(*writer_args, **writer_kwargs)**\n\n   ``CsvWriterTextIO`` which accepts row data in the form of ``dict``.\n\n   Data is passed to ``csv.DictWriter``.\n\n   See also: ``ohio.CsvWriterTextIO``.\n\n\niterio\n------\n\nProvide a readable file-like interface to any iterable.\n\n**class ohio.IteratorTextIO(iterable)**\n\n   Readable file-like interface for iterable text streams.\n\n   ``IteratorTextIO`` wraps any iterable of text for consumption like\n   a file, offering methods ``readline()``, ``read([size])``, *etc.*,\n   (implemented via base class ``ohio.StreamTextIOBase``).\n\n   For example, given a consumer which expects to ``read()``:\n\n   ::\n\n      \u003e\u003e\u003e def read_chunks(fdesc, chunk_size=1024):\n      ...     get_chunk = lambda: fdesc.read(chunk_size)\n      ...     yield from iter(get_chunk, '')\n\n   …And either streamed or in-memory text (*i.e.* which is not simply\n   on a file system):\n\n   ::\n\n      \u003e\u003e\u003e def all_caps(fdesc):\n      ...     for line in fdesc:\n      ...         yield line.upper()\n\n   …We can connect these two interfaces via ``IteratorTextIO``:\n\n   ::\n\n      \u003e\u003e\u003e with open('/usr/share/dict/words') as fdesc:\n      ...     louder_words_lines = all_caps(fdesc)\n      ...     with IteratorTextIO(louder_words_lines) as louder_words_desc:\n      ...         louder_words_chunked = read_chunks(louder_words_desc)\n\n\npipeio\n------\n\nEfficiently connect ``read()`` and ``write()`` interfaces.\n\n``PipeTextIO`` provides a *readable* and iterable interface to text\nwhose producer requires a *writable* interface.\n\nIn contrast to first writing such text to memory and then consuming\nit, ``PipeTextIO`` only allows write operations as necessary to fill\nits buffer, to fulfill read operations, asynchronously. As such,\n``PipeTextIO`` consumes a stable minimum of memory, and may\nsignificantly boost speed, with a minimum of boilerplate.\n\n**ohio.pipe_text(writer_func, *args, buffer_size=None, **kwargs)**\n\n   Iteratively stream output written by given function through\n   readable file-like interface.\n\n   Uses in-process writer thread, (which runs the given function), to\n   mimic buffered text transfer, such as between the standard output\n   and input of two piped processes.\n\n   Calls to ``write`` are blocked until required by calls to ``read``.\n\n   Note: If at all possible, use a generator! Your iterative text-\n   writing function can most likely be designed as a generator, (or as\n   some sort of iterator). Its output can then, far more simply and\n   easily, be streamed to some input. If your input must be ``read``\n   from a file-like object, see ``ohio.IteratorTextIO``. If your\n   output must be CSV-encoded, see ``ohio.encode_csv``,\n   ``ohio.CsvTextIO`` and ``ohio.CsvWriterTextIO``.\n\n   ``PipeTextIO`` is suitable for situations where output *must* be\n   written to a file-like object, which is made blocking to enforce\n   iterativity.\n\n   ``PipeTextIO`` is not “seekable,” but supports all other typical,\n   read-write file-like features.\n\n   For example, consider the following callable, (artificially)\n   requiring a file-like object, to which to write:\n\n   ::\n\n      \u003e\u003e\u003e def write_output(file_like):\n      ...     file_like.write(\"Hi there.\\r\\n\")\n      ...     print('[writer]', 'Yay I wrote one line')\n      ...     file_like.write(\"Cool, right?\\r\\n\")\n      ...     print('[writer]', 'Finally ... I wrote a second line!')\n      ...     file_like.write(\"All right, later :-)\\r\\n\")\n      ...     print('[writer]', \"Done.\")\n\n   Most typically, we might *read* this content as follows, using\n   either the ``PipeTextIO`` constructor or its ``pipe_text`` helper:\n\n   ::\n\n      \u003e\u003e\u003e with PipeTextIO(write_output) as pipe:\n      ...     for line in pipe:\n      ...         ...\n\n   And, this syntax is recommended. However, for the sake of example,\n   consider the following:\n\n   ::\n\n      \u003e\u003e\u003e pipe = PipeTextIO(write_output, buffer_size=1)\n\n      \u003e\u003e\u003e pipe.read(5)\n      [writer] Yay I wrote one line\n      'Hi th'\n      [writer] Finally ... I wrote a second line!\n\n      \u003e\u003e\u003e pipe.readline()\n      'ere.\\r\\n'\n\n      \u003e\u003e\u003e pipe.readline()\n      'Cool, right?\\r\\n'\n      [writer] Done.\n\n      \u003e\u003e\u003e pipe.read()\n      'All right, later :-)\\r\\n'\n\n   In the above example, ``write_output`` requires a file-like\n   interface to which to write its output; (and, we presume that there\n   is no alternative to this implementation – such as a generator –\n   that its output is large enough that we don’t want to hold it in\n   memory **and** that we don’t need this output written to the file\n   system). We are enabled to read it directly, in chunks:\n\n   ..\n\n      1. Initially, nothing is written.\n\n      2. 1. Upon requesting to read – in this case, only the first 5\n              bytes – the writer is initialized, and permitted to\n              write its first chunk, (which happens to be one full\n              line). This is retrieved from the write buffer, and\n              sufficient to satisfy the read request.\n\n          2. Having removed the first chunk from the write buffer,\n              the writer is permitted to eagerly write its next chunk,\n              (the second line), (but, no more than that).\n\n      3. The second read request – for the remainder of the line – is\n          fully satisfied by the first chunk retrieved from the write\n          buffer. No more writing takes place.\n\n      4. The third read request, for another line, retrieves the\n          second chunk from the write buffer. The writer is permitted\n          to write its final chunk to the write buffer.\n\n      5. The final read request returns all remaining text,\n          (retrieved from the write buffer).\n\n   Concretely, this is commonly useful with the PostgreSQL COPY\n   command, for efficient data transfer, (and without the added\n   complexity of the file system). While your database interface may\n   vary, ``PipeTextIO`` enables the following syntax, for example to\n   copy data into the database:\n\n   ::\n\n      \u003e\u003e\u003e def write_csv(file_like):\n      ...     writer = csv.writer(file_like)\n      ...     ...\n\n      \u003e\u003e\u003e with PipeTextIO(write_csv) as pipe, \\\n      ...      connection.cursor() as cursor:\n      ...     cursor.copy_from(pipe, 'my_table', format='csv')\n\n   …or, to copy data out of the database:\n\n   ::\n\n      \u003e\u003e\u003e with connection.cursor() as cursor:\n      ...     writer = lambda pipe: cursor.copy_to(pipe,\n      ...                                          'my_table',\n      ...                                          format='csv')\n      ...\n      ...     with PipeTextIO(writer) as pipe:\n      ...         reader = csv.reader(pipe)\n      ...         ...\n\n   Alternatively, writer arguments may be passed to ``PipeTextIO``:\n\n   ::\n\n      \u003e\u003e\u003e with connection.cursor() as cursor:\n      ...     with PipeTextIO(cursor.copy_to,\n      ...                     args=['my_table'],\n      ...                     kwargs={'format': 'csv'}) as pipe:\n      ...         reader = csv.reader(pipe)\n      ...         ...\n\n   (But, bear in mind, the signature of the callable passed to\n   ``PipeTextIO`` must be such that its first, anonymous argument is\n   the ``PipeTextIO`` instance.)\n\n   Consider also the above example with the helper ``pipe_text``:\n\n   ::\n\n      \u003e\u003e\u003e with connection.cursor() as cursor:\n      ...     with pipe_text(cursor.copy_to,\n      ...                    'my_table',\n      ...                    format='csv') as pipe:\n      ...         reader = csv.reader(pipe)\n      ...         ...\n\n   Finally, note that copying *to* the database is likely best\n   performed via ``ohio.CsvTextIO``, (though copying *from* requires\n   ``PipeTextIO``, as above):\n\n   ::\n\n      \u003e\u003e\u003e with ohio.CsvTextIO(data_rows) as csv_buffer, \\\n      ...      connection.cursor() as cursor:\n      ...     cursor.copy_from(csv_buffer, 'my_table', format='csv')\n\n\nbaseio\n------\n\nLow-level primitives.\n\n**class ohio.StreamTextIOBase**\n\n   Readable file-like abstract base class.\n\n   Concrete classes must implement method ``__next_chunk__`` to return\n   chunk(s) of the text to be read.\n\n**exception ohio.IOClosed(*args)**\n\n   Exception indicating an attempted operation on a file-like object\n   which has been closed.\n\n.. _extensions:\n\n\nExtensions\n----------\n\nModules integrating Ohio with the toolsets that need it.\n\n\nExtensions for NumPy\n~~~~~~~~~~~~~~~~~~~~\n\nThis module enables writing NumPy array data to database and\npopulating arrays from database via PostgreSQL ``COPY``. The operation\nis ensured, by Ohio, to be memory-efficient.\n\n**Note**: This integration is intended for NumPy, and attempts to\n``import numpy``. NumPy must be available (installed) in your\nenvironment.\n\n**ohio.ext.numpy.pg_copy_to_table(arr, table_name, connectable,\ncolumns=None, fmt=None)**\n\n   Copy ``array`` to database table via PostgreSQL ``COPY``.\n\n   ``ohio.PipeTextIO`` enables the direct, in-process “piping” of\n   ``array`` CSV into the “standard input” of the PostgreSQL ``COPY``\n   command, for quick, memory-efficient database persistence, (and\n   without the needless involvement of the local file system).\n\n   For example, given a SQLAlchemy ``connectable`` – either a database\n   connection ``Engine`` or ``Connection`` – and a NumPy ``array``:\n\n   ::\n\n      \u003e\u003e\u003e from sqlalchemy import create_engine\n      \u003e\u003e\u003e engine = create_engine('postgresql://')\n\n      \u003e\u003e\u003e arr = numpy.array([1.000102487, 5.982, 2.901, 103.929])\n\n   We may persist this data to an existing table – *e.g.* “data”:\n\n   ::\n\n      \u003e\u003e\u003e pg_copy_to_table(arr, 'data', engine, columns=['value'])\n\n   ``pg_copy_to_table`` utilizes ``numpy.savetxt`` and supports its\n   ``fmt`` parameter.\n\n**ohio.ext.numpy.pg_copy_from_table(table_name, connectable, dtype,\ncolumns=None)**\n\n   Construct ``array`` from database table via PostgreSQL ``COPY``.\n\n   ``ohio.PipeTextIO`` enables the in-process “piping” of the\n   PostgreSQL ``COPY`` command into NumPy’s ``fromiter``, for quick,\n   memory-efficient construction of ``array`` from database, (and\n   without the needless involvement of the local file system).\n\n   For example, given a SQLAlchemy ``connectable`` – either a database\n   connection ``Engine`` or ``Connection``:\n\n   ::\n\n      \u003e\u003e\u003e from sqlalchemy import create_engine\n      \u003e\u003e\u003e engine = create_engine('postgresql://')\n\n   We may construct a NumPy ``array`` from the contents of a specified\n   table:\n\n   ::\n\n      \u003e\u003e\u003e arr = pg_copy_from_table(\n      ...     'data',\n      ...     engine,\n      ...     float,\n      ... )\n\n**ohio.ext.numpy.pg_copy_from_query(query, connectable, dtype)**\n\n   Construct ``array`` from database ``query`` via PostgreSQL\n   ``COPY``.\n\n   ``ohio.PipeTextIO`` enables the in-process “piping” of the\n   PostgreSQL ``COPY`` command into NumPy’s ``fromiter``, for quick,\n   memory-efficient construction of ``array`` from database, (and\n   without the needless involvement of the local file system).\n\n   For example, given a SQLAlchemy ``connectable`` – either a database\n   connection ``Engine`` or ``Connection``:\n\n   ::\n\n      \u003e\u003e\u003e from sqlalchemy import create_engine\n      \u003e\u003e\u003e engine = create_engine('postgresql://')\n\n   We may construct a NumPy ``array`` from a given query:\n\n   ::\n\n      \u003e\u003e\u003e arr = pg_copy_from_query(\n      ...     'select value0, value1, value3 from data',\n      ...     engine,\n      ...     float,\n      ... )\n\n\nExtensions for Pandas\n~~~~~~~~~~~~~~~~~~~~~\n\nThis module extends ``pandas.DataFrame`` with methods ``pg_copy_to``\nand ``pg_copy_from``.\n\nTo enable, simply import this module anywhere in your project, (most\nlikely – just once, in its root module):\n\n::\n\n   \u003e\u003e\u003e import ohio.ext.pandas\n\nFor example, if you have just one module – in there – or, in a Python\npackage:\n\n::\n\n   ohio/\n       __init__.py\n       baseio.py\n       ...\n\nthen in its ``__init__.py``, to ensure that extensions are loaded\nbefore your code, which uses them, is run.\n\n**Note**: These extensions are intended for Pandas, and attempt to\n``import pandas``. Pandas must be available (installed) in your\nenvironment.\n\n**class ohio.ext.pandas.DataFramePgCopyTo(data_frame)**\n\n   ``pg_copy_to``: Copy ``DataFrame`` to database table via PostgreSQL\n   ``COPY``.\n\n   ``ohio.CsvTextIO`` enables the direct reading of ``DataFrame`` CSV\n   into the “standard input” of the PostgreSQL ``COPY`` command, for\n   quick, memory-efficient database persistence, (and without the\n   needless involvement of the local file system).\n\n   For example, given a SQLAlchemy ``connectable`` – either a database\n   connection ``Engine`` or ``Connection`` – and a Pandas\n   ``DataFrame``:\n\n   ::\n\n      \u003e\u003e\u003e from sqlalchemy import create_engine\n      \u003e\u003e\u003e engine = create_engine('postgresql://')\n\n      \u003e\u003e\u003e df = pandas.DataFrame({'name' : ['User 1', 'User 2', 'User 3']})\n\n   We may simply invoke the ``DataFrame``’s Ohio extension method,\n   ``pg_copy_to``:\n\n   ::\n\n      \u003e\u003e\u003e df.pg_copy_to('users', engine)\n\n   ``pg_copy_to`` supports all the same parameters as ``to_sql``,\n   (excepting parameter ``method``).\n\n**ohio.ext.pandas.to_sql_method_pg_copy_to(table, conn, keys,\ndata_iter)**\n\n   Write pandas data to table via stream through PostgreSQL ``COPY``.\n\n   This implements a pandas ``to_sql`` “method”, utilizing\n   ``ohio.CsvTextIO`` for performance stability.\n\n**ohio.ext.pandas.data_frame_pg_copy_from(sql, connectable,\nschema=None, index_col=None, parse_dates=False, columns=None,\ndtype=None, nrows=None, buffer_size=100)**\n\n   ``pg_copy_from``: Construct ``DataFrame`` from database table or\n   query via PostgreSQL ``COPY``.\n\n   ``ohio.PipeTextIO`` enables the direct, in-process “piping” of the\n   PostgreSQL ``COPY`` command into Pandas ``read_csv``, for quick,\n   memory-efficient construction of ``DataFrame`` from database, (and\n   without the needless involvement of the local file system).\n\n   For example, given a SQLAlchemy ``connectable`` – either a database\n   connection ``Engine`` or ``Connection``:\n\n   ::\n\n      \u003e\u003e\u003e from sqlalchemy import create_engine\n      \u003e\u003e\u003e engine = create_engine('postgresql://')\n\n   We may simply invoke the ``DataFrame``’s Ohio extension method,\n   ``pg_copy_from``:\n\n   ::\n\n      \u003e\u003e\u003e df = DataFrame.pg_copy_from('users', engine)\n\n   ``pg_copy_from`` supports many of the same parameters as\n   ``read_sql`` and ``read_csv``.\n\n   In addition, ``pg_copy_from`` accepts the optimization parameter\n   ``buffer_size``, which controls the maximum number of CSV-encoded\n   results written by the database cursor to hold in memory prior to\n   their being read into the ``DataFrame``. Depending on use-case,\n   increasing this value may speed up the operation, at the cost of\n   additional memory – and vice-versa. ``buffer_size`` defaults to\n   ``100``.\n\n\nBenchmarking\n~~~~~~~~~~~~\n\nOhio extensions for pandas were benchmarked to test their speed and\nmemory-efficiency relative both to pandas built-in functionality and\nto custom implementations which do not utilize Ohio.\n\nInterfaces and syntactical niceties aside, Ohio generally features\nmemory stability. Its tools enable pipelines which may also improve\nspeed, (and which do so in standard use-cases).\n\nIn the below benchmark, Ohio extensions ``pg_copy_from`` \u0026\n``pg_copy_to`` reduced memory consumption by 84% \u0026 61%, and completed\nin 39% \u0026 91% less time, relative to pandas built-ins ``read_sql`` \u0026\n``to_sql``, (respectively).\n\nCompared to purpose-built extensions – which utilized PostgreSQL\n``COPY``, but using ``io.StringIO`` in place of ``ohio.PipeTextIO``\nand ``ohio.CsvTextIO`` – ``pg_copy_from`` \u0026 ``pg_copy_to`` also\nreduced memory consumption by 60% \u0026 32%, respectively.\n``pg_copy_from`` \u0026 ``pg_copy_to`` also completed in 16% \u0026 13% less\ntime than the ``io.StringIO`` versions.\n\nThe benchmarks plotted below were produced from averages and standard\ndeviations over 3 randomized trials per target. Input data consisted\nof 896,677 rows across 83 columns: 1 of these of type timestamp, 51\nintegers and 31 floats. The benchmarking package, ``prof``, is\npreserved in `Ohio's repository \u003chttps://github.com/dssg/ohio\u003e`_.\n\n.. image:: https://raw.githubusercontent.com/dssg/ohio/0.6.0/doc/img/profile-copy-from-database-to-datafram-1554345457.svg?sanitize=true\n\nohio_pg_copy_from_X\n   ``pg_copy_from(buffer_size=X)``\n\n   A PostgreSQL database-connected cursor writes the results of\n   ``COPY`` to a ``PipeTextIO``, from which pandas constructs a\n   ``DataFrame``.\n\npandas_read_sql\n   ``pandas.read_sql()``\n\n   Pandas constructs a ``DataFrame`` from a given database query.\n\npandas_read_sql_chunks_100\n   ``pandas.read_sql(chunksize=100)``\n\n   Pandas is instructed to generate ``DataFrame`` slices of the\n   database query result, and these slices are concatenated into a\n   single frame, with: ``pandas.concat(chunks, copy=False)``.\n\npandas_read_csv_stringio\n   ``pandas.read_csv(StringIO())``\n\n   A PostgreSQL database-connected cursor writes the results of\n   ``COPY`` to a ``StringIO``, from which pandas constructs a\n   ``DataFrame``.\n\n.. image:: https://raw.githubusercontent.com/dssg/ohio/0.6.0/doc/img/profile-copy-from-dataframe-to-databas-1555458507.svg?sanitize=true\n\nohio_pg_copy_to\n   ``pg_copy_to()``\n\n   ``DataFrame`` data are encoded through a ``CsvTextIO``, and read by\n   a PostgreSQL database-connected cursor’s ``COPY`` command.\n\npandas_to_sql\n   ``pandas.DataFrame.to_sql()``\n\n   Pandas inserts ``DataFrame`` data into the database row by row.\n\npandas_to_sql_multi_100\n   ``pandas.DataFrame.to_sql(method='multi', chunksize=100)``\n\n   Pandas inserts ``DataFrame`` data into the database in chunks of\n   rows.\n\ncopy_stringio_to_db\n   ``DataFrame`` data are written and encoded to a ``StringIO``, and\n   then read by a PostgreSQL database-connected cursor’s ``COPY``\n   command.\n\n.. _recipes:\n\n\nRecipes\n-------\n\nStand-alone modules implementing functionality which depends upon Ohio\nprimitives.\n\n\ndbjoin\n~~~~~~\n\nJoin the “COPY” results of arbitrary database queries in Python,\nwithout unnecessary memory overhead.\n\nThis is largely useful to work around databases’ per-query column\nlimit.\n\n**ohio.recipe.dbjoin.pg_join_queries(queries, engine, sep=', ',\nend='\\n', copy_options=('CSV', 'HEADER'))**\n\n   Join the text-encoded result streams of an arbitrary number of\n   PostgreSQL database queries to work around the database’s per-query\n   column limit.\n\n   Query results are read via PostgreSQL ``COPY``, streamed through\n   ``PipeTextIO``, and joined line-by-line into a singular stream.\n\n   For example, given a set of database queries whose results cannot\n   be combined into a single PostgreSQL query, we might join these\n   queries’ results and write these results to a file-like object:\n\n   ::\n\n      \u003e\u003e\u003e queries = [\n      ...     'SELECT a, b, c FROM a_table',\n      ...     ...\n      ... ]\n\n      \u003e\u003e\u003e with open('results.csv', 'w', newline='') as fdesc:\n      ...     for line in pg_join_queries(queries, engine):\n      ...         fdesc.write(line)\n\n   Or, we might read these results into a single Pandas DataFrame:\n\n   ::\n\n      \u003e\u003e\u003e csv_lines = pg_join_queries(queries, engine)\n      \u003e\u003e\u003e csv_buffer = ohio.IteratorTextIO(csv_lines)\n      \u003e\u003e\u003e df = pandas.read_csv(csv_buffer)\n\n   By default, ``pg_join_queries`` requests CSV-encoded results, with\n   an initial header line indicating the result columns. These\n   options, which are sent directly to the PostgreSQL ``COPY``\n   command, may be controlled via ``copy_options``. For example, to\n   omit the CSV header:\n\n   ::\n\n      \u003e\u003e\u003e pg_join_queries(queries, engine, copy_options=['CSV'])\n\n   Or, to request PostgreSQL’s tab-delimited text format via the\n   syntax of PostgreSQL v9.0+:\n\n   ::\n\n      \u003e\u003e\u003e pg_join_queries(\n      ...     queries,\n      ...     engine,\n      ...     sep='\\t',\n      ...     copy_options={'FORMAT': 'TEXT'},\n      ... )\n\n   In the above example, we’ve instructed PostgreSQL to use its\n   ``text`` results encoder, (and we’ve omitted the instruction to\n   include a header).\n\n   **NOTE**: In the last example, we also explicitly specified the\n   separator used in the results’ encoding. This is not passed to the\n   database; rather, it is necessary for ``pg_join_queries`` to\n   properly join queries’ results.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdssg%2Fohio","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdssg%2Fohio","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdssg%2Fohio/lists"}