{"id":14067795,"url":"https://github.com/russHyde/polyply","last_synced_at":"2025-07-30T02:31:27.968Z","repository":{"id":109720015,"uuid":"136173468","full_name":"russHyde/polyply","owner":"russHyde","description":"`polyply` allows you to manipulate multiple data-frames within a single magrittr / dplyr pipeline","archived":false,"fork":false,"pushed_at":"2019-07-12T15:22:15.000Z","size":85,"stargazers_count":8,"open_issues_count":7,"forks_count":2,"subscribers_count":1,"default_branch":"master","last_synced_at":"2025-04-20T20:19:59.246Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"R","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/russHyde.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2018-06-05T12:23:51.000Z","updated_at":"2024-06-04T11:59:30.000Z","dependencies_parsed_at":"2023-07-11T22:15:32.125Z","dependency_job_id":null,"html_url":"https://github.com/russHyde/polyply","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/russHyde/polyply","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/russHyde%2Fpolyply","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/russHyde%2Fpolyply/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/russHyde%2Fpolyply/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/russHyde%2Fpolyply/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/russHyde","download_url":"https://codeload.github.com/russHyde/polyply/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/russHyde%2Fpolyply/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":267798625,"owners_count":24145727,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-07-30T02:00:09.044Z","response_time":70,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-13T07:05:47.129Z","updated_at":"2025-07-30T02:31:27.738Z","avatar_url":"https://github.com/russHyde.png","language":"R","funding_links":[],"categories":["R"],"sub_categories":[],"readme":"# `polyply` - an R package for manipulation of multiple data-frames in a single magrittr pipeline\n\n[![Build Status](\n  https://travis-ci.org/russHyde/polyply.svg?branch=master\n  )\n](\n  https://travis-ci.org/russHyde/polyply\n)\n\n[![Coverage Status](\n  https://img.shields.io/codecov/c/github/russHyde/polyply/master.svg\n  )\n](\n  https://codecov.io/github/russHyde/polyply?branch=master\n)\n\nContributors are more than welcome - but you've got to be nice - see the\ncode-of-conduct (CONDUCT.md)\n\n----\n\nI've talked about these ideas on twitter and biostars recently.\n\n`dplyr` really shines in the manipulation of single data-frames and has\nfunctions for merging existing data-frames together, similar to how tables are\ncombined in relational databases. The `dplyr` syntax gets rather heavy-handed\nwhen you need to both manipulate and merge more than one data-frame in a single\npipeline.\n\nAn example might be illustrative:\n\nSuppose data-frames `A`, `B`, and `C` exist that contain information related to\na gene-expression experiment. Any number of other applications could have been\nchosen.\n\n`A` might contain annotation information for the genes that were studied (IDs\nin external databases, gene lengths etc). `B` might contain information about\nthe experimental samples that were studied (was the sample treated with a\nparticular treatment; where were the samples sourced etc). `C` might contain\nthe expression level for each gene in each sample.\n\nAssume the three datasets are 'tidy'. There's a single row for each gene in\n`A`, there's a single row for each sample in `B` and there's a single row for\neach gene/sample combination in `C`.\n\nSuppose I want to extract the expression information for all patients sourced\nfrom Glasgow, for genes that are longer than 2 kilobases. Then I want to pass\nthat expression data into ggplot2 and make a well-annotated scatter plot (the\nannotations using information from both the gene-metadata `A` and the\nbiological-samples-metadata `B`).\n\nThere's many ways to do this using standard tidyverse approaches.\n\n~~~~\n# prefilter, constructing superfluous data-frames\nglasgow_samples \u003c- filter(B, source == \"Glasgow\")\nlong_genes \u003c- filter(A, gene_length \u003e 2000)\nexpression_data \u003c- C %\u003e%\n  inner_join(glasgow_samples, by = \"sample_id\") %\u003e%\n  inner_join(long_genes, by = \"feature_id\") %\u003e%\n  ggplot(...)\n~~~~\n\n~~~~\n# filter within the join\nC %\u003e%\n  inner_join(filter(B, source == \"Glasgow\"), by = \"sample_id\") %\u003e%\n  inner_join(filter(A, gene_length \u003e 2000)) %\u003e%\n  ggplot(...)\n~~~~\n\n~~~~\n# post-filter, making a huge temporary data-frame\nC %\u003e%\n  inner_join(B, by = \"sample_id\") %\u003e%\n  inner_join(A, by = \"feature_id\") %\u003e%\n  filter(source == \"Glasgow\" \u0026 gene_length \u003e 2000) %\u003e%\n  ggplot(...)\n~~~~\n\nAll of the above are perfectly valid approaches.\n\nIf you do them once.\n\nAnd if your data-frames are sufficently small.\n\nBut 'doing things just once' always seems like the exception rather than the\nrule, and working with manageably-sized data-frames is another rarity.\n\nSo, to  mitigate against duplication, which bits of the above code should be\nabstracted away?\n\nSince the 'datasets used' will change less rapidly than the 'questions asked',\nI'm more likely to need to change the filters / selections / mutations that are\napplied to the individual data-frames than I am to change the pipeline for\njoining-together the different datasets. Hell, there's a logical connection\nbetween the different data-frames that is unaffected by filtering any given\ndata-frame - so surely joining on that logical connection should be abstracted\nout first.\n\nWith-respect-to-the-above, we want to define a single merging function that\nwill take the collestion of data-frames (assume a list of data-frames for now),\nand return a single data-frame for filtering / selection etc or for use in\n`ggplot`. This might look like:\n\n~~~~\nlist(genes = A, samples = B, expressions = C) %\u003e%\n  my_merging_function() %\u003e%\n  filter(source == \"Glasgow\" \u0026 gene_length \u003e 2000) %\u003e%\n  ggplot(...)\n~~~~\n\nHere, the merging function works similarly to how a View works in SQL - it's a\nvirtual specification for how data-tables should be combined together. But\nunlike in SQL, there's no optimisation performed in R and so that query would\ncreate the same huge temporary inner-join data-frame described above. Given\nthat, something like the following would be more memory efficient in R:\n\n~~~~\nlist(\n  genes = filter(A, ...),\n  samples = filter(B, ...),\n  expressions = C\n) %\u003e%\n  my_merging_function() %\u003e%\n  ggplot(...)\n~~~~\n\nBut but, what if you could do this:\n\n~~~~\nlist(\n  genes = A,\n  samples = B,\n  expressions = C\n) %\u003e%\n  filter_a_specific_dataframe_within_that_collection(...) %\u003e%\n  filter_a_different_dataframe(...) %\u003e%\n  my_merging_function() %\u003e%\n  ggplot(...)\n~~~~\n\n... and have it behave algebraically identically to the two previous calls.\nThis would limit memory burden, limit the need to store intermediate results\nbut also allow you store intermediate results etc.\n\nBut but but, doesn't tidygraph do something similar already? in tidygraph,\nthere are a couple of data-frames (one for edges and one for nodes) stored\ninside a tbl_graph object, and you can mutate / filter / etc each of these\ndata-frames independently. To indicate which of the tables you want to work\nwith, you use 'activate'. So we could generalise the code from tidygraph in a\nway that allows the following workflow:\n\n~~~~\nsome_collection(genes = A, samples = B, expressions = C) %\u003e%\n  activate(genes) %\u003e%\n    filter(gene_length \u003e 2000) %\u003e%\n  activate(samples) %\u003e%\n    filter(source == \"Glasgow\") %\u003e%\n  my_merging_function() %\u003e%\n  [... other filtering / mutation / selection steps ...] %\u003e%\n  [... downstream output maker ...]\n~~~~\n\nHappy to take criticism of the idea.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FrussHyde%2Fpolyply","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FrussHyde%2Fpolyply","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FrussHyde%2Fpolyply/lists"}