{"id":16707412,"url":"https://github.com/xiaodaigh/data-wrangling-puzzles","last_synced_at":"2026-02-10T08:30:52.931Z","repository":{"id":73048501,"uuid":"285134730","full_name":"xiaodaigh/data-wrangling-puzzles","owner":"xiaodaigh","description":null,"archived":false,"fork":false,"pushed_at":"2020-09-23T00:29:34.000Z","size":30,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-08-29T05:33:45.159Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/xiaodaigh.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2020-08-05T00:30:15.000Z","updated_at":"2020-09-23T00:29:36.000Z","dependencies_parsed_at":"2023-03-13T20:18:48.101Z","dependency_job_id":null,"html_url":"https://github.com/xiaodaigh/data-wrangling-puzzles","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/xiaodaigh/data-wrangling-puzzles","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xiaodaigh%2Fdata-wrangling-puzzles","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xiaodaigh%2Fdata-wrangling-puzzles/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xiaodaigh%2Fdata-wrangling-puzzles/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xiaodaigh%2Fdata-wrangling-puzzles/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/xiaodaigh","download_url":"https://codeload.github.com/xiaodaigh/data-wrangling-puzzles/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xiaodaigh%2Fdata-wrangling-puzzles/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29294550,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-10T03:42:42.660Z","status":"ssl_error","status_checked_at":"2026-02-10T03:42:41.897Z","response_time":65,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-10-12T19:39:25.320Z","updated_at":"2026-02-10T08:30:52.911Z","avatar_url":"https://github.com/xiaodaigh.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# data-wrangling-puzzles\n\n## Puzzle 1 - Messed up column names need distribution\n\n[Orig link](https://discourse.julialang.org/t/how-would-i-remove-a-column-from-a-dataframe-by-distributing-its-values-among-existing-columns/44265).\n\nFrom \n\n```julia\ndf = DataFrame(:person=\u003e[\"bob\",\"phil\",\"nick\"],:london=\u003e[1,1,0],:spain=\u003e[1,0,0],Symbol(\"london,spain\")=\u003e[1,1,1])\n```\nto\n\n```julia\ndf = DataFrame(:person=\u003e[\"bob\",\"phil\",\"nick\"],:london=\u003e[2,2,1],:spain=\u003e[2,1,1])\n```\n\n### Julia solution\n\n\u003cdetails\u003e\n\n```julia\nfunction row_spread(row)\n    dict = Dict{String, Int}()\n    for (colname, val) in zip(keys(row), values(row))\n        if colname == :person\n            continue\n        end\n        for country in split(string(colname), \",\")\n            dict[country] = get(dict, country, 0) + val\n        end\n    end\n    new_row = hcat(DataFrame(person = row.person), DataFrame(dict))\n    new_row\nend\n\nnew_df = reduce(vcat, row_spread(row) for row in eachrow(df))\n```\n\n\u003c/details\u003e\n\n## Puzzle 2 - Pivot a dataframe to wide format with values in multiple columns\n\nSee https://discourse.julialang.org/t/pivot-a-dataframe-to-wide-format-with-values-in-multiple-columns/45916\n\n\u003cdetails\u003e\n```\nwide = DataFrame(x = 1:12,\n       a  = 2:13,\n       b  = 3:14,\n       val1  = randn(12),\n       val2  = randn(12),\n       cname = repeat([\"c\", \"d\"], inner =6)\n       )\n\n12×6 DataFrame\n│ Row │ x     │ a     │ b     │ val1      │ val2      │ cname  │\n│     │ Int64 │ Int64 │ Int64 │ Float64   │ Float64   │ String │\n├─────┼───────┼───────┼───────┼───────────┼───────────┼────────┤\n│ 1   │ 1     │ 2     │ 3     │ 1.51014   │ -1.18548  │ c      │\n│ 2   │ 2     │ 3     │ 4     │ 0.0845411 │ -0.370083 │ c      │\n│ 3   │ 3     │ 4     │ 5     │ 0.826283  │ -1.00423  │ c      │\n│ 4   │ 4     │ 5     │ 6     │ -0.53175  │ -1.16659  │ c      │\n│ 5   │ 5     │ 6     │ 7     │ -1.77975  │ 0.336333  │ c      │\n│ 6   │ 6     │ 7     │ 8     │ 0.632577  │ 0.236621  │ c      │\n│ 7   │ 7     │ 8     │ 9     │ -0.681532 │ 1.14869   │ d      │\n│ 8   │ 8     │ 9     │ 10    │ -0.775619 │ 0.393475  │ d      │\n│ 9   │ 9     │ 10    │ 11    │ -0.533034 │ 0.059624  │ d      │\n│ 10  │ 10    │ 11    │ 12    │ 0.496152  │ -1.23507  │ d      │\n│ 11  │ 11    │ 12    │ 13    │ 0.834099  │ 2.12115   │ d      │\n│ 12  │ 12    │ 13    │ 14    │ 0.532357  │ -0.369267 │ d      │\n```\n\nI am trying to mimic the pivot_wider function in R:\n\n`wide %\u003e% pivot_wider(names_from = cname, values_from = c(val1,val2))`\n\n===  ===  ===  ==========  ==========  ==========  ==========\n  x    a    b      val1_c      val1_d      val2_c      val2_d\n===  ===  ===  ==========  ==========  ==========  ==========\n  1    2    3   1.0174232          NA  -0.6611959          NA\n  2    3    4   0.6590795          NA  -2.0954505          NA\n  3    4    5   1.2939581          NA   1.6350356          NA\n  4    5    6  -1.9395356          NA   0.7813238          NA\n  5    6    7   0.3558087          NA   0.9789414          NA\n  6    7    8   0.9859100          NA  -0.9803336          NA\n  7    8    9          NA   0.4949224          NA  -0.0659333\n  8    9   10          NA   0.5024755          NA  -0.2317832\n  9   10   11          NA   1.6926897          NA  -0.3840687\n 10   11   12          NA  -0.4324705          NA  -0.0901276\n 11   12   13          NA  -0.6415260          NA   0.0014151\n 12   13   14          NA   1.2406868          NA  -2.1959740\n===  ===  ===  ==========  ==========  ==========  ==========\n```\n\u003c/details\u003e\n\n\n## Puzzle 3\n\n\u003cdetails\u003e\n\u003csummary\u003eKeep only certain rows based on data in group\u003c/summary\u003e\n\nI need to select groups of observations from a large dataframe (about 2.9 mio rows) using a number of conditions which apply to different observations in each group (so I cannot select on individual rows only).\n\nUsing a small dataframe as a starting point, I wrote an algorithm which applies the conditions and generates the result dataframe with my desired groups within a loop. See mwe below.\n\nIf I apply this loop to the large dataframe, performance becomes a (serious) problem.\n\nI haven’t used the split/apply/combine approach before so I am learning about it right now. I worked through the documentation but I haven’t been able to write code for my problem yet.\n\nFor example, I am struggling to understand how to select different rows of groupeddataframes. I figured out how to get the age of the status1 row select(combine(first, gdf), :status =\u003e :obs1_status) but not for the status2 row.\n\nAny hints/guidance on how to implement selection conditions on groupeddataframes using combine/select/transform commands? (I.e. generate df_result in a faster way?)\n\nThanks a lot!\n\n```\nusing DataFrames\n\n# generate sample dataframe\ndf = DataFrame(id = [1,1,1,2,2,3,4], age = [53,52,17,31,29,22,71], status = [1,2,3,1,2,1,1])\n\n# initialize result dataframe\ndf_result = copy(df[1:2,:]; copycols=true);\n\nfor k = 1:maximum(df.id)\n```\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n    \u003csummary\u003e Sample Data \u003c/summary\u003e\n\n```julia\ndf = DataFrame(id = [1,1,1,2,2,3,4], age = [53,52,17,31,29,22,71], status = [1,2,3,1,2,1,1])\n```\n\n```\n7×3 DataFrame\n│ Row │ id    │ age   │ status │\n│     │ Int64 │ Int64 │ Int64  │\n├─────┼───────┼───────┼────────┤\n│ 1   │ 1     │ 53    │ 1      │\n│ 2   │ 1     │ 52    │ 2      │\n│ 3   │ 1     │ 17    │ 3      │\n│ 4   │ 2     │ 31    │ 1      │\n│ 5   │ 2     │ 29    │ 2      │\n│ 6   │ 3     │ 22    │ 1      │\n│ 7   │ 4     │ 71    │ 1      │\n```\n\u003c/details\u003e\n\n\u003cdetails\u003e\n    \u003csummary\u003e Julia Solutions \u003c/summary\u003e\n \n```julia\nusing Pipe, PairAsPipe, DataFramesMeta\ndf_result = @pipe df |\u003e\n    groupby(_, :id) |\u003e\n    combine(_,\n        @pap(status1and2 = sum(in(1:2), :status)),\n        @pap(not_wokring_age = sum(:status .== 1 .\u0026 (:age .\u003c 25 .| :age .\u003e 61)))\n    ) |\u003e\n    @where(_, :status1and2 .== 2, :not_wokring_age .== 0) |\u003e\n    @select(_, :id) |\u003e\n    innerjoin(_, df; on = :id)\n```\n\n\u003c/details\u003e\n\n## Pivot wider two columns\n\n```\njulia\u003e df = DataFrame(t = [:a, :b, :c, :a, :b, :c], x = 1:6, y = 11:16)\n6×3 DataFrame\n│ Row │ t      │ x     │ y     │\n│     │ Symbol │ Int64 │ Int64 │\n├─────┼────────┼───────┼───────┤\n│ 1   │ a      │ 1     │ 11    │\n│ 2   │ b      │ 2     │ 12    │\n│ 3   │ c      │ 3     │ 13    │\n│ 4   │ a      │ 4     │ 14    │\n│ 5   │ b      │ 5     │ 15    │\n│ 6   │ c      │ 6     │ 16    │\nso that it becomes\n│ Row │ x_a   │ x_b   │ x_c   │ y_a   │ y_b   │ y_c   │\n│     │ Int64 │ Int64 │ Int64 │ Int64 │ Int64 │ Int64 │\n├─────┼───────┼───────┼───────┼───────┼───────┼───────┤\n│ 1   │ 1     │ 2     │ 3     │ 11    │ 12    │ 13    │\n│ 2   │ 4     │ 5     │ 6     │ 14    │ 15    │ 16    │\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fxiaodaigh%2Fdata-wrangling-puzzles","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fxiaodaigh%2Fdata-wrangling-puzzles","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fxiaodaigh%2Fdata-wrangling-puzzles/lists"}