{"id":17801397,"url":"https://github.com/evamaerey/ggsmoothfit","last_synced_at":"2026-02-22T19:03:21.732Z","repository":{"id":198545691,"uuid":"690232840","full_name":"EvaMaeRey/ggsmoothfit","owner":"EvaMaeRey","description":"extending stat_smooth to return fitted values and residuals","archived":false,"fork":false,"pushed_at":"2025-03-20T20:42:35.000Z","size":1942,"stargazers_count":2,"open_issues_count":1,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-04-02T04:47:12.998Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"R","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/EvaMaeRey.png","metadata":{"files":{"readme":"README.Rmd","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-09-11T19:46:49.000Z","updated_at":"2025-03-20T20:42:39.000Z","dependencies_parsed_at":null,"dependency_job_id":"f73c95a2-f50d-41ff-801f-35b9d25f0ad2","html_url":"https://github.com/EvaMaeRey/ggsmoothfit","commit_stats":null,"previous_names":["evamaerey/ggsmoothfit"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/EvaMaeRey/ggsmoothfit","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/EvaMaeRey%2Fggsmoothfit","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/EvaMaeRey%2Fggsmoothfit/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/EvaMaeRey%2Fggsmoothfit/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/EvaMaeRey%2Fggsmoothfit/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/EvaMaeRey","download_url":"https://codeload.github.com/EvaMaeRey/ggsmoothfit/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/EvaMaeRey%2Fggsmoothfit/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29723574,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-22T15:10:41.462Z","status":"ssl_error","status_checked_at":"2026-02-22T15:10:04.636Z","response_time":110,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-10-27T12:38:08.067Z","updated_at":"2026-02-22T19:03:21.692Z","avatar_url":"https://github.com/EvaMaeRey.png","language":"R","funding_links":[],"categories":[],"sub_categories":[],"readme":"---\noutput: \n  github_document:\n    toc: TRUE\n---\n\n\u003c!-- README.md is generated from README.Rmd. Please edit that file --\u003e\n\n```{r, include = FALSE}\nknitr::opts_chunk$set(\n  collapse = TRUE,\n  comment = \"#\u003e\"\n)\nlibrary(tidyverse, warn.conflicts = F)\nggplot2::theme_set(theme_gray(base_size = 18))\n```\n\n# ggsmoothfit \n\n\u003c!-- badges: start --\u003e\n\u003c!-- badges: end --\u003e\n\nThe goal of {ggsmoothfit} is to let you visualize model fitted values and residuals easily!\n\n``` r\nlibrary(tidyverse, warn.conflicts = F)\nlibrary(ggsmoothfit)\nmtcars %\u003e% \n  ggplot() + \n  aes(wt, mpg) +\n  geom_point() +\n  geom_smooth() +\n  ggsmoothfit:::geom_fit() + \n  ggsmoothfit:::geom_residuals() + \n  ggsmoothfit:::geom_smooth_fit(xseq = 2:3, size = 5) +\n  ggsmoothfit:::geom_smooth_step(xseq = 2:3) + \n  ggsmoothfit:::geom_smooth_fit(xseq = 0, size = 5, method = lm)\n```\n\n\n# Let's build this functionality\n\n\n# Step 0. Examine ggplot2::StatSmooth$compute_group, and a dataframe that it returns\n\n\nKey take away: this function allows you to set values of x with the *xseq argument*.  Although the default is to create an evenly spaced sequence. \n\n```{r}\nggplot2::StatSmooth$compute_group %\u003e% capture.output() %\u003e% .[1:10]\n\nlibrary(dplyr)\nmtcars %\u003e%\n  rename(x = wt, y = mpg, cat = am) %\u003e%\n  ggplot2::StatSmooth$compute_group(method = lm, \n                           formula = y ~ x, n = 7)\n```\n\n# Step 1. create compute_group_smooth_fit \n\nHere we'll piggy back on StatSmooth$compute_group, to create a function, compute_group_smooth_fit.  We ask that function to compute predictions at the values of x observed in our data set.  We also preserve the values of y (as yend) so that we can draw in the residual error.\n\nxend and yend are computed to draw the segments visualizing the error.\n\n```{r compute_group_smooth_fit}\ncompute_group_smooth_fit \u003c- function(data, scales, method = NULL, formula = NULL,\n                           xseq = NULL,\n                           level = 0.95, method.args = list(),\n                           na.rm = FALSE, flipped_aes = NA){\n  \n  if(is.null(xseq)){ # predictions based on observations \n\n  ggplot2::StatSmooth$compute_group(data = data, scales = scales, \n                       method = method, formula = formula, \n                       se = FALSE, n= 80, span = 0.75, fullrange = FALSE,\n                       xseq = data$x, \n                       level = .95, method.args = method.args, \n                       na.rm = na.rm, flipped_aes = flipped_aes) |\u003e\n      dplyr::mutate(xend = data$x,\n                    yend = data$y)\n  \n  }else{  # predict specific input values\n    \n  ggplot2::StatSmooth$compute_group(data = data, scales = scales, \n                       method = method, formula = formula, \n                       se = FALSE, n= 80, span = 0.75, fullrange = FALSE,\n                       xseq = xseq, \n                       level = .95, method.args = method.args, \n                       na.rm = na.rm, flipped_aes = flipped_aes)   \n    \n  }\n  \n}\n```\n\n\n```{r}\nmtcars %\u003e% \n  slice(1:10) %\u003e% \n  rename(x = wt, y = mpg) %\u003e% \n  compute_group_smooth_fit(method = lm, formula = y ~ x)\n```\n\n\nWe'll also create compute_group_smooth_sq_error, further piggybacking, this time on the function we just build.  This creates the ymin, ymax, xmin and xmax columns needed to show the *squared* error.  Initially, I'd included this computation above, but the plot results can be bad, as the 'flags' that come off of the residuals effect the plot spacing even when they aren't used.  Preferring to avoid this side-effect, we create two functions (and later two ggproto objects).  Note too that xmax is computed in the units of y, and initial plotting can yield squares that do not look like squares. Standardizing both variables, with coord_equal will get us to squares. \n\n```{r compute_group_smooth_sq_error}\ncompute_group_smooth_sq_error \u003c- function(data, scales, method = NULL, \n                                          formula = NULL,\n                          \n                           level = 0.95, method.args = list(),\n                           na.rm = FALSE, flipped_aes = NA){\n  \n compute_group_smooth_fit(data = data, scales = scales, \n                       method = method, formula = formula, \n                       level = .95, method.args = method.args, \n                       na.rm = na.rm, flipped_aes = flipped_aes) %\u003e% \n    dplyr::mutate(ymin = y,\n           xmin = x,\n           ymax = yend,\n           xmax = x + (ymax - ymin))\n  \n}\n```\n\n\n\n# Step 1.1 test compute group \n\n```{r}\nmtcars %\u003e% \n  slice(1:10) %\u003e% \n  rename(x = wt, y = mpg) %\u003e% \n  compute_group_smooth_fit(method = lm, formula = y ~ x)\n\nmtcars %\u003e% \n  slice(1:10) %\u003e% \n  rename(x = wt, y = mpg) %\u003e% \n  compute_group_smooth_sq_error(method = lm, formula = y ~ x)\n```\n\n# Step 2. Pass to ggproto\n\n```{r ggproto_objects}\nStatSmoothFit \u003c- ggplot2::ggproto(\"StatSmoothFit\", \n                                  ggplot2::StatSmooth,\n                                  compute_group = compute_group_smooth_fit,\n                                  required_aes = c(\"x\", \"y\"))\n\nStatSmoothErrorSq \u003c- ggplot2::ggproto(\"StatSmoothErrorSq\", \n                                      ggplot2::StatSmooth,\n                                      compute_group = compute_group_smooth_sq_error,\n                                      required_aes = c(\"x\", \"y\"))\n```\n\n\n# Try Out Stat\n\n```{r}\nmtcars %\u003e% \n  ggplot() + \n  aes(x = wt, y = mpg) + \n  geom_point() + \n  geom_smooth() + \n  geom_point(stat = StatSmoothFit, color = \"blue\") + \n  geom_segment(stat = StatSmoothFit, color = \"blue\")\n```\n\n\n# Step 3. Pass to stat_*/ geom_ functions\n\n```{r stat_fit}\nlibrary(statexpress)\n\nstat_smooth_fit \u003c- function(geom = \"point\", ...){\n  \n  qlayer(geom = geom,\n         stat = StatSmoothFit, ...)\n  \n}\n\n\ngeom_smooth_fit \u003c- function(...){\n  \n  qlayer(geom = qproto_update(GeomPoint, aes(colour = from_theme(accent))),\n         stat = StatSmoothFit, ...)\n  \n}\n\ngeom_smooth_residuals \u003c- function(...){\n  \n  qlayer(geom = qproto_update(GeomSegment, aes(colour = from_theme(accent))),\n         stat = StatSmoothFit, ...)\n  \n}\n\nmtcars %\u003e% \n  ggplot() + \n  aes(x = wt, y = mpg) + \n  geom_point() + \n  geom_smooth() + \n  geom_smooth_fit() + \n  geom_smooth_residuals() + \n  geom_smooth_fit(xseq = 2:3, size = 8)\n\n```\n\n\n```{r stat_errorsq}\ngeom_smooth_residuals_squared \u003c- function(...){\n  \n  qlayer(geom = qproto_update(GeomRect, \n                              aes(fill = from_theme(accent), \n                                  alpha = .2,\n                                  color = from_theme(accent),\n                                  linewidth = from_theme(linewidth*.2))),\n         stat = StatSmoothErrorSq,\n         ...)\n  \n}\n\nstandardize \u003c- function(x){\n  \n  var_mean \u003c- mean(x) \n  var_sd \u003c- sd(x)\n  \n  (x-var_mean)/var_sd\n  \n}\n\n```\n\n\nFor best results, use standardized x, y and coord_equal() as shown below\n\n\n```{r}\nmtcars %\u003e% \n  ggplot() + \n  aes(x = wt, y = mpg) + \n  geom_point() + \n  geom_smooth() + \n  geom_smooth_fit() + \n  geom_smooth_residuals() +\n  geom_smooth_residuals_squared()\n\nlast_plot() + \n  coord_equal()\n\nlast_plot() +\n  aes(standardize(wt), standardize(mpg)) \n\n```\n\n\n\n# And with lm\n\n```{r}\nmtcars %\u003e% \n  ggplot() +\n  aes(wt, mpg) + \n  geom_point() +\n  geom_smooth(alpha = .2, se = FALSE, method = lm) + \n  geom_smooth_fit(method = lm) + # wrap as geom_smooth_fit()\n  geom_smooth_residuals(method = lm) + \n  geom_smooth_fit(xseq = 0, method = lm)\n```\n\n# Contrast to an empty model...\n\n- show mean of y, residuals and squares (variance)\n\n```{r}\nmtcars %\u003e% \n  ggplot() +\n  aes(standardize(wt), standardize(mpg)) + \n  geom_point() +\n  geom_smooth(method = lm, formula = y ~ 1) + \n  geom_smooth_fit(method = lm, formula = y ~ 1) + # wrap as geom_smooth_fit()\n  geom_smooth_residuals(method = lm, formula = y ~ 1) + \n  geom_smooth_residuals_squared(method = lm, formula = y ~ 1) + \n  coord_equal()\n\n```\n\n\n\n\n```{r geom_smooth_step}\ngeom_smooth_step \u003c- function(method = NULL, formula = y ~ x,\n                              color = \"darkred\", xseq = 0:1){\n  \n  stat_smooth(method = method,\n              formula = formula,\n              geom = \"segment\", # draw fitted values as points\n              color = color,\n              xseq = xseq, # 'from', 'to' value pair\n              aes(yend = after_stat(y[1]), # 'from' value of y\n                  xend = xseq[2]), # 'to' value of x\n              arrow = arrow(ends = c(\"last\", \"first\"), \n                            length = unit(.1, \"in\")))\n} \n```\n\n\n```{r}\nmtcars |\u003e\n  ggplot(aes(wt, mpg)) +\n  geom_point() +\n  geom_smooth(method = lm) +\n  geom_smooth_step(method = lm, xseq = 2:3)\n\n```\n\n\n\n# Via @friendly [ggplot2 extenders ggsprings discussion]( https://github.com/ggplot2-extenders/ggplot-extension-club/discussions/83) and [springs extension case study](https://ggplot2-book.org/ext-springs.html)\n\n```{r}\nlibrary(ggplot2)\n\ncreate_spring \u003c- function(x, \n                          y, \n                          xend, \n                          yend, \n                          diameter = 1, \n                          tension = 0.75, \n                          n = 50) {\n  \n  # Validate the input arguments\n  if (tension \u003c= 0) {\n    rlang::abort(\"`tension` must be larger than zero.\")\n  }\n  if (diameter == 0) {\n    rlang::abort(\"`diameter` can not be zero.\")\n  }\n  if (n == 0) {\n    rlang::abort(\"`n` must be greater than zero.\")\n  }\n  \n  # Calculate the direct length of the spring path\n  length \u003c- sqrt((x - xend)^2 + (y - yend)^2)\n  \n  # Calculate the number of revolutions and points we need\n  n_revolutions \u003c- length / (diameter * tension)\n  n_points \u003c- n * n_revolutions\n  \n  # Calculate the sequence of radians and the x and y offset values\n  radians \u003c- seq(0, n_revolutions * 2 * pi, length.out = n_points)\n  x \u003c- seq(x, xend, length.out = n_points)\n  y \u003c- seq(y, yend, length.out = n_points)\n  \n  # Create and return the transformed data frame\n  data.frame(\n    x = cos(radians) * diameter/2 + x,\n    y = sin(radians) * diameter/2 + y\n  )\n}\n\nGeomSmoothSpring \u003c- ggproto(\"GeomSmoothSpring\", Geom,\n  \n  # Ensure that each row has a unique group id\n  setup_data = function(data, params) {\n    if (is.null(data$group)) {\n      data$group \u003c- seq_len(nrow(data))\n    }\n    if (anyDuplicated(data$group)) {\n      data$group \u003c- paste(data$group, seq_len(nrow(data)), sep = \"-\")\n    }\n    data\n  },\n  \n  # Transform the data inside the draw_panel() method\n  draw_panel = function(data, \n                        panel_params, \n                        coord, \n                        n = 50, \n                        arrow = NULL,\n                        lineend = \"butt\", \n                        linejoin = \"round\", \n                        linemitre = 10,\n                        na.rm = FALSE) {\n    \n    # Transform the input data to specify the spring paths\n    cols_to_keep \u003c- setdiff(names(data), c(\"x\", \"y\", \"xend\", \"yend\"))\n    \n    data$diameter \u003c- data$diameter %||%  (.025 * abs(min(data$x)-max(data$x)))\n    data$springlength \u003c- sqrt((data$x-data$xend)^2 + (data$y-data$yend)^2)\n    data$tension \u003c- data$tension %||% (1 * data$springlength)\n    \n    springs \u003c- lapply(seq_len(nrow(data)), function(i) {\n      spring_path \u003c- create_spring(\n        data$x[i], \n        data$y[i], \n        data$xend[i], \n        data$yend[i], \n        data$diameter[i], \n        data$tension[i], \n        n\n      )\n      cbind(spring_path, unclass(data[i, cols_to_keep]))\n    })\n    springs \u003c- do.call(rbind, springs)\n    \n    # Use the draw_panel() method from GeomPath to do the drawing\n    GeomPath$draw_panel(\n      data = springs, \n      panel_params = panel_params, \n      coord = coord, \n      arrow = arrow, \n      lineend = lineend, \n      linejoin = linejoin, \n      linemitre = linemitre, \n      na.rm = na.rm\n    )\n  },\n  \n  # Specify the default and required aesthetics\n  required_aes = c(\"x\", \"y\", \"xend\", \"yend\"),\n  default_aes = aes(\n    colour = from_theme(accent), \n    linewidth = 0.5, \n    linetype = 1L, \n    alpha = NA\n  )\n  \n)\n\n\nanscombe |\u003e \n  ggplot() + \n  aes(x = x1, y = y1) + \n  geom_point() + \n  geom_smooth(method = lm) +\n  stat_smooth_fit(geom = GeomSmoothSpring, method = lm) + \n  ggchalkboard:::theme_blackboard()\n\nlast_plot() + \n  aes(x = x2, y = y2)\n\nlast_plot() + \n  aes(tension = 1)\n\nlast_plot() + \n  aes(x = x3, y = y3) + \n  aes(tension = NULL)\n\nlast_plot() + \n  aes(x = x4, y = y4)\n\n\nanscombe |\u003e \n  ggplot() + \n  aes(x = x2, y = y2) + \n  geom_point() + \n  geom_smooth(method = lm, formula = y ~ 1) +\n  stat_smooth_fit(geom = GeomSmoothSpring, method = lm, formula = y ~ 1) + \n  ggchalkboard:::theme_blackboard()\n\nlast_plot() + \n  aes(tension = 1)\n\n```\n\n\n\n\n\n\n# Part 2. Packaging and documentation  🚧 ✅ \n\n\n\n## minimal requirements for github package.  Have you:\n\n### Created files for package archetecture with `devtools::create(\"./ggbarlabs\")` ✅ \n\n### Moved functions R folder? ✅  \n\n```{r}\nknitr::knit_code$get() |\u003e names()\n```\n\n\n```{r}\nknitrExtra::chunk_to_dir(c(\n                         \"geom_smooth_step\",\n                         \"compute_group_smooth_fit\", \n                         \"compute_group_smooth_sq_error\",\n                         \"ggproto_objects\",\n                         \"stat_fit\", \n                         \"stat_errorsq\"))\n```\n\n\n\n### Added roxygen skeleton? ✅ \n\n for auto documentation and making sure proposed functions are *exported*\n\n### Managed dependencies ? ✅ \n\npackage dependencies managed, i.e. `depend::function()` in proposed functions and declared in the DESCRIPTION\n\n### Chosen a license? ✅ \n\n\n```{r, eval = F}\nusethis::use_package(\"ggplot2\")\nusethis::use_mit_license()\n```\n\n\n\n## `devtools::check()` report\n\n```{r, error = T, eval = F}\n# rm(list = c(\"geom_barlab_count\", \"geom_barlab_count_percent\"))\ndevtools::check(pkg = \".\")\n```\n\n\n---\n\n# Don't want to use ggsmoothfit?  Here are some ways to get it done with base ggplot2!\n\n## Option 1. Verbal description and move on...\n\n\"image a line that drops down from the observation to the model line\" use vanilla geom_smooth\n\n```{r stat-smooth}\nmtcars %\u003e% \n  ggplot() +\n  aes(wt, mpg) + \n  geom_point() + \n  geom_smooth()\n```\n\n\n## Option 2: precalculate and plot \n\n[stack overflow example goes here.]\n\n\n## Option 3: little known xseq argument and geom = \"point\"\n\nFirst a bit of under-the-hood thinking about geom_smooth/stat_smooth.\n\n```{r}\nmtcars %\u003e% \n  ggplot() +\n  aes(wt, mpg) + \n  geom_smooth(n = 12) +\n  stat_smooth(geom = \"point\", \n              color = \"blue\", \n              n = 12)\n```\n\nSpecify xseq... Almost surely new to you (and probably more interesting to stats instructors): predicting at observed values of x..\n\nxseq has only recently been advertised, but possibly of interest.. https://ggplot2.tidyverse.org/reference/geom_smooth.html\n\n```{r, warning= F, message=F}\n# fit where the support is in the data... \nmtcars %\u003e% \n  ggplot() +\n  aes(wt, mpg) + \n  geom_point() +\n  geom_smooth() + \n  stat_smooth(geom = \"point\",  color = \"blue\", # fitted values\n              xseq = mtcars$wt) +\n  stat_smooth(geom = \"segment\", color = \"darkred\", # residuals\n              xseq = mtcars$wt,\n              xend = mtcars$wt,\n              yend = mtcars$mpg)\n```\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fevamaerey%2Fggsmoothfit","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fevamaerey%2Fggsmoothfit","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fevamaerey%2Fggsmoothfit/lists"}