{"id":13520188,"url":"https://github.com/cfusting/fast-symbolic-regression","last_synced_at":"2025-03-31T16:30:52.654Z","repository":{"id":77710387,"uuid":"108316605","full_name":"cfusting/fast-symbolic-regression","owner":"cfusting","description":"Blazing fast symbolic regresison","archived":false,"fork":false,"pushed_at":"2019-10-10T01:02:47.000Z","size":167,"stargazers_count":76,"open_issues_count":4,"forks_count":20,"subscribers_count":5,"default_branch":"master","last_synced_at":"2024-08-02T05:22:59.658Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/cfusting.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2017-10-25T19:32:05.000Z","updated_at":"2024-05-29T20:35:51.000Z","dependencies_parsed_at":"2023-04-28T12:54:33.245Z","dependency_job_id":null,"html_url":"https://github.com/cfusting/fast-symbolic-regression","commit_stats":null,"previous_names":[],"tags_count":1,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cfusting%2Ffast-symbolic-regression","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cfusting%2Ffast-symbolic-regression/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cfusting%2Ffast-symbolic-regression/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cfusting%2Ffast-symbolic-regression/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/cfusting","download_url":"https://codeload.github.com/cfusting/fast-symbolic-regression/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":222670691,"owners_count":17020513,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-01T05:02:13.633Z","updated_at":"2024-11-02T03:30:33.771Z","avatar_url":"https://github.com/cfusting.png","language":"Python","funding_links":[],"categories":["Python","software toolls"],"sub_categories":[],"readme":"# Fast Symbolic Regression\n\nSymbolic Regression is a non-linear, non-parametric Machine Learning method capable of modeling complex data sets. fastsr aims at providing the most simple, powerful models possible by optimizing not only for error but also for model complexity.\nfastsr is built on top of [fastgp](https://github.com/cfusting/fastgp), a numpy implementation of [genetic programming](https://en.wikipedia.org/wiki/Genetic_programming) built on top of [deap](https://github.com/DEAP/deap).\nAll estimators adhere to the [sklearn](http://scikit-learn.org/stable/) estimator interface and can thus be used in pipelines.\n\nfastsr was designed and developed by the [Morphology, Evolution \u0026 Cognition Laboratory](http://www.meclab.org/) at the University of Vermont. It extends research code which can be found [here](https://github.com/mszubert/gecco_2016).\n\nInstallation\n------------\nfastsr is compatible with Python 2.7+.\n```bash\npip install fastsr\n```\n\nExample Usage\u003ca name=\"ex\"\u003e\u003c/a\u003e\n------------------------------\n[Symbolic Regression](https://en.wikipedia.org/wiki/Symbolic_regression) is really good at fitting nonlinear functions. Let's try to fit the third order polynomial x^3 + x^2 + x. This is the \"regression\" example from the examples folder.\n```python\nimport matplotlib.pyplot as plt\n\nimport numpy as np\n\nfrom fastsr.estimators.symbolic_regression import SymbolicRegression\n\nfrom fastgp.algorithms.fast_evaluate import fast_numpy_evaluate\nfrom fastgp.parametrized.simple_parametrized_terminals import get_node_semantics\n\nfrom sklearn.model_selection import train_test_split\n```\n```python\ndef target(x):\n    return x**3 + x**2 + x\n```\nNow we'll generate some data on the domain \\[-10, 10\\].\n```python\nX = np.linspace(-10, 10, 100, endpoint=True)\ny = target(X)\nX_train, X_test, y_train, y_test = train_test_split(X, y)\n```\nFinally we'll create and fit the Symbolic Regression estimator and check the score.\n```python\nsr = SymbolicRegression(ngen=100, pop_size=100, stop_time=60)\nsr.fit(X_train, y_train)\nscore = sr.score(X_test, y_test)\nprint('Score: {}'.format(score))\n```\n```\nScore: 0.0\n```\nWhoa! That's not much error. Don't get too used to scores like that though, real data sets aren't usually as simple as a third order polynomial.\n\nfastsr uses Genetic Programming to fit the data. That means equations are evolving to fit the data better and better each generation. Let's have a look at the best individuals and their respective scores.\n```python\nprint('Best Individuals:')\nsr.print_best_individuals()\n```\n```\nBest Individuals:\n0.0 : add(add(square(X0), cube(X0)), X0)\n34.006734006733936 : add(square(X0), cube(X0))\n2081.346746380927 : add(cube(X0), X0)\n2115.3534803876605 : cube(X0)\n137605.24466869785 : add(add(X0, add(X0, X0)), add(X0, X0))\n141529.89102341252 : add(add(X0, X0), add(X0, X0))\n145522.55084614072 : add(add(X0, X0), X0)\n149583.22413688237 : add(X0, X0)\n151203.96034032793 : numpy_protected_sqrt(cube(numpy_protected_log_abs(exp(X0))))\n151203.96034032793 : cube(numpy_protected_sqrt(X0))\n153711.91089563753 : numpy_protected_log_abs(exp(X0))\n153711.91089563753 : X0\n155827.26437602515 : square(X0)\n156037.81673350732 : add(numpy_protected_sqrt(X0), cbrt(X0))\n157192.02956807753 : numpy_protected_sqrt(exp(cbrt(X0)))\n```\nAt the top we find our best individual, which is exactly the third order polynomial we defined our target function to be. You might be confused as to why we consider all these other individuals, some with very large errors be be \"best\".\nWe can look through the history object to see some of the equations that led up to our winning model by ordering by error.\n```python\nhistory = sr.history_\npopulation = list(filter(lambda x: hasattr(x, 'error'), list(sr.history_.genealogy_history.values())))\npopulation.sort(key=lambda x: x.error, reverse=True)\n```\nLet's get a sample of the unique solutions. There are quite a few so the print statements have been omitted.\n```python\nX = X.reshape((len(X), 1))\ni = 1\nprevious_errror = population[0]\nunique_individuals = []\nwhile i \u003c len(population):\n    ind = population[i]\n    if ind.error != previous_errror:\n        print(str(i) + ' | ' + str(ind.error) + ' | ' + str(ind))\n        unique_individuals.append(ind)\n    previous_errror = ind.error\n    i += 1\n\n```\nNow we can plot the equations over the target functions.\n```python\ndef plot(index):\n    plt.plot(X, y, 'r')\n    plt.axis([-10, 10, -1000, 1000])\n    y_hat = fast_numpy_evaluate(unique_individuals[index], sr.pset_.context, X, get_node_semantics)\n    plt.plot(X, y_hat, 'g')\n    plt.savefig(str(i) + 'ind.png')\n    plt.gcf().clear()\n\ni = 0\nwhile i \u003c len(unique_individuals):\n    plot(i)\n    i += 10\ni = len(unique_individuals) - 1\nplot(i)\n```\nStitched together into a gif we get a view into the evolutionary process.\n\n![Convergence Gif](docs/converge.gif)\n\nFitness Age Size Complexity Pareto Optimization\n-----------------------------------------------\nIn addition to minimizing the error when creating an interpretable model it's often useful to minimize the size of the equations and their complexity (as defined by the order of an approximating polynomial\u003ca href=\"#lc-1\"\u003e\\[1\\]\u003c/a\u003e). In [Multi-Objective optimization](https://en.wikipedia.org/wiki/Multi-objective_optimization) we keep all individuals that are not dominated by any other individuals and call this group the Pareto Front. These are the individuals printed in the \u003ca href=\"#ex\"\u003eExample Usage\u003c/a\u003e above. The age component helps prevent the population of equations from falling into a local optimum and was introduced in AFPO \u003ca href=\"#lc-2\"\u003e\\[2\\]\u003ca\u003e but is out of the scope of this readme.\n\nThe result of this optimization technique is that a range of solutions are considered \"best\" individuals. Although in practice you will probably be interested in the top or several top individuals, be aware that the population as a whole was pressured into keeping individual equations as simple as possible in addition to keeping error as low as possible.\n\nLiterature Cited\n----------------\n1. Ekaterina J Vladislavleva, Guido F Smits, and Dick Den Hertog. 2009. Order of nonlinearity as a complexity measure for models generated by symbolic regression via pareto genetic programming. IEEE Transactions on Evolutionary Computation 13, 2 (2009), 333–349.\u003ca name=\"lc-1\"\u003e\u003c/a\u003e\n2. Michael Schmidt and Hod Lipson. 2011. Age-fitness pareto optimization. In Genetic Programming Theory and Practice VIII. Springer, 129–146.\u003ca name=\"lc-2\"\u003e\u003c/a\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcfusting%2Ffast-symbolic-regression","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcfusting%2Ffast-symbolic-regression","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcfusting%2Ffast-symbolic-regression/lists"}