{"id":13583468,"url":"https://github.com/viasite/site-audit-seo","last_synced_at":"2026-03-14T22:17:17.758Z","repository":{"id":37172743,"uuid":"244975082","full_name":"viasite/site-audit-seo","owner":"viasite","description":"Web service and CLI tool for SEO site audit: crawl site, lighthouse all pages, view public reports in browser. Also output to console, json, csv, xlsx","archived":false,"fork":false,"pushed_at":"2025-06-24T09:21:54.000Z","size":10981,"stargazers_count":274,"open_issues_count":11,"forks_count":37,"subscribers_count":8,"default_branch":"master","last_synced_at":"2025-11-27T10:37:01.385Z","etag":null,"topics":["audit","cli","crawl-site","crawler","lighthouse","puppeteer","scraper","seo","seo-audit","seo-site-audit","site-audit","xlsx"],"latest_commit_sha":null,"homepage":"http://json-viewer.popstas.pro/scan","language":"JavaScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/viasite.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2020-03-04T18:31:06.000Z","updated_at":"2025-11-17T16:44:14.000Z","dependencies_parsed_at":"2025-04-09T19:16:47.610Z","dependency_job_id":"f5a59103-9445-47ce-a182-dcd0b0cd208b","html_url":"https://github.com/viasite/site-audit-seo","commit_stats":null,"previous_names":["viasite/sites-scraper"],"tags_count":38,"template":false,"template_full_name":null,"purl":"pkg:github/viasite/site-audit-seo","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/viasite%2Fsite-audit-seo","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/viasite%2Fsite-audit-seo/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/viasite%2Fsite-audit-seo/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/viasite%2Fsite-audit-seo/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/viasite","download_url":"https://codeload.github.com/viasite/site-audit-seo/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/viasite%2Fsite-audit-seo/sbom","scorecard":{"id":919893,"data":{"date":"2025-08-11","repo":{"name":"github.com/viasite/site-audit-seo","commit":"6a6d92525c8917c19cff4a98be80f31ce4f954b1"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":1.7,"checks":[{"name":"Code-Review","score":0,"reason":"Found 2/28 approved changesets -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Maintained","score":1,"reason":"2 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 1","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Dangerous-Workflow","score":-1,"reason":"no workflows found","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Token-Permissions","score":-1,"reason":"No tokens found","details":null,"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":9,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Warn: project license file does not contain an FSF or OSI license."],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: containerImage not pinned by hash: Dockerfile:1: pin your Docker image by updating node:18-slim to node:18-slim@sha256:f9ab18e354e6855ae56ef2b290dd225c1e51a564f87584b9bd21dd651838830e","Warn: npmCommand not pinned by hash: Dockerfile:38","Info:   0 out of   1 containerImage dependencies pinned","Info:   0 out of   1 npmCommand dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'master'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 5 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}},{"name":"Vulnerabilities","score":0,"reason":"21 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: GHSA-7v5v-9h63-cj86","Warn: Project is vulnerable to: GHSA-jr5f-v2jv-69x6","Warn: Project is vulnerable to: GHSA-v6h2-p8h4-qcjw","Warn: Project is vulnerable to: GHSA-grv7-fg5c-xmjg","Warn: Project is vulnerable to: GHSA-pxg6-pf52-xh8x","Warn: Project is vulnerable to: GHSA-3xgq-45jj-v275","Warn: Project is vulnerable to: GHSA-fjxv-7rqg-78g4","Warn: Project is vulnerable to: GHSA-952p-6rrq-rcjv","Warn: Project is vulnerable to: GHSA-c2qf-rxjj-qqgw","Warn: Project is vulnerable to: GHSA-pq67-2wwv-3xjx","Warn: Project is vulnerable to: GHSA-8cj5-5rvv-wf4v","Warn: Project is vulnerable to: GHSA-3h5v-q93c-6h6q","Warn: Project is vulnerable to: GHSA-m95q-7qp3-xv42","Warn: Project is vulnerable to: GHSA-8hc4-vh64-cxmj","Warn: Project is vulnerable to: GHSA-qwcr-r2fm-qrc7","Warn: Project is vulnerable to: GHSA-qw6h-vgh9-j6wx","Warn: Project is vulnerable to: GHSA-9wv6-86v2-598j","Warn: Project is vulnerable to: GHSA-rhx6-c78j-4q9w","Warn: Project is vulnerable to: GHSA-35q2-47q7-3pc3","Warn: Project is vulnerable to: GHSA-m6fv-jmcg-4jfg","Warn: Project is vulnerable to: GHSA-cm22-4g7w-348p"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}}]},"last_synced_at":"2025-08-25T00:47:18.982Z","repository_id":37172743,"created_at":"2025-08-25T00:47:18.983Z","updated_at":"2025-08-25T00:47:18.983Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":30519575,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-14T19:51:21.629Z","status":"ssl_error","status_checked_at":"2026-03-14T19:51:12.959Z","response_time":57,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audit","cli","crawl-site","crawler","lighthouse","puppeteer","scraper","seo","seo-audit","seo-site-audit","site-audit","xlsx"],"created_at":"2024-08-01T15:03:30.079Z","updated_at":"2026-03-14T22:17:17.750Z","avatar_url":"https://github.com/viasite.png","language":"JavaScript","funding_links":[],"categories":["JavaScript"],"sub_categories":[],"readme":"[![npm](https://img.shields.io/npm/v/site-audit-seo)](https://www.npmjs.com/package/site-audit-seo) [![npm](https://img.shields.io/npm/dt/site-audit-seo)](https://www.npmjs.com/package/site-audit-seo)\n\nWeb service and CLI tool for SEO site audit: crawl site, lighthouse all pages, view public reports in browser. Also output to console, json, csv.\n\nWeb view report - [json-viewer](https://json-viewer.popstas.pro/).\n\nDemo:\n- [Default report](https://json-viewer.popstas.pro/?url=https://site-audit.viasite.ru/reports/blog.popstas.ru-default.json)\n- [Lighthouse report](https://json-viewer.popstas.pro/?url=https://site-audit.viasite.ru/reports/blog.popstas.ru-lighthouse.json)\n- [Default + Basic Lighthouse report](https://json-viewer.popstas.pro/?url=https://site-audit.viasite.ru/reports/blog.popstas.ru-default-plus-lighthouse.json)\n\nРусское описание [ниже](#русский)\n\n![site-audit-demo](assets/site-audit-demo.gif)\n\n## Using without install\nOpen https://json-viewer.popstas.pro/. Public server allow to scan up to 100 pages at once.\n\n## Features:\n- Crawls the entire site, collects links to pages and documents\n- Does not follow links outside the scanned domain (configurable)\n- Analyse each page with Lighthouse (see below)\n- Analyse main page text with Mozilla Readability and Yake\n- Search pages with SSL mixed content\n- Scan list of urls, `--url-list`\n- Set default report fields and filters\n- Scan presets\n- Documents with the extensions `doc`,` docx`, `xls`,` xlsx`, `ppt`,` pptx`, `pdf`,` rar`, `zip` are added to the list with a depth == 0\n\n## Technical details:\n- Does not load images, css, js (configurable)\n- Each site is saved to a file with a domain name in `~/site-audit-seo/`\n- Some URLs are ignored ([`preRequest` in `src/scrap-site.js`](src/scrap-site.js#L98))\n\n### Web viewer features:\n- Fixed table header and url column\n- Add/remove columns\n- Column presets\n- Field groups by categories\n- Filters presets (ex. `h1_count != 1`)\n- Color validation\n- Verbose page details (`+` button)\n- Direct URL to same report with selected fields, filters, sort\n- Stats for whole scanned pages, validation summary\n- Persistent URL to report when `--upload` using\n- Switch between last uploaded reports\n- Rescan current report\n\n\n### Fields list (18.08.2020):\n- url\n- mixed_content_url\n- canonical\n- is_canonical\n- previousUrl\n- depth\n- status\n- request_time\n- redirects\n- redirected_from\n- title\n- h1\n- page_date\n- description\n- keywords\n- og_title\n- og_image\n- schema_types\n- h1_count\n- h2_count\n- h3_count\n- h4_count\n- canonical_count\n- google_amp\n- images\n- images_without_alt\n- images_alt_empty\n- images_outer\n- links\n- links_inner\n- links_outer\n- text_ratio_percent\n- dom_size\n- html_size\n- html_size_rendered\n- lighthouse_scores_performance\n- lighthouse_scores_pwa\n- lighthouse_scores_accessibility\n- lighthouse_scores_best-practices\n- lighthouse_scores_seo\n- lighthouse_first-contentful-paint\n- lighthouse_speed-index\n- lighthouse_largest-contentful-paint\n- lighthouse_interactive\n- lighthouse_total-blocking-time\n- lighthouse_cumulative-layout-shift\n- and 150 more lighthouse tests!\n\n\n## Install\n\n## Zero-knowledge install\nRequires Docker.\n\n### Windows: download and run `install-run.bat`.\nScript will clone repository to `%LocalAppData%\\Programs\\site-audit-seo` and run service on http://localhost:5302.\n\n\n### Linux/MacOS:\n```\ncurl https://raw.githubusercontent.com/viasite/site-audit-seo/master/install-run.sh | bash\n```\n\nScript will clone repository to `$HOME/.local/share/programs/site-audit-seo` and run service on http://localhost:5302.\n\nService will available on http://localhost:5302\n\n##### Default ports:\n- Backend: `5301`\n- Frontend: `5302`\n- Yake: `5303`\n\nYou can change it in `.env` file or in `docker-compose.yml`.\n\n## Install with NPM:\n``` bash\nnpm install -g site-audit-seo\n```\n\n#### For linux users\n``` bash\nnpm install -g site-audit-seo --unsafe-perm=true\n```\n\nAfter installing on Ubuntu, you may need to change the owner of the Chrome directory from root to user.\n\nRun this (replace `$USER` to your username or run from your user, not from `root`):\n``` bash\nsudo chown -R $USER:$USER \"$(npm prefix -g)/lib/node_modules/site-audit-seo/node_modules/puppeteer/.local-chromium/\"\n```\n\n## Install developer instanse with docker-compose\n``` bash\ngit clone https://github.com/viasite/site-audit-seo\ncd site-audit-seo\ngit clone https://github.com/viasite/site-audit-seo-viewer data/front\ndocker-compose pull # for skip build step\ndocker-compose up -d\n```\n\nError details [Invalid file descriptor to ICU data received](https://github.com/puppeteer/puppeteer/issues/2519).\n\n## Command line usage:\n```\n$ site-audit-seo --help\nUsage: site-audit-seo -u https://example.com\n\nOptions:\n  -u --urls \u003curls\u003e                  Comma separated url list for scan\n  -p, --preset \u003cpreset\u003e             Table preset (minimal, seo, seo-minimal, headers, parse, lighthouse,\n                                    lighthouse-all) (default: \"seo\")\n  -t, --timeout \u003ctimeout\u003e           Timeout for page request, in ms (default: 10000)\n  -e, --exclude \u003cfields\u003e            Comma separated fields to exclude from results\n  -d, --max-depth \u003cdepth\u003e           Max scan depth (default: 10)\n  -c, --concurrency \u003cthreads\u003e       Threads number (default: by cpu cores)\n  --lighthouse                      Appends base Lighthouse fields to preset\n  --delay \u003cms\u003e                      Delay between requests (default: 0)\n  -f, --fields \u003cjson\u003e               Field in format --field 'title=$(\"title\").text()' (default: [])\n  --default-filter \u003cdefaultFilter\u003e  Default filter when JSON viewed, example: depth\u003e1\n  --no-skip-static                  Scan static files\n  --no-limit-domain                 Scan not only current domain\n  --docs-extensions \u003cext\u003e           Comma-separated extensions that will be add to table (default:\n                                    doc,docx,xls,xlsx,ppt,pptx,pdf,rar,zip)\n  --follow-xml-sitemap              Follow sitemap.xml (default: false)\n  --ignore-robots-txt               Ignore disallowed in robots.txt (default: false)\n  --url-list                        assume that --url contains url list, will set -d 1 --no-limit-domain\n                                    --ignore-robots-txt (default: false)\n  --remove-selectors \u003cselectors\u003e    CSS selectors for remove before screenshot, comma separated (default:\n                                    \".matter-after,#matter-1,[data-slug]\")\n  -m, --max-requests \u003cnum\u003e          Limit max pages scan (default: 0)\n  --influxdb-max-send \u003cnum\u003e         Limit send to InfluxDB (default: 5)\n  --no-headless                     Show browser GUI while scan\n  --remove-csv                      Delete csv after json generate (default: true)\n  --remove-json                     Delete json after serve (default: true)\n  --no-remove-csv                   No delete csv after generate\n  --no-remove-json                  No delete json after serve\n  --out-dir \u003cdir\u003e                   Output directory (default: \"~/site-audit-seo/\")\n  --out-name \u003cname\u003e                 Output file name, default: domain\n  --csv \u003cpath\u003e                      Skip scan, only convert existing csv to json\n  --json                            Save as JSON (default: true)\n  --no-json                         No save as JSON\n  --upload                          Upload JSON to public web (default: false)\n  --no-color                        No console colors\n  --partial-report \u003cpartialReport\u003e\n  --lang \u003clang\u003e                     Language (en, ru, default: system language)\n  --no-console-validate             Don't output validate messages in console\n  --disable-plugins \u003cplugins\u003e       Comma-separated plugin list (default: [])\n  --screenshot                      Save page screenshot (default: false)\n  -V, --version                     output the version number\n  -h, --help                        display help for command\n```\n\n\n\n## Custom fields\n\n### Linux/Mac:\n``` bash\nsite-audit-seo -d 1 -u https://example -f 'title=$(\"title\").text()' -f 'h1=$(\"h1\").text()'\nsite-audit-seo -d 1 -u https://example -f noindex=$('meta[content=\"noindex,%20nofollow\"]').length\n```\n\n### Windows:\n``` bash\nsite-audit-seo -d 1 -u https://example -f title=$('title').text() -f h1=$('h1').text()\n```\n\n## Remove fields from results\nThis will output fields from `seo` preset excluding canonical fields:\n``` bash\nsite-audit-seo -u https://example.com --exclude canonical,is_canonical\n```\n\n## Lighthouse\n### Analyse each page with Lighthouse\n``` bash\nsite-audit-seo -u https://example.com --preset lighthouse\n```\n\n### Analyse seo + Lighthouse\n``` bash\nsite-audit-seo -u https://example.com --lighthouse\n```\n\n## Config file\nYou can copy [.site-audit-seo.conf.js](.site-audit-seo.conf.js) to your home directory and tune options.\n\n## Send to InfluxDB\nIt is beta feature. How to config:\n\n1. Add this to `~/.site-audit-seo.conf`:\n\n``` js\nmodule.exports = {\n  influxdb: {\n    host: 'influxdb.host',\n    port: 8086,\n    database: 'telegraf',\n    measurement: 'site_audit_seo', // optional\n    username: 'user',\n    password: 'password',\n    maxSendCount: 5, // optional, default send part of pages\n  }\n};\n```\n\n2. Use `--influxdb-max-send` in terminal.\n\n3. Create command for scan your urls:\n\n```\nsite-audit-seo -u https://page-with-url-list.txt --url-list --lighthouse --upload --influxdb-max-send 100 \u003e\u003e ~/log/site-audit-seo.log\n```\n\n4. Add command to cron.\n\n\n## Plugins\n- [Readability](https://github.com/popstas/site-audit-seo-readability) - main page text length, reading time\n- [Yake](https://github.com/popstas/site-audit-seo-yake) - keywords extraction from main page text\n\nSee [CONTRIBUTING.md](CONTRIBUTING.md) for details about plugin development.\n\n#### Install plugins:\n```\ncd data\nnpm install site-audit-seo-readability\nnpm install site-audit-seo-yake\n```\n\n#### Disable plugins:\nYou can add argument such: `--disable-plugins readability,yake`. It more faster, but less data extracted.\n\n## Credentials\nBased on [headless-chrome-crawler](https://github.com/yujiosaka/headless-chrome-crawler) (puppeteer). Used forked version [@popstas/headless-chrome-crawler](https://github.com/popstas/headless-chrome-crawler).\n\n## Bugs\n1. Sometimes it writes identical pages to csv. This happens in 2 cases:\n1.1. Redirect from another page to this (solved by setting `skipRequestedRedirect: true`, hardcoded).\n1.2. Simultaneous request of the same page in parallel threads.\n\n\n## Free audit tools alternatives\n- [WebSite Auditor (Link Assistant)](https://www.link-assistant.com/) - desktop app, 500 pages\n- [Screaming Frog SEO Spider](https://www.screamingfrog.co.uk/seo-spider/) - desktop app, same as site-audit-seo, 500 pages\n- [Seobility](https://www.seobility.net/) - 1 project up to 1000 pages free\n- [Neilpatel (Ubersuggest)](https://app.neilpatel.com/) - 1 project, 150 pages\n- [Semrush](https://semrush.com/) - 1 project, 100 pages per month free\n- [Seoptimer](https://www.seoptimer.com/) - good for single page analysis\n\n\n## Free data scrapers\n- [Web Scraper](https://webscraper.io/) - free for local use extension\n- [Portia](https://github.com/scrapinghub/portia) - self-hosted visual scraper builder, scrapy based\n- [Crawlab](https://github.com/crawlab-team/crawlab) - distributed web crawler admin platform, self-hosted with Docker\n- [OutWit Hub](https://www.outwit.com/#hub) - free edition, pro edition for $99\n- [Octoparse](https://www.octoparse.com/) - 10 000 records free\n- [Parsers.me](https://parsers.me/) - 1 000 pages per run free\n- [website-scraper](https://www.npmjs.com/package/website-scraper) - opensource, CLI, download site to local directory\n- [website-scraper-puppeteer](https://www.npmjs.com/package/website-scraper-puppeteer) - same but puppeteer based\n- [Gerapy](https://github.com/Gerapy/Gerapy) - distributed Crawler Management Framework Based on Scrapy, Scrapyd, Django and Vue.js\n- [DAXRM Rank Tracker](https://www.daxrm.com/integrations/rank-tracker/) - Real-time Google \u0026 Bing SEO Rank tracking\n\n## Русский\nСканирование одного или несколько сайтов в json файл с веб-интерфейсом.\n\n## Особенности:\n- Обходит весь сайт, собирает ссылки на страницы и документы\n- Сводка результатов после сканирования\n- Документы с расширениями `doc`, `docx`, `xls`, `xlsx`, `pdf`, `rar`, `zip` добавляются в список с глубиной 0\n- Поиск страниц с SSL mixed content\n- Каждый сайт сохраняется в файл с именем домена\n- Не ходит по ссылкам вне сканируемого домена (настраивается)\n- Не загружает картинки, css, js (настраивается)\n- Некоторые URL игнорируются ([`preRequest` в `src/scrap-site.js`](src/scrap-site.js#L112))\n- Можно прогнать каждую страницу по Lighthouse (см. ниже)\n- Сканирование произвольного списка URL, `--url-list`\n\n## Установка:\n``` bash\nnpm install -g site-audit-seo\n```\n\n#### Если у вас Ubuntu\n``` bash\nnpm install -g site-audit-seo --unsafe-perm=true\n```\n\n```\nnpm run postinstall-puppeteer-fix\n```\n\nИли запустите это (замените `$USER` на вашего юзера, либо запускайте под юзером, не под `root`):\n``` bash\nsudo chown -R $USER:$USER \"$(npm prefix -g)/lib/node_modules/site-audit-seo/node_modules/puppeteer/.local-chromium/\"\n```\n\nПодробности ошибки [Invalid file descriptor to ICU data received](https://github.com/puppeteer/puppeteer/issues/2519).\n\n\n## Использование\n```\nsite-audit-seo -u https://example.com\n```\n\n\n## Кастомные поля\nМожно передать дополнительные поля так:\n``` bash\nsite-audit-seo -d 1 -u https://example -f \"title=$('title').text()\" -f \"h1=$('h1').text()\"\n```\n\n## Lighthouse\n### Прогнать каждую страницу по Lighthouse\n``` bash\nsite-audit-seo -u https://example.com --preset lighthouse\n```\n\n### Обычный seo аудит + Lighthouse\n``` bash\nsite-audit-seo -u https://example.com --lighthouse\n```\n\n## Как посчитать контент по csv\n1. Открыть в блокноте\n2. Документы посчитать поиском `,0`\n3. Листалки исключить поиском `?`\n4. Вычесть 1 (шапка)\n\n\n## Баги\n1. Иногда пишет в csv одинаковые страницы. Это бывает в 2 случаях:\n1.1. Редирект с другой страницы на эту (решается установкой `skipRequestedRedirect: true`, сделано).\n1.2. Одновременный запрос одной и той же страницы в параллельных потоках.\n\n\n## TODO:\n- Unique links\n- [Offline w3c validation](https://www.npmjs.com/package/html-validator)\n- [Words count](https://github.com/IonicaBizau/count-words)\n- [Sentences count](https://github.com/NaturalNode/natural)\n- Do not load image with non-standard URL, like [this](https://lh3.googleusercontent.com/pw/ACtC-3dd9Ng2Jdq713vsFqqTrNT6j_nyH3mFsRAzPbIAzWvDoRkiKSW2MIQOxrtpPVab4e9BElcL_Rlr8eGT68R7ZBnLCHpnHHJNRcd8JadddrxpVVClu1iOnkxPUQXOx-7OoNDmeEtH0xyg7NkEI8VF0oJRXQ=w1423-h1068-no?authuser=0)\n- External follow links\n- Broken images\n- Breadcrumbs - https://github.com/glitchdigital/structured-data-testing-tool\n- joeyguerra/schema.js - https://gist.github.com/joeyguerra/7740007\n- smhg/microdata-js - https://github.com/smhg/microdata-js\n- indicate page scan error\n- Find broken encoding like `СЂРµРіРёРѕРЅР°Р»СЊРЅРѕРіРѕ`\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fviasite%2Fsite-audit-seo","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fviasite%2Fsite-audit-seo","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fviasite%2Fsite-audit-seo/lists"}