{"id":13843745,"url":"https://github.com/p1g3/JSINFO-SCAN","last_synced_at":"2025-07-11T19:33:30.518Z","repository":{"id":43522846,"uuid":"193940777","full_name":"p1g3/JSINFO-SCAN","owner":"p1g3","description":"递归式寻找域名和api。","archived":false,"fork":false,"pushed_at":"2023-08-03T16:40:40.000Z","size":45,"stargazers_count":697,"open_issues_count":3,"forks_count":93,"subscribers_count":10,"default_branch":"master","last_synced_at":"2024-08-05T17:39:07.728Z","etag":null,"topics":["js-info","src-hunter"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/p1g3.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2019-06-26T16:26:58.000Z","updated_at":"2024-08-05T04:37:48.000Z","dependencies_parsed_at":"2022-07-18T18:42:12.487Z","dependency_job_id":null,"html_url":"https://github.com/p1g3/JSINFO-SCAN","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/p1g3%2FJSINFO-SCAN","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/p1g3%2FJSINFO-SCAN/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/p1g3%2FJSINFO-SCAN/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/p1g3%2FJSINFO-SCAN/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/p1g3","download_url":"https://codeload.github.com/p1g3/JSINFO-SCAN/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":225755125,"owners_count":17519206,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["js-info","src-hunter"],"created_at":"2024-08-04T17:02:26.065Z","updated_at":"2024-11-21T15:31:28.640Z","avatar_url":"https://github.com/p1g3.png","language":"Python","funding_links":[],"categories":["Python","Python (1887)","信息搜集","others"],"sub_categories":[],"readme":"# JSINFO-SCAN\r\n\r\n### 前言\r\n\r\n很早以前就想写一款对网站进行爬取，并且对网站中引入的JS进行信息搜集的一个工具，之前一直没有思路，因为对正则的熟悉程度没有到可以对js中的info进行匹配的地步，最近实验室的朋友写了一款工具：[JSFinder](https://github.com/Threezh1/JSFinder \"JSFinder\")，借用了他的思路，写了一款递归爬取域名(netloc/domain)，以及递归从JS中获取信息的工具。\r\n\r\n### 思路\r\n\r\n写这个工具主要有这么几个点：\r\n\r\n- 如何爬取域名\r\n- 如何爬取JS\r\n- 如何从JS中获取信息\r\n\r\n三个点的解决方案：\r\n\r\n- 正则匹配href属性中的链接\r\n- 正则匹配src属性中的链接，并判断链接是否为js。正则匹配`\u003cscript\u003e\u003c/script\u003e`标签中的文本。\r\n-   使用LinkFinder的正则从JS中匹配敏感信息\r\n\r\n### 优点\r\n\r\n- 实现递归爬取\r\n- 对JS中匹配到的info进行了处理，直观的展示出来\r\n\r\n### 缺点\r\n\r\n使用的是单线程，因为目前多线程还没学够，如果冒昧使用担心引起数据混乱的问题，也并不熟悉使用Lock()函数，怕使用的地方多了，多线程也变成了单线程。\r\n\r\n### Usage\r\n\r\n```\r\npython3 jsinfo.py -d jd.com --keyword jd --save jd.api.txt --savedomain jd.domain.txt\r\n```\r\n\r\n- -d/-f\r\n\r\n对单个域名或对域名文件进行扫描，文件需要一行一个域名。\r\n\r\n- --keyword\r\n\r\n设置爬取关键字，会使用该关键字对域名进行匹配，必选项。\r\n\r\n- --savedomain\r\n\r\n设置爬取出来的域名保存路径\r\n\r\n- --save\r\n\r\n设置api保存路径\r\n\r\n### 实例\r\n\r\n- 对京东进行爬取\r\n\r\n只爬取jd.com，设置keyword为jd,joybuy,360：\r\n\r\n![enter image description here](https://s2.ax1x.com/2019/06/27/Zma4Tf.png)\r\n\r\n- 对百度进行爬取\r\n\r\n![enter image description here](https://s2.ax1x.com/2019/06/27/Zmablj.png)\r\n\r\n### Update\r\n\r\n2019-7-5：重构代码，加入了爬行深度的设定，深度为1~2，默认为1，2即为深度爬取，同时增加了url的存储，即深度爬取爬取到的链接。\r\n\r\n2020-2-10：重构代码，整体使用了协程，使用队列的方式作为递归标准，默认递归深度为8，可根据自身需要进行修改。\r\n\r\n经过测试，速度是v1版本的十倍不止，并且获取到的域名也是v1版本的两倍，但是这一版取消了获取api。\r\n\r\n效果：\r\n\r\n![](https://s2.ax1x.com/2020/02/10/1ImiUe.jpg)\r\n\r\n\r\n2020-2-16：增加搜集其他信息的功能，比如邮箱，代码作者，ip等。\r\n\r\n效果：\r\n\r\n![3Ufg5n.png](https://s2.ax1x.com/2020/02/26/3Ufg5n.png)\r\n\r\n2020-8-1：重构代码，具体看下面的README。\r\n\r\n#### JSINFO流程图\r\n\r\n```bash\r\n_____  ___    _  _   _  ___    _____ \r\n(___  )(  _`\\ (_)( ) ( )(  _`\\ (  _  )\r\n    | || (_(_)| || `\\| || (_(_)| ( ) |\r\n _  | |`\\__ \\ | || , ` ||  _)  | | | |\r\n( )_| |( )_) || || |`\\ || |    | (_) |\r\n`\\___/'`\\____)(_)(_) (_)(_)    (_____)\r\n        Author：P1g3#p1g3cyx@gmail.com\r\n```\r\n\r\n![jsinfo.jpg](https://i.loli.net/2020/08/01/QRMeW2HABCxVamk.png)\r\n\r\n目的：\r\n\r\n- 扩充资产（主要体现在爬取根域名这块）\r\n- 爬取敏感信息（新增了大量敏感信息正则，对邮箱的提取进行了优化）\r\n- 爬取api（将前面版本中去除的api功能恢复）\r\n\r\n##### update\r\n\r\n- 新增banner信息\r\n- 新增敏感信息正则\r\n- 代码优化\r\n- 对Ctrl+C退出进行优化（当退出时，会自动将已爬取到的信息保存到当前目录下）\r\n- 无需自动指定输出结果，最终输出为四个文件，为str(int(time.time()))_xxx（xxx为root_domains、sub_domains、leak_infos、apis）\r\n- 错误处理优化\r\n- 减少传入参数\r\n\r\n##### Usage\r\n\r\n```\r\npython3 jsinfo.py --target www.baidu.com --keywords baidu\r\n```\r\n\r\n- target（域名 ==\u003e 可传入单个域名或域名文件）\r\n- keywords（域名中的关键字，用于搜集根域名以及扩充子域名）\r\n- black_keywords（黑名单关键字，当返回包中含有这些关键字则不再进行二次爬取，用于某些商城页面避免爬到无用链接）\r\n\r\n##### 使用效果\r\n\r\n![image-20200801011132874.png](https://i.loli.net/2020/08/01/e5rxjcQCFhLBdlU.png)\r\n\r\n\r\n\r\n阿里的资产还在跑，目前获取到了如下信息：\r\n\r\n- 2k+ 子域\r\n- 80+ 根域\r\n- 80000+ api\r\n- 5000+ 敏感信息\r\n\r\nPS：欢迎反馈Bug，本项目将持续更新，如有问题请联系wx：p1g3___，如果需要下载历史版本的jsinfo，请从commit中寻找...\r\n\r\n##### 部分正则来源\r\n\r\n- https://github.com/m4ll0k/SecretFinder/blob/master/SecretFinder.py\r\n- https://github.com/GerbenJavado/LinkFinder","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fp1g3%2FJSINFO-SCAN","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fp1g3%2FJSINFO-SCAN","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fp1g3%2FJSINFO-SCAN/lists"}