Ecosyste.ms: Awesome
An open API service indexing awesome lists of open source software.
https://github.com/ayakashi-io/ayakashi
:zap: Ayakashi.io - The next generation web scraping framework
https://github.com/ayakashi-io/ayakashi
automation data-mining headless-chrome web-crawling web-scraping
Last synced: 6 days ago
JSON representation
:zap: Ayakashi.io - The next generation web scraping framework
- Host: GitHub
- URL: https://github.com/ayakashi-io/ayakashi
- Owner: ayakashi-io
- License: other
- Created: 2019-04-12T23:01:07.000Z (over 5 years ago)
- Default Branch: master
- Last Pushed: 2023-06-29T12:45:36.000Z (over 1 year ago)
- Last Synced: 2024-04-23T19:38:29.115Z (7 months ago)
- Topics: automation, data-mining, headless-chrome, web-crawling, web-scraping
- Language: TypeScript
- Homepage: https://ayakashi-io.github.io
- Size: 1.24 MB
- Stars: 197
- Watchers: 6
- Forks: 8
- Open Issues: 8
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
README
## The next generation web scraping framework
The web has changed. Gone are the days that raw html parsing scripts were the proper tool for the job.
Javascript and single page applications are now the norm.
Demand for data scraping and automation is higher than ever,
from business needs to data science and machine learning.
Our tools need to evolve.### Ayakashi helps you build scraping and automation systems that are
* easy to build
* simple or sophisticated
* highly performant
* maintainable and built for change### Powerful querying and data models
Ayakashi's way of finding things in the page and using them is done with [props](https://ayakashi-io.github.io/docs/guide/tour.html#props)
and [domQL](https://ayakashi-io.github.io/docs/guide/querying-with-domql.html).
Directly inspired by the relational database world (and SQL), domQL makes
DOM access easy and readable no matter how obscure the page's structure is.
Props are the way to package domQL expressions as re-usable structures which
can then be passed around to [actions](https://ayakashi-io.github.io/docs/guide/tour.html#actions) or to be used as models for [data
extraction](https://ayakashi-io.github.io/docs/guide/data-extraction.html).![domql](https://ayakashi-io.github.io/assets/img/domql.png)
### High level builtin actions
Ready made actions so you can focus on what matters.
Easily handle infinite scrolling, single page navigation, events
and [more](https://ayakashi-io.github.io/docs/reference/builtin-actions.html).
Plus, you can always [build your own actions](https://ayakashi-io.github.io/docs/advanced/creating-your-own-actions.html),
either from scratch or by composing other actions.### Preload code on pages
Need to include a bunch of code, a library you [made](https://ayakashi-io.github.io/docs/advanced/creating-your-own-preloaders.html)
or a [3rd party module](https://ayakashi-io.github.io/docs/going_deeper/loading-libraries-as-preloaders.html)
and make it available on a page?
[Preloaders](https://ayakashi-io.github.io/docs/guide/tour.html#preloaders) have you covered.### Control how you save your data
Automatically save your extracted data
to [all major SQL engines, JSON and CSV.](https://ayakashi-io.github.io/docs/guide/builtin-saving-scripts.html)
Need something more exotic or the ability to control exactly how the data is persisted?
Package and plug your custom logic as a script.### Manage the flow with pipelines
Scraping the data is only one part of the deal.
How about something like this:![pipelines](https://ayakashi-io.github.io/assets/img/diagram.png)
Need it to also be clean, readable and performant?
If so, [pipelines](https://ayakashi-io.github.io/docs/guide/tour.html#pipelines) can help.### Utilize all your cores
Ayakashi can utilize available cores as needed. Especially useful for projects that need
to run multiple operations in parallel.### Extend it as you like
All APIs used to build the builtin functionality are properly exposed.
All core entities are composable and extensible.### Use the language of the web
Many argue about javascript and its quirkiness as a language but the truth is:
If you want to scrape the web, you should speak its language.### Great editor support
Ayakashi comes bundled with a fully documented public API that you can explore
directly in your editor.
Autocomplete any method, check signatures and examples or follow links to more documentation.![editor support](https://ayakashi-io.github.io/assets/img/editor.png)
Sounds cool?
Just head over to the [getting started guide](https://ayakashi-io.github.io)!