{"id":20356546,"url":"https://github.com/madhurimarawat/web-scrapper-functions","last_synced_at":"2025-04-12T02:51:42.357Z","repository":{"id":210892289,"uuid":"727697397","full_name":"madhurimarawat/Web-Scrapper-Functions","owner":"madhurimarawat","description":"Streamlit-based Python web scraper for text, images, and PDFs. User-friendly interface for quick data extraction from websites. Simplify your web scraping tasks effortlessly.","archived":false,"fork":false,"pushed_at":"2024-11-30T13:33:32.000Z","size":149,"stargazers_count":8,"open_issues_count":0,"forks_count":3,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-03-25T22:35:58.145Z","etag":null,"topics":["automation","beautifulsoup","complete-pdf-text-data","complete-text-downloader","image-downloader-python","pdf-data-extraction","pdf-downloader","python","requests","streamlit-deployment","streamlit-webapp","text-data-website","text-file-rendering","user-input-link","web-scraper","web-scraping","web-scraping-automated","web-scraping-functions","zip-file-download","zip-file-rendering"],"latest_commit_sha":null,"homepage":"https://web-scrapper-functions-h6phqofpkjtaylwyn9uvzf.streamlit.app/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/madhurimarawat.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-12-05T11:47:22.000Z","updated_at":"2025-02-15T10:55:21.000Z","dependencies_parsed_at":"2024-01-15T13:18:08.853Z","dependency_job_id":"3fbb4529-218c-4050-8365-f1ad4fd7e48d","html_url":"https://github.com/madhurimarawat/Web-Scrapper-Functions","commit_stats":null,"previous_names":["madhurimarawat/web-scrapper-functions"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/madhurimarawat%2FWeb-Scrapper-Functions","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/madhurimarawat%2FWeb-Scrapper-Functions/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/madhurimarawat%2FWeb-Scrapper-Functions/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/madhurimarawat%2FWeb-Scrapper-Functions/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/madhurimarawat","download_url":"https://codeload.github.com/madhurimarawat/Web-Scrapper-Functions/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248509382,"owners_count":21116013,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["automation","beautifulsoup","complete-pdf-text-data","complete-text-downloader","image-downloader-python","pdf-data-extraction","pdf-downloader","python","requests","streamlit-deployment","streamlit-webapp","text-data-website","text-file-rendering","user-input-link","web-scraper","web-scraping","web-scraping-automated","web-scraping-functions","zip-file-download","zip-file-rendering"],"created_at":"2024-11-14T23:16:59.356Z","updated_at":"2025-04-12T02:51:42.349Z","avatar_url":"https://github.com/madhurimarawat.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Web-Scrapper-Functions\nStreamlit-based Python web scraper for text, images, and PDFs. User-friendly interface for quick data extraction from websites. Simplify your web scraping tasks effortlessly.\n\n\u003ca href = \"https://web-scrapper-functions-h6phqofpkjtaylwyn9uvzf.streamlit.app/\"\u003e\u003cimg width=\"927\" src=\"https://github.com/user-attachments/assets/f935ec1b-b973-49fe-9fae-5fc9340134f1\" title = \"Website Image\" alt=\"Website Image\"\u003e\u003c/a\u003e\n\u003cbr\u003e\u003cbr\u003e\n\n\u003ca href = \"https://web-scrapper-functions-h6phqofpkjtaylwyn9uvzf.streamlit.app/\"\u003e\u003cimg width=\"927\" src=\"https://github.com/user-attachments/assets/ac0346dc-d798-4e6b-8ce4-883abd46755c\" title = \"Website Image\" alt=\"Website Image\"\u003e\u003c/a\u003e\n\u003cbr\u003e\u003cbr\u003e\n\n\u003ca href = \"https://web-scrapper-functions-h6phqofpkjtaylwyn9uvzf.streamlit.app/\"\u003e\u003cimg width=\"927\" src=\"https://github.com/madhurimarawat/Web-Scrapper-Functions/assets/105432776/885e212d-42dc-4530-a4c3-f4a3c8df59ab\" title = \"Text File Image\" alt=\"Text File Image\"\u003e\u003c/a\u003e\n\n---\n# Mode of Execution Used \u003cimg src=\"https://th.bing.com/th/id/R.c936445e15a65dfdba20a63e14e7df39?rik=fqWqO9kKIVlK7g\u0026riu=http%3a%2f%2fassets.stickpng.com%2fimages%2f58481537cef1014c0b5e4968.png\u0026ehk=dtrTKn1QsJ3%2b2TFlSfLR%2fxHdNYHdrqqCUUs8voipcI8%3d\u0026risl=\u0026pid=ImgRaw\u0026r=0\" title=\"PyCharm\" alt=\"PyCharm\" width=\"40\" height=\"40\"\u003e\u0026nbsp;\u003cimg src=\"https://seeklogo.com/images/S/streamlit-logo-1A3B208AE4-seeklogo.com.png\" title=\"Streamlit\" alt=\"Streamlit\" width=\"40\" height=\"40\"\u003e\n\n\n\u003ch2\u003ePycharm\u003c/h2\u003e\n\n- Visit the official website of pycharm: \u003ca href=\"https://www.jetbrains.com/pycharm/\"\u003e\u003cimg src=\"https://th.bing.com/th/id/R.c936445e15a65dfdba20a63e14e7df39?rik=fqWqO9kKIVlK7g\u0026riu=http%3a%2f%2fassets.stickpng.com%2fimages%2f58481537cef1014c0b5e4968.png\u0026ehk=dtrTKn1QsJ3%2b2TFlSfLR%2fxHdNYHdrqqCUUs8voipcI8%3d\u0026risl=\u0026pid=ImgRaw\u0026r=0\" title=\"PyCharm\" alt=\"PyCharm\" width=\"40\" height=\"40\"\u003e\u003c/a\u003e\u003cbr\u003e\n- Download according to the platform that will be used like Linux, Macos or Windows.\u003cbr\u003e\nTwo versions of Pycharm are available:\u003cbr\u003e\n1. **Community version** \u003cbr\u003e\n- Community version is open source and we can use it for free without any paid plan.\u003cbr\u003e\n- We can download this at the end of pycharm website.\u003cbr\u003e\n- After downloading community version we can directly follow the setup wizard and it will be setup.\u003cbr\u003e\u003cbr\u003e\n2. **Professional Version**\u003cbr\u003e\n- This is available at the top of website, we can directly download from there.\u003cbr\u003e\n- After downloading professional version, follow the below steps.\u003cbr\u003e\n- Follow the setup wizard and sign up for the free version (trial version) or else continue with the premium or paid version.\u003cbr\u003e\n\n### Using Pycharm\n- First, in pycharm we have the concept of virtual environment. In virtual environment we can install all the required libraries or frameworks.\u003cbr\u003e\n- Each project has its own virtual environment, so thath we can install requirements like Libraries or Framworks for that project only.\u003cbr\u003e\n- After this we can create a new file, various file types are available in pycharm like script files, text files and also Jupyter Notebooks.\u003cbr\u003e\n- After selecting the required file type, we can continue the execution of that file by saving it and using this shortcut shift+F10 (In Windows).\u003cbr\u003e\n- Output is given in Console while installation happens in terminal in Pycharm.\n\n## Streamlit Server\n\n- Streamlit is a python framework through which we can deploy any machine learning model and any python project with ease and without worrying about the frontend.\u003cbr\u003e\n- Streamlit is very user-friendly.\u003cbr\u003e\n- Streamlit has pre defined functions for all frontend components and we can directly use them.\u003cbr\u003e\n- To install streamlit in your system, just run this command-\n\n```\npip install streamlit\n```\n\n## Running Project in Streamlit Server\n\u003cp\u003eMake Sure all dependencies are already satisfied before running the app.\u003c/p\u003e\n\n1. We can Directly run streamlit app  with the following command-\u003cbr\u003e\n```\nstreamlit run app.py\n```\nwhere app.py is the name of file containing streamlit code.\u003cbr\u003e\n\nBy default, streamlit will run on port 8501.\u003cbr\u003e\n\nAlso we can execute multiple files simultaneously and it will be executed in next ports like 8502 and so on.\u003cbr\u003e\n\n2. Navigate to URL http://localhost:8501\n\nYou should be able to view the homepage of your app.\n\n🌟 Project and Models will change but this process will remain the same for all Streamlit projects.\u003cbr\u003e\n\n## Deploying using Streamlit\n\n1. Visit the official website of streamlit : \u003ca href=\"https://streamlit.io/\"\u003e\u003cimg src=\"https://seeklogo.com/images/S/streamlit-logo-1A3B208AE4-seeklogo.com.png\" title=\"Streamlit\" alt=\"Streamlit\" width=\"40\" height=\"40\"\u003e\u003c/a\u003e \u003cbr\u003e\n2. Now make an account with GitHub.\u003cbr\u003e\n3. Now add all the code in Github repository.\u003cbr\u003e\n4. Go to streamlit and there is an option for new deployment.\u003cbr\u003e\n5. Type your Github repository name and specify the file name. If you name your file as streamlit_app it will directly access it else you have to specify the path.\u003cbr\u003e\n6. Now also make sure you upload all your libraries and requirement name in a requirement.txt file.\u003cbr\u003e\n7. Version can also be mentioned like this python==3.9.\u003cbr\u003e\n8. When we mention version in the requirement file streamlit install all dependencies from there.\u003cbr\u003e\n9. If everything went well our app will be deployed on web and you can share the link and access the app from all browsers.\n\n---\n\n## About Project :\n\n\u003cp\u003eComplete Description about the project and resources used.\u003c/p\u003e\n\n- **Embedded Links:**\n  Extracts and provides embedded links within a website.\u003cbr\u003e\n\n- **Main Website Text Data:**\n  Gathers and presents the primary textual content from the main website.\u003cbr\u003e\n\n- **Main Website Text Data along with Embedded Links Text Data:**\n  Combines main website text data with text data from embedded links.\u003cbr\u003e\n\n- **Complete Website Text Data:**\n  Retrieves and displays the entire textual content of the website.\u003cbr\u003e\n\n- **Extract Text from PDF Link:**\n  Retrieves and extract data of PDF file using the PDF Link Provided.\u003cbr\u003e\n\n- **Main Website PDF Data along with Embedded Links PDF Data:**\n Merge Text data extracted from PDF file in the main website with PDF data extracted from embedded links.\u003cbr\u003e\n\n- **Complete Website PDF Data:**\n  Captures and retrieves the PDF data available across the entire website.\u003cbr\u003e\n\n- **Complete Website Text and PDF Data:**\n  Presents a comprehensive view of both textual and PDF data from the entire website.\u003cbr\u003e\n\n- **Download PDF Files From Main Website:**\n  Facilitates the selective download of PDF files from the main website.\u003cbr\u003e\n\n- **Download All PDF Files From Website:**\n  Enables the bulk download of all PDF files available on the website.\u003cbr\u003e\n\n- **Download Image Files From Main Website:**\n  Allows the selective download of image files from the main website.\u003cbr\u003e\n\n- **Download All Image Files From Website:**\n  Supports the bulk download of all image files present on the website.\u003cbr\u003e\n\n- Visit Website from : \u003ca href=\"https://web-scrapper-functions-h6phqofpkjtaylwyn9uvzf.streamlit.app/\"\u003eWeb Scraper\u003c/a\u003e\n\n---\n## Libraries Used 📚 💻\n\u003cp\u003eShort Description about all libraries used.\u003c/p\u003e\nTo install python library this command is used- \u003cbr\u003e\u003cbr\u003e\n\n```\npip install library_name\n```\n\n\u003cul\u003e\n  \u003cli\u003e\u003cb\u003eStreamlit:\u003c/b\u003e Simplifies the creation of Python web applications, ideal for data scientists to effortlessly share interactive data visualizations.\u003c/li\u003e\n\n  \u003cli\u003e\u003cb\u003eRequests:\u003c/b\u003e Powerful Python module for HTTP requests, offering a streamlined API to integrate web services with ease.\u003c/li\u003e\n\n  \u003cli\u003e\u003cb\u003eLxml:\u003c/b\u003eLxml is a powerful and efficient Python library for processing XML and HTML documents, providing a comprehensive toolkit for parsing, validating, and manipulating structured data.\u003c/li\u003e\n\n  \u003cli\u003e\u003cb\u003eBeautiful Soup (bs4):\u003c/b\u003e Python library for web scraping tasks, facilitating the extraction of valuable information from HTML and XML documents.\u003c/li\u003e\n\n  \u003cli\u003e\u003cb\u003ePyPDF2:\u003c/b\u003e Focuses on handling PDF documents in Python, enabling tasks such as merging, splitting, and text/image extraction from PDF files.\u003c/li\u003e\n\n  \u003cli\u003e\u003cb\u003eio:\u003c/b\u003e The Python \u003ccode\u003eio\u003c/code\u003e module provides a versatile set of tools for managing input and output streams, supporting files, strings, and memory buffers.\u003c/li\u003e\n\n  \u003cli\u003e\u003cb\u003eZipfile:\u003c/b\u003e Python module for creating, extracting, and manipulating ZIP archives, essential for efficient data storage and transfer tasks.\u003c/li\u003e\n\u003c/ul\u003e\n\n---\n## Web Scraper Limitations\n\nThe web app has a limit of 100MB. Once this limit is reached, the app will not be able to scrape additional content. If you need to scrape more content, please run the app locally.\n\n\n---\n\n## Thanks for Visiting 😄\n\n- Drop a 🌟 if you find this repository useful.\u003cbr\u003e\u003cbr\u003e\n- If you have any doubts or suggestions, feel free to reach me.\u003cbr\u003e\u003cbr\u003e\n📫 How to reach me:  \u0026nbsp; [![Linkedin Badge](https://img.shields.io/badge/-madhurima-blue?style=flat\u0026logo=Linkedin\u0026logoColor=white)](https://www.linkedin.com/in/madhurima-rawat/) \u0026nbsp; \u0026nbsp;\n\u003ca href =\"mailto:rawatmadhurima@gmail.com\"\u003e\u003cimg src=\"https://github.com/madhurimarawat/Machine-Learning-Using-Python/assets/105432776/b6a0873a-e961-42c0-8fbf-ab65828c961a\" height=35 width=30 title=\"Mail Illustration\" alt=\"Mail Illustration📫\" \u003e \u003c/a\u003e\u003cbr\u003e\u003cbr\u003e\n- **Contribute and Discuss:** Feel free to open \u003ca href= \"https://github.com/madhurimarawat/Web-Scrapper-Functions/issues\"\u003eissues 🐛\u003c/a\u003e, submit \u003ca href = \"https://github.com/madhurimarawat/Web-Scrapper-Functions/pulls\"\u003epull requests 🛠️\u003c/a\u003e, or start \u003ca href = \"https://github.com/madhurimarawat/Web-Scrapper-Functions/discussions\"\u003ediscussions 💬\u003c/a\u003e to help improve this repository!\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmadhurimarawat%2Fweb-scrapper-functions","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmadhurimarawat%2Fweb-scrapper-functions","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmadhurimarawat%2Fweb-scrapper-functions/lists"}