{"id":23418544,"url":"https://github.com/tezcatlipoca0000/trevi-spider","last_synced_at":"2026-04-30T14:36:51.866Z","repository":{"id":223136468,"uuid":"759407655","full_name":"Tezcatlipoca0000/trevi-spider","owner":"Tezcatlipoca0000","description":"A tailored program that scrapes data from the website and updates my DB","archived":false,"fork":false,"pushed_at":"2024-09-07T14:16:43.000Z","size":29,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-02-15T02:30:53.208Z","etag":null,"topics":["gmail-api","pandas","python3","scraping-python","selenium-python"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Tezcatlipoca0000.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-02-18T14:12:45.000Z","updated_at":"2024-09-07T14:16:47.000Z","dependencies_parsed_at":"2024-09-07T15:40:43.631Z","dependency_job_id":"f0cca495-8d1b-4992-9bd3-523a1e3a03c3","html_url":"https://github.com/Tezcatlipoca0000/trevi-spider","commit_stats":null,"previous_names":["tezcatlipoca0000/trevi-spider"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tezcatlipoca0000%2Ftrevi-spider","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tezcatlipoca0000%2Ftrevi-spider/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tezcatlipoca0000%2Ftrevi-spider/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tezcatlipoca0000%2Ftrevi-spider/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Tezcatlipoca0000","download_url":"https://codeload.github.com/Tezcatlipoca0000/trevi-spider/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248000094,"owners_count":21031102,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["gmail-api","pandas","python3","scraping-python","selenium-python"],"created_at":"2024-12-23T00:20:14.788Z","updated_at":"2026-04-30T14:36:46.835Z","avatar_url":"https://github.com/Tezcatlipoca0000.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Trevi Spider\r\n\r\n## What it does:\r\nIt can execute three distinct functions to populate the local business database with carefully extracted data from the provider's website:\r\n\t- It scrapes the price data from the provider's website.\r\n\t- Retrieves the placed ordered via email to the provider. \r\n\t- Creates a dataframe with the new data when available.  \r\n\r\n## What technologies does it need:\r\n\t- Python\r\n \t- google-api-python-client \r\n \t- google-auth-httplib2 \r\n \t- google-auth-oauthlib\r\n \t- numpy\r\n  \t- Selenium\r\n   \t- Pandas\r\n   \t- openpyxl \r\n\r\n\r\n## What files does it need (stdin):\r\n\t- To get placed-order:\r\n\t\t- token.json ~ GMAIL API **git ignored**\r\n\t\t- credentials.json ~ GMAIL.API **git ignored**\r\n\t\t- mygmailaccount.inbox.message_with_pedido.xlsx\r\n\r\n\t- To update:\r\n\t\t- Provedores Todos.xlsm ~ Local database **git ignored**\r\n\t\t- pedido.xlsx ~ Placed-order to provider, retrieved from GMAIL **git ignored**\r\n\t\t- trevi_full.xlsx ~ Scraped data from provider's website **git ignored**\r\n\r\n\r\n## What files does it create (stdout):\r\n\t- After scraping the data:\r\n\t\t- trevi_full.xlsx ~ Scraped data from provider's website **git ignored**\r\n\r\n\t- After retrieving placed-order:\r\n\t\t- pedido.xlsx ~ Placed-order to provider, retrieved from GMAIL **git ignored**\r\n\r\n\t- After creating an updtaded dataframe:\r\n\t\t- Final.xlsx ~ Updated dataframe with necesary information **git ignored**\r\n\r\n## What I'm thinking about:\r\n\t- Uploading all the necessary sample_files. \r\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftezcatlipoca0000%2Ftrevi-spider","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ftezcatlipoca0000%2Ftrevi-spider","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftezcatlipoca0000%2Ftrevi-spider/lists"}