{"id":18243779,"url":"https://github.com/paulchen2713/scrap-nstc-html-files","last_synced_at":"2026-07-27T07:31:39.842Z","repository":{"id":261193078,"uuid":"883555737","full_name":"paulchen2713/Scrap-NSTC-HTML-Files","owner":"paulchen2713","description":"從國科會網站 (.aspx) 找清大每位教師的補助研究計畫資料 (.html)，抓取 年度、姓名、系所、計畫名稱、執行年限、金額 等資訊整理成一個檔案。","archived":false,"fork":false,"pushed_at":"2024-11-05T07:24:32.000Z","size":9471,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-02-14T14:43:28.864Z","etag":null,"topics":["html-parser","nstc","nthu","python","scraping-data"],"latest_commit_sha":null,"homepage":"","language":"HTML","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/paulchen2713.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-11-05T07:08:22.000Z","updated_at":"2024-11-05T07:24:34.000Z","dependencies_parsed_at":"2024-11-05T08:29:18.984Z","dependency_job_id":null,"html_url":"https://github.com/paulchen2713/Scrap-NSTC-HTML-Files","commit_stats":null,"previous_names":["paulchen2713/scrap-nstc-html-files"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/paulchen2713%2FScrap-NSTC-HTML-Files","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/paulchen2713%2FScrap-NSTC-HTML-Files/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/paulchen2713%2FScrap-NSTC-HTML-Files/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/paulchen2713%2FScrap-NSTC-HTML-Files/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/paulchen2713","download_url":"https://codeload.github.com/paulchen2713/Scrap-NSTC-HTML-Files/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247900855,"owners_count":21015139,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["html-parser","nstc","nthu","python","scraping-data"],"created_at":"2024-11-05T09:03:03.850Z","updated_at":"2025-10-14T17:07:11.761Z","avatar_url":"https://github.com/paulchen2713.png","language":"HTML","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Scrap-NSTC-HTML-Files\n從[國科會網站](https://wsts.nstc.gov.tw/STSWeb/Award/AwardMultiQuery.aspx) (.aspx) 找清大每位教師的[國家科學及技術委員會補助研究計畫資料](https://wsts.nstc.gov.tw/STSWeb/Award/AwardMultiQuery.aspx?year=107\u0026code=QS01\u0026organ=A%2CFA04%2C\u0026name=) (.html)，抓取 107-111 年度、姓名、系所、計畫名稱、執行年限、金額 等資訊整理成一個檔案。\n\n註: 這不是爬蟲，只是從靜態 html 網頁內容把想要的資料撈出來而已，且要自己手動 Ctrl + Shift + C 把每 1 頁 (1 頁 200 筆) 共 14 頁的 html 網頁內容存下來。\n\n\n### Tag examples\n```htmlembedded\n\u003ctr class=\"Grid_Row\"\u003e\n    \u003ctd align=\"center\"\u003e110\u003c/td\u003e\n    \u003ctd align=\"left\"\u003e鍾偉和\u003c/td\u003e\n    \u003ctd align=\"left\"\u003e國立清華大學通訊工程研究所\u003c/td\u003e\n    \u003ctd align=\"left\"\u003e\n        \u003cspan id=\"wUctlAwardQueryPage_grdResult_lblAWARD_PLAN_CHI_DESCc_198\"\u003e運用機器學習於巨量多天線傳輸系統之設計\u003c/span\u003e\u003cbr\u003e\n        \u003cspan id=\"wUctlAwardQueryPage_grdResult_lblAWARD_ST_ENDc_198\"\u003e2021/08/01~2024/07/31\u003c/span\u003e\u003cbr\u003e\n        \u003cspan id=\"wUctlAwardQueryPage_grdResult_lblAWARD_TOT_AUD_AMTc_198\"\u003e3,036,000元\u003c/span\u003e\u0026nbsp;\n```\n```htmlembedded\n\u003ctr class=\"Grid_AlternatingRow\"\u003e\n    \u003ctd align=\"center\"\u003e107\u003c/td\u003e\n    \u003ctd align=\"left\"\u003e鍾偉和\u003c/td\u003e\n    \u003ctd align=\"left\"\u003e國立清華大學電機工程學系(所)\u003c/td\u003e\n    \u003ctd align=\"left\"\u003e\n        \u003cspan id=\"wUctlAwardQueryPage_grdResult_lblAWARD_PLAN_CHI_DESCc_191\"\u003e適用於智慧型物聯人聯網中之多天線系統訊號處理\u003c/span\u003e\u003cbr\u003e\n        \u003cspan id=\"wUctlAwardQueryPage_grdResult_lblAWARD_ST_ENDc_191\"\u003e2018/08/01~2021/10/31\u003c/span\u003e\u003cbr\u003e\n        \u003cspan id=\"wUctlAwardQueryPage_grdResult_lblAWARD_TOT_AUD_AMTc_191\"\u003e2,598,000元\u003c/span\u003e\u0026nbsp;\n```\n\n### Sample code\n```python\nimport os\nimport csv\nimport openpyxl\nfrom openpyxl import Workbook\nfrom bs4 import BeautifulSoup\n\n# Prepare the Excel workbook\nwb = Workbook()\nws = wb.active\nws.title = \"Extracted Data\"\nws.append(['Fiscal Year', 'Professor Name', 'Department', 'Project Name', 'Project Duration', 'Project Cost'])\n\n# Prepare CSV output\ncsv_data = []\ncsv_data.append(['Fiscal Year', 'Professor Name', 'Department', 'Project Name', 'Project Duration', 'Project Cost'])\n\n# Get all HTML filenames from the data folder\ndata_folder = 'data'\nhtml_files = [f for f in os.listdir(data_folder) if f.endswith('.html')]\n\n# Iterate through all HTML files in the data folder\nfor file_name in html_files:\n    file_path = os.path.join(data_folder, file_name)\n    with open(file_path, 'r', encoding='utf-8') as file:\n        soup = BeautifulSoup(file, 'html.parser')\n\n    # Find all rows of the table\n    rows = soup.find_all('tr', class_=['Grid_AlternatingRow', 'Grid_Row'])\n\n    # Iterate through each row and extract the required information\n    for row in rows:\n        fiscal_year = row.find_all('td')[0].get_text(strip=True)\n        professor_name = row.find_all('td')[1].get_text(strip=True)\n        department = row.find_all('td')[2].get_text(strip=True)\n        project_name = row.find('span', id=lambda x: x and 'lblAWARD_PLAN_CHI_DESCc' in x).get_text(strip=True).replace('\\n', '').replace('\\r', '').replace('\\t', '').replace(' ', '')\n        project_duration = row.find('span', id=lambda x: x and 'lblAWARD_ST_ENDc' in x).get_text(strip=True)\n        project_cost = row.find('span', id=lambda x: x and 'lblAWARD_TOT_AUD_AMTc' in x).get_text(strip=True)\n\n        # Write to Excel\n        ws.append([fiscal_year, professor_name, department, project_name, project_duration, project_cost])\n\n        # Add to CSV data\n        csv_data.append([fiscal_year, professor_name, department, project_name, project_duration, project_cost])\n\n# Set the output file name\nexcel_filename = \"combined_extracted_data\"\ncsv_filename = excel_filename\n\n# Save the Excel file\nif os.path.isfile(f'{excel_filename}.xlsx') is False:\n    wb.save(f'{excel_filename}.xlsx')\n\n# Save the CSV file\nif os.path.isfile(f'{csv_filename}.csv') is False:\n    with open(f'{csv_filename}.csv', 'w', newline='', encoding='utf-8-sig') as csvfile:\n        writer = csv.writer(csvfile)\n        writer.writerows(csv_data)\n\nprint(f\"Data extraction complete. Check '{excel_filename}.xlsx' and '{csv_filename}.csv' for the output.\")\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpaulchen2713%2Fscrap-nstc-html-files","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fpaulchen2713%2Fscrap-nstc-html-files","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpaulchen2713%2Fscrap-nstc-html-files/lists"}