{"id":13715375,"url":"https://github.com/soloice/Chinese-Character-Recognition","last_synced_at":"2025-05-07T04:30:44.621Z","repository":{"id":83994407,"uuid":"81660663","full_name":"soloice/Chinese-Character-Recognition","owner":"soloice","description":"This project shows how to use CNN to perform Chinese character recognition, a much more complicated task compared to MNIST digit recognition.","archived":false,"fork":false,"pushed_at":"2017-05-06T06:41:55.000Z","size":190,"stargazers_count":200,"open_issues_count":1,"forks_count":77,"subscribers_count":12,"default_branch":"master","last_synced_at":"2024-11-14T03:34:28.631Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/soloice.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2017-02-11T15:10:33.000Z","updated_at":"2024-11-13T10:06:37.000Z","dependencies_parsed_at":null,"dependency_job_id":"ee947e61-c535-4312-861c-2fef62b68284","html_url":"https://github.com/soloice/Chinese-Character-Recognition","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soloice%2FChinese-Character-Recognition","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soloice%2FChinese-Character-Recognition/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soloice%2FChinese-Character-Recognition/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soloice%2FChinese-Character-Recognition/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/soloice","download_url":"https://codeload.github.com/soloice/Chinese-Character-Recognition/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":252813646,"owners_count":21808362,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-03T00:00:58.246Z","updated_at":"2025-05-07T04:30:44.275Z","avatar_url":"https://github.com/soloice.png","language":"Python","funding_links":[],"categories":["Łamanie"],"sub_categories":["Chińskie"],"readme":"# Chinese-Character-Recognition\nThis tensorflow project shows how to use CNN to perform Chinese character recognition, a much more complicated task compared to digit recognition.\n\nThis code is based on a similar [repo](https://github.com/burness/tensorflow-101/tree/master/chinese_hand_write_rec/src), but is cleaner and outperforms that one: This code defines a data iterator class to deal with the data preprocessing and use 2 different input pipeline for training and testing data. \n\nMoreover, Training 16000 steps (~30 hour on my workstation with 8 CPU cores and 32 GB RAM) achieves a top 1 accuracy of 92.50% and a top 3 accuracy of 97.48%. If you have a GPU card, it should be much more faster.\n\nThe repo owner above also kindly shared the [preprocessed dataset](https://pan.baidu.com/s/1o84jIrg#list/path=%2F)\n\nFor training, run:\n---------------\n`python chinese_character_recognition_bn.py --mode=train --max_steps=16002 --eval_steps=100 --save_steps=500`\n\nFor validation, run:\n--------------\n`python chinese_character_recognition_bn.py --mode=validation`\n\n\nNote that the network hasn't fully convergenced yet after 16000 mini-batches of training (though almost), so training for a longer time should be able to improve the performance furthur. As reported in [this project report](http://cs231n.stanford.edu/reports/zyh_project.pdf), the network has the potential to achieve a top 1 accuracy of 95% when properly trained (maybe even better, since in the aforementioned paper they didn't use batch normalization, which should improve generalization capacity of the network). So if you have a powerful GPU, run more training steps.\n\nThis is [my checkpoint](https://pan.baidu.com/s/1o7CJrBW) at step 16000. If you want to try it on your laptop or something else, feel free to do so!\n\n\nI also attached the learning curve of `chinese_character_recognition_bn.py`:\n\n![accuracy-bn](https://github.com/soloice/Chinese-Character-Recognition/blob/master/accuracy-bn-16k.PNG)\n\n![loss-bn](https://github.com/soloice/Chinese-Character-Recognition/blob/master/loss-bn-16k.PNG)\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsoloice%2FChinese-Character-Recognition","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsoloice%2FChinese-Character-Recognition","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsoloice%2FChinese-Character-Recognition/lists"}