{"id":41195763,"url":"https://github.com/liserjrqlxue/anno2xlsx","last_synced_at":"2026-01-22T20:36:09.861Z","repository":{"id":37750627,"uuid":"165582818","full_name":"liserjrqlxue/anno2xlsx","owner":"liserjrqlxue","description":"汇总临床全外分析结果，进行部分注释处理，生成下游解读系统需要的表格文件","archived":false,"fork":false,"pushed_at":"2023-07-05T06:02:10.000Z","size":2480,"stargazers_count":5,"open_issues_count":0,"forks_count":2,"subscribers_count":2,"default_branch":"master","last_synced_at":"2024-06-21T00:13:07.883Z","etag":null,"topics":["acmg","aes-json","annotate","xlsx"],"latest_commit_sha":null,"homepage":"","language":"Go","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/liserjrqlxue.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-01-14T02:33:56.000Z","updated_at":"2024-02-29T07:24:18.000Z","dependencies_parsed_at":"2024-06-19T17:11:15.540Z","dependency_job_id":"d1ea1199-a158-4667-907f-ba937ac302b5","html_url":"https://github.com/liserjrqlxue/anno2xlsx","commit_stats":null,"previous_names":[],"tags_count":79,"template":false,"template_full_name":null,"purl":"pkg:github/liserjrqlxue/anno2xlsx","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liserjrqlxue%2Fanno2xlsx","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liserjrqlxue%2Fanno2xlsx/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liserjrqlxue%2Fanno2xlsx/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liserjrqlxue%2Fanno2xlsx/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/liserjrqlxue","download_url":"https://codeload.github.com/liserjrqlxue/anno2xlsx/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liserjrqlxue%2Fanno2xlsx/sbom","scorecard":{"id":592391,"data":{"date":"2025-08-11","repo":{"name":"github.com/liserjrqlxue/anno2xlsx","commit":"4869139fd7be2adbc21f4866bb57199d3f4da43a"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":3.8,"checks":[{"name":"Code-Review","score":0,"reason":"Found 0/1 approved changesets -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Dangerous-Workflow","score":10,"reason":"no dangerous workflow patterns detected","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Token-Permissions","score":0,"reason":"detected GitHub workflow tokens with excessive permissions","details":["Warn: no topLevel permission defined: .github/workflows/gitlab.rsync.yml:1","Info: no jobLevel write permissions found"],"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/gitlab.rsync.yml:12: update your workflow using https://app.stepsecurity.io/secureworkflow/liserjrqlxue/anno2xlsx/gitlab.rsync.yml/master?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/gitlab.rsync.yml:15: update your workflow using https://app.stepsecurity.io/secureworkflow/liserjrqlxue/anno2xlsx/gitlab.rsync.yml/master?enable=pin","Info:   0 out of   1 GitHub-owned GitHubAction dependencies pinned","Info:   0 out of   1 third-party GitHubAction dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: GNU General Public License v3.0: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":-1,"reason":"internal error: error during branchesHandler.setup: internal error: githubv4.Query: Resource not accessible by integration","details":null,"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 30 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}},{"name":"Vulnerabilities","score":10,"reason":"0 existing vulnerabilities detected","details":null,"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}}]},"last_synced_at":"2025-08-20T22:17:53.110Z","repository_id":37750627,"created_at":"2025-08-20T22:17:53.110Z","updated_at":"2025-08-20T22:17:53.110Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28670385,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-22T19:36:09.361Z","status":"ssl_error","status_checked_at":"2026-01-22T19:36:05.567Z","response_time":144,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["acmg","aes-json","annotate","xlsx"],"created_at":"2026-01-22T20:36:09.209Z","updated_at":"2026-01-22T20:36:09.848Z","avatar_url":"https://github.com/liserjrqlxue.png","language":"Go","funding_links":[],"categories":[],"sub_categories":[],"readme":"# anno2xlsx \u003c!-- omit in toc --\u003e\n\n[![Gitpod ready-to-code](https://img.shields.io/badge/Gitpod-ready--to--code-blue?logo=gitpod)](https://gitpod.io/#https://github.com/liserjrqlxue/anno2xlsx)\n[![GoDoc](https://godoc.org/github.com/liserjrqlxue/anno2xlsx?status.svg)](https://pkg.go.dev/github.com/liserjrqlxue/anno2xlsx)\n[![Go Report Card](https://goreportcard.com/badge/github.com/liserjrqlxue/anno2xlsx)](https://goreportcard.com/report/github.com/liserjrqlxue/anno2xlsx)\n\n- [USAGE](#usage)\n  - [PARAM](#param)\n- [处理逻辑](#处理逻辑)\n  - [CNV](#cnv)\n  - [Extra](#extra)\n  - [fam\\_info](#fam_info)\n  - [FV](#fv)\n    - [遗传相符](#遗传相符)\n    - [familyTag](#familytag)\n    - [筛选标签](#筛选标签)\n  - [LOH](#loh)\n  - [QC](#qc)\n- [AES加密数据库](#aes加密数据库)\n  - [DATABASE](#database)\n  - [疾病库/基因频谱](#疾病库基因频谱)\n  - [ACMG Secondary Finding](#acmg-secondary-finding)\n    - [注释校验](#注释校验)\n  - [VIPHL](#viphl)\n- [CHPO](#chpo)\n- [注意](#注意)\n  - [特殊位点库](#特殊位点库)\n- [特性](#特性)\n  - [结合性处理](#结合性处理)\n    - [格式转换](#格式转换)\n    - [Het-\\\u003eHom修正](#het-hom修正)\n    - [Hom-\\\u003eHemi修正](#hom-hemi修正)\n- [复用拼接字段](#复用拼接字段)\n- [注意](#注意-1)\n- [UTIL](#util)\n  - [`tier1tags`](#tier1tags)\n    - [TO-DO](#to-do)\n\n## USAGE\n\n### PARAM\n\n| arg          | type    | example                                         | note                                                                                 |\n|--------------|---------|-------------------------------------------------|--------------------------------------------------------------------------------------|\n| -json        | boolean |                                                 | 输出json格式结果                                                                           |\n| -save        | boolean |                                                 | 保存excel                                                                              |\n| -academic    | boolean |                                                 | 学术使用，比如REVEL数据库注释                                                                    |\n| -hl          | boolean |                                                 | 使用耳聋变异库                                                                              |\n| -nb          | boolean |                                                 | 使用新筛变异库                                                                              |\n| -pp          | boolean |                                                 | 使用孕前变异库                                                                              |\n| -sf          | boolean |                                                 | 使用ACMG SF变异库                                                                         |\n| -acmg        | boolean |                                                 | 使用ACMG2015计算证据项PVS1, PS1,PS4, PM1,PM2,PM4,PM5 PP2,PP3, BA1, BS1,BS2, BP1,BP3,BP4,BP7 |\n| -autoPVS1    | boolean |                                                 | 使用autoPVS1结果处理证据项PVS1                                                                |\n| -allTier1    | boolean |                                                 | 不进行tier1过滤                                                                           |\n| -allgene     | boolean |                                                 | tier1过滤不过滤基因                                                                         |\n| -cnvAnnot    | boolean |                                                 | 重新进行UpdateCnvAnnot                                                                   |\n| -cnvFlter    | boolean |                                                 | 进行 cnv 结果过滤                                                                          |\n| -warn        | boolean |                                                 | 警告基因名无法识别问题，而非中断                                                                     |\n| -wgs         | boolean |                                                 | wgs模式                                                                                |\n| -couple      | boolean |                                                 | 夫妻模式                                                                                 |\n| -trio        | boolean |                                                 | 标准trio模式                                                                             |\n| -trio2       | boolean |                                                 | 非标准trio，但是保持先证者、父亲、母亲顺序                                                              |\n| -redis       | boolean |                                                 | 使用redis服务注释本地频率                                                                      |\n| -redisAddr   | string  | 127.0.0.1:6380                                  | redis服务器地址                                                                           |\n| -cfg         | string  | etc/config.toml                                 | toml配置文件                                                                             |\n| -geneId      | string  | db/gene.id.txt                                  | 基因名-基因ID 对应数据库                                                                       |\n| -list        | string  | sample1,sample2,sample3                         | 样品编号，逗号分割，**有顺序**                                                                    |\n| -gender      | string  | M,M,F                                           | 样品性别，逗号分割，与 -list 顺序一致                                                               |\n| -qc          | string  | sample1.coverage.report,sample2.coverage.report | bamdst质控文件，逗号分割，与 -list 顺序一致                                                         |\n| -filterStat  | string  | L01.filter.stat,L02.filter.stat                 | 计算reads QC的文件，逗号分割                                                                   |\n| -imqc        | string  | sample1.QC.txt,sample2.QC.txt                   | 一体机QC.txt格式QC输入，逗号分割，过滤 -list 内样品列表                                                  |\n| -mtqc        | string  | sample1.MT.QC.txt,sample2.MT.QC.txt             | 线粒体QC.txt，逗号分割，过滤 -list 内样品列表                                                        |\n| -karyotype   | string  | sample1.karyotpye.txt,sample2.karyotype.txt     | 核型信息，逗号分割                                                                            |\n| -loh         | string  | loh1.xlsx,loh2.xlsx                             | loh结果excel，逗号分割，按 -list 样品编号顺序创建sheet                                                |\n| -lohSheet    | string  | LOH_annotation                                  | sheet name                                                                           |\n| -kinship     | string  | kinship.txt                                     | trio的亲缘关系                                                                            |\n| -snv         | string  | snv1.txt,snv2.txt                               | snv注释结果，逗号分割                                                                         |\n| -exon        | string  | sample1.exon.txt,sample2.exon.txt               | exon CNV 输入文件，逗号分割，过滤 -list 内样品列表                                                    |\n| -large       | string  | sample1.large.txt,sample2.large.txt             | large CNV注释结果，逗号分割                                                                   |\n| -extra       | string  | extra1.txt,extra2.txt                           | 额外简单放入excel的额外sheets中                                                                |\n| -extraSheet  | string  | sheet1,sheet2                                   | -extra对应sheet name                                                                   |\n| -prefix      | string  | outputPrefix                                    | 输出前缀，默认 -snv 第一个输入                                                                   |\n| -log         | string  | prefix.log                                      | log输出文件                                                                              |\n| -tag         | string  | .tag                                            | tier1结果文件名加入额外标签，[prefix].Tier1[tag].xlsx                                            |\n| -product     | string  | DX1516                                          | 产品编号                                                                                 |\n| -seqType     | string  | SEQ2000                                         | redis 查询关键词，区分频率库                                                                    |\n| -specVarList | string  | etc/spec.var.lite.txt                           | 特殊变异库                                                                                |\n\n#### `-trio`\n\n- `-wesim` 模式 输出 `.result.tsv` 时 额外 处理\n    - 拆分 `Zygosity`\n    - 额外 输出 两列: `Genotype of Family Member 1` `Genotype of Family Member 2`\n- 额外 `Tier1` 统计\n- 额外 `Tier1` 判断规则\n- 额外 `遗传相符` 判断规则\n- 与 `-trio2` 相同\n    - 额外 `筛选标签` 标签规则\n    - `变异来源`\n    - `familyTag` 标签规则\n\n## 处理逻辑\n\n### CNV\n\n1. 读取 `annotation.Gene.transcript` , 得到基因-转录本对应关系\n2. 读取 `CNV` 数据文件列表到 `paths`\n3. 使用 `addCnv2Sheet(sheet, title, paths, sampleMap, filterSize, filterGene, stats, key, gender, cnvFile)` 写入 sheet\n    - 读取 `paths` 到 `cnvDb`，遍历 `item`\n        1. 跳过其他样品编号\n      2. 一体机模式，进行表头处理\n      3. `anno.CnvPrimer`进行引物设计\n      4. `item[\"OMIM_Gene\"]` 作为 基因，获取 `geneIDs`\n      5. 基于 `geneIDs` 注释 `CHPO` `基因-疾病` `突变频谱`\n      6. `-cnvAnnot` 时\n         1. 转换 `json` 格式写入 `item[\"CNV_annot\"]`\n      7. `item[\"OMIM_Phenotype_ID\"]` -\u003e `item[\"OMIM\"]`\n      8. `-cnvFilter` 时\n         1. `exon cnv` 跳过 `item[\"OMIM\"]` 为空\n         2. `large cnv` 跳过 `长度 \u003c 1M`\n      9. 基于 `title` 写入 `sheet row`\n      10. `-wesim` 时\n         1. 基于 `wesim.cnvColumn` 写入 `cnvFile`\n\n### Extra\n\n1. 解析 `-extra` 和 `-extraSheet` 到 `extraArray` `extraSheetArray`\n2. 校验 `长度` 相等\n3. 遍历 -\u003e `i`\n   1. 识别 `extraArray[i]` 后缀\n      1. 后缀 `xlsx` 时，`AppendSheet` `extraSheetArray[i]` -\u003e `extraSheetArray[i]`\n      2. 其他时，作为 `txt` 读取 `slice` 写入 `sheet` `extraSheetArray[i]`\n\n### fam_info\n\n1. 写表头 \"SampleID\"\n2. 遍历 `sampleList` -\u003e `sampleID`\n   1. `AddRow().AddCell().SetString(sampleID)` `sampleID` 写入 行\n\n### FV\n\n1. `loadData` 读取到 哈希数值 `data`\n   1. 文件名后缀识别 是否 `gzip` 文件，分别读取\n2. `cycle1`\n   1. `data` 循环1 -\u003e `item`\n      1. `annotate1(item)`\n         1. inhouse_AF -\u003e frequency\n         2. 历史验证假阳次数 \u003c- \"重复数\"\n         3. score to prediction\n         4. update Function\n         5. update FuncRegion\n         6. gene symbol -\u003e geneID\n         7. 基于 `geneIDs` 注释 `CHPO` `基因-疾病` `突变频谱`\n         8. 注释 孕前数据库\n         9. 注释 新生儿数据库\n         10. 注释 耳聋数据库\n         11. \"Omim Gene\" -\u003e \"Gene\"\n         12. \"OMIM_Phenotype_ID\" -\u003e \"OMIM\"\n         13. `-acmg` 时\n         14. `anno.UpdateSnv(item, *gender)`\n         15. \"引物设计\"\n         16. `item[\"flank\"] += \" \" + item[\"HGVSc\"]`\n         17. \"变异来源\"\n         18. 注释 Tier\n         19. Tier1 时\n           1. `anno.UpdateSnvTier1(item)`\n           2. `anno.UpdateAutoRule(item)`\n           3. `anno.UpdateManualRule(item)`\n           4. `annotate1Tier1(item)`\n3. `delDupVar`\n   1. `data` 循环1 -\u003e `item`\n      1. `Tier1` 时\n         1. \"#Chr\"+\"Start\"+\"Stop\"+\"Ref\"+\"Call\"+\"Gene Symbol\"+\"Transcript\" -\u003e `key`\n         2. `countVar[key] \u003e 1` 时\n            1. 加入 `duplicateVar[key]`\n   2. 遍历 `duplicateVar` -\u003e `key,items`\n      1. 最大 `score = anno.FuncInfo[item[\"Function\"]]` -\u003e `maxFunc`\n      2. 遍历 `items` -\u003e `item`\n         1. `score \u003c maxFunc` 时\n            1. `item[\"delete\"] = \"Y\"`\n            2. `deleteVar[key+\"\\t\"+transcript] = true`\n            3. `countVar[key]--`\n            4. `continue`\n         2. 最小 `transcriptLevel[transcript]` -\u003e `minTrans`\n      3. 遍历 `items` -\u003e `item`\n         1. `item[\"delete\"] != \"Y\"` 且 `transcriptLevel[transcript] != minTrans`\n            1. `item[\"delete\"] = \"Y\"`\n            2. `deleteVar[key+\"\\t\"+transcript] = true`\n            3. `countVar[key]--`\n4. `cycle2`\n   1. `data` 循环1 -\u003e `item`\n      1. `Tier1` 时 `anno.InheritCheck(item,inheritDb)`\n         1. 统计 基因-转录本 维度 \"Zygosity\" \"ModeInheritance\" 分布计数\n   2. `data` 循环2 -\u003e `item`\n      1. \"ClinVar Significance\" 拼接 \"CLNSIGCONF\"\n      2. `Tier1` 时\n         1. \"MutationName\" 记入 `tier1Db`\n         2. \"#Chr\"+\"Start\"+\"Stop\"+\"Ref\"+\"Call\"+\"Gene Symbol\"+\"Transcript\" -\u003e `key`\n         3. 非 `deleteVar[key]`\n            1. `annotate2(item)`\n               1. `item[\"遗传相符\"] = anno.InheritCoincide(item, inheritDb, *trio)`\n               2. `item[\"遗传相符-经典trio\"] = anno.InheritCoincide(item, inheritDb, true)`\n               3. `item[\"遗传相符-非经典trio\"] = anno.InheritCoincide(item, inheritDb, false)`\n               4. `familyTag = \"single\"`\n                  1. `-trio || -trio2` 时 `familyTag = \"trio\"`\n                  2. `-couple` 时 `familyTag = \"couple\"`\n               5. `item[\"familyTag\"] = anno.FamilyTag(item, inheritDb, \"trio\")`\n               6. `item[\"筛选标签\"] = anno.UpdateTags(item, specVarDb, *trio, *trio2)`\n               7. `anno.Format(item)`\n                  1. `FloatFormat`\n                     1. 6 位精度\n                  2. `NewLineFormat`\n                     1. `\u003cbr/\u003e` -\u003e `\\n`\n               8. `-wesim` 时 `annotate2IM`\n                  1. \"Gene Symbol\" 属于 `acmgSFGene` 时 item[\"IsACMG59\"] = \"Y\"\n                  2. \"DiseaseName/ModeInheritance\" 拼接 \"DiseaseNameCH\"或\"DiseaseNameEN\"与\"ModeInheritance\"\n                  3. `-trio` 时，拆分 \"Zygosity\" 给 \"Genotype of Family Member 1\" \"Genotype of Family Member 2\"\n                  4. 基于 `wesim.resultColumn` 写入 `resultArray`\n            2. 写入 Tier1 \"filter_variants\"\n            3. 非 `-wgs` 时 写入 Tier2\n      3. `outputTier3` 时\n         1. `SteamWriterSetStringMap2Row(tier3SW, 1, tier3RowID, item, tier3Titles)`\n            1. 流式写入 `Tier3`\n5. `-wgs` 时 `wgsCycle`\n   1. 初始化 `wgsXlsx`\n      1. MTTitle 写入 \"MT\" Sheet，filterVaristsTitle 写入 \"intron\" Sheet\n   2. `data` 循环1 -\u003e `item`\n      1. 重新计算 `wgs` 模式 `Tier1`\n      2. `Tier1` 时 `anno.InheritCheck(item,inheritDb)`\n         1. 统计 基因-转录本 维度 \"Zygosity\" \"ModeInheritance\" 分布计数\n   3. `data` 循环2 -\u003e `item`\n      1. `Tier1` 时\n         1. `annotate4(item)`\n           1. `item[\"遗传相符\"] = anno.InheritCoincide(item, inheritDb, *trio)`\n         2. `item[\"遗传相符-经典trio\"] = anno.InheritCoincide(item, inheritDb, true)`\n         3. `item[\"遗传相符-非经典trio\"] = anno.InheritCoincide(item, inheritDb, false)`\n         4. `-trio`时 `item[\"familyTag\"] = anno.FamilyTag(item, inheritDb, \"trio\")`\n         5. `item[\"筛选标签\"] = anno.UpdateTags(item, specVarDb, *trio, *trio2)`\n         2. 写入 Tier2\n         3. 额外（不在 `tier1Db` ） 非 \"no-change\" 变异\n         1. 写入 `intronSheet`\n      2. 线粒体 写入 \"MT\" Sheet\n\n#### Tier\n\n- [功能集](etc/function.level.json)\n\n- trio\n  - 先证者有检出\n    - 自动化判断 不是 B/LB || 有证据项 PM2\n      - denovo\n        - 公共频率\u003c=0.01\n          - 基因集\n            - 功能集\n              - Tier1\n            - ( WGS || SpliceAI Pred D ) \u0026\u0026 非 \"no-change\"\n              - 本地频率\u003c=0.01\n                - Tier1\n              - Tier2\n            - Tier2\n          - Tier2\n        - Tier2\n      - 非denovo\n        - 公共频率\u003c=0.01\n          - 基因集\n            - 功能集\n              - Tier1\n            - ( WGS || SpliceAI Pred D ) \u0026\u0026 非 \"no-change\"\n              - 本地频率\u003c=0.01\n                - Tier1\n              - Tier3\n            - Tier3\n          - Tier3\n        - Tier3\n- single\n  - 自动化判断 不是 B/LB || 有证据项 PM2\n    - 公共频率\u003c=0.01\n      - 基因集\n        - 功能集\n          - Tier1\n        - ( WGS || SpliceAI Pred D ) \u0026\u0026 非 \"no-change\"\n          - 本地频率\u003c=0.01\n            - Tier1\n          - Tier3\n        - Tier3\n      - Tier3\n    - Tier3\n- HGMD DM || ClinVar P/LP\n  - 非线粒体\n    - 公共频率 \u003c= 0.05\n      - Tier1\n    - 非 Tier1\n      - Tier2\n- 特殊位点库\n  - Tier1\n\n#### 遗传相符\n\n#### familyTag\n\n#### 烈性突变:\n\n- \"splice-3\"\n- \"splice-5\"\n- \"init-loss\"\n- \"start_lost\"\n- \"alt-start\"\n- \"frameshift\"\n- \"nonsense\"\n- \"stop-gain\"\n- \"stop_gained\"\n- \"span\"\n\n#### 筛选标签\n\n- 标签拼接，分号 `;` 分割\n- Tag1AFThreshold = 0.05\n\n- tag1:\n  - 本地频率\u003c=Tag1AFThreshold || 特殊变异库 || HGMD DM || ClinVar P/LP\n    - trio || trio2\n        - 遗传相符-经典trio == 相符\n            - 标签 T1\n        - 遗传相符-非经典trio == 相符\n            - 遗传模式: AR || XL || YL\n                - 标签 1\n            - 遗传模式: AD\n                - 纯合记录 均为 0\n                    - 标签 1\n    - single\n        - 遗传相符 == 相符\n        - 遗传模式: AR || XL || YL\n          - 标签 1\n        - 遗传模式: AD\n          - 纯合记录 均为 0\n            - 标签 1\n\n- tag2:\n  - 特殊变异库 || HGMD DM || ClinVar P/LP\n    - 标签 2\n\n- tag3:\n  - 本地频率\u003c=0.01\n    - 烈性突变\n      - 标签 3\n\n- tag4:\n  - 本地频率\u003c=0.01\n    - PP3\n      - 标签 4\n      - PP3:\n        - 非 PVS1\n          - 保守性: {\"GERP++_RS_pred\",\"PhyloP Vertebrates Pred\",\"PhyloP Placental Mammals Pred\"}\n            - 无 不保守\n              - 至少2个 保守\n                - 有害性b1: {\"Ens Condel Pred\",\"SIFT Pred\",\"MutationTaster Pred\",\"Polyphen2 HVAR Pred\"}\n                  - 无 良性/多态\n                    - 至少2个 有害\n                      - PP3\n                - 有害性b2: {\"dbscSNV_RF_pred\",\"dbscSNV_ADA_pred\",\"SpliceAI Pred\"}\n                  - 无 良性/多态\n                    - 至少2个 有害\n                      - PP3\n                    - spliceAI 有害\n                      - PP3\n    - 非重复区域 \u0026\u0026 特定功能\n      - 标签 4\n      - 特定功能:\n        - stop-loss\n        - cds-ins\n        - cds-del\n        - cds-indel\n    - SpliceAI Pred == D \u0026\u0026 功能非 no-change\n      - 标签 4\n\n### LOH\n\n1. `appendLOHs(excel, lohs, lohSheetName, sampleList)`\n1. 遍历 `lohs` -\u003e `i,path`\n  1. 对应 `sampleID = sampleList[i]`\n  2. `AppendSheet` `lohSheetName` -\u003e `sampleID+\"-loh\"`\n\n### QC\n\n1. `parseQC`\n   1. `-karyotype` -\u003e `karyotypeMap`\n   2. `-qc`\n      1. `loadQC(*qc, *kinship, qualitys, *wgs)`\n         1. `kinship` -\u003e `kinshipHash`\n         2. `sep=\"\\t\"`\n            1. `-wgs` 时 `sep=\": \"`\n         3. 遍历 `-qc` -\u003e `i,path`\n            1. 遍历 `paht` -\u003e `line`\n               1. `^## Files : (\\S+)` -\u003e \"bamPath\"\n               2. `sep` 分割， `TrimSpace`，填充 `quality[i]`\n            2. `-wgs` 时 修正 \"bamPath\"\n            3. `kinshipHash[quality[i][\"样本编号\"]]` 更新 `quality[i]`\n      2. 遍历 `qualitys` -\u003e `quality`\n         1. `qualityKeyMap` 更新列\n         2. `karyotypeMap[quality[\"样本编号\"]]` 更新 \"核型预测\"\n         3. \"原始数据产出（Gb）\" 更新 \"原始数据产出（Mb）\"\n         4. `-wesim` 时\n            1. 基于 `wesim.qcColumn` 写入 `qcFile`\n      3. `-filterStat` -\u003e `loadFilterStat(*filterStat, qualitys[0])`\n         1. 计算 \"Q20 碱基的比例\" \"Q30 碱基的比例\" \"测序数据的 GC 含量\" \"低质量 reads 比例\"\n      4. `-imqc`\n         1. 遍历 `-imqc` -\u003e `path` 读入 `imqc`\n         2. 遍历 `qualitys` -\u003e `sampleID,quality`\n            1. `imqc[sampleID]` 更新 `quality`\n               1. `quality[\"Q20 碱基的比例\"] = qcMap[\"Q20_clean\"] + \"%\"`\n               2. `quality[\"Q30 碱基的比例\"] = qcMap[\"Q30_clean\"] + \"%\"`\n               3. `quality[\"测序数据的 GC 含量\"] = qcMap[\"GC_clean\"] + \"%\"`\n               4. `quality[\"低质量 reads 比例\"] = qcMap[\"lowQual\"] + \"%\"`\n      5. `mtqc`\n         1. 遍历 `-mtqc` -\u003e `i,path`\n            1. `path` 读入 `qcMap`\n            2. `qcMap` 更新 `qualitys[i]`\n2. `updateQC`\n   1. \"罕见变异占比（Tier1/总）\"\n   2. \"罕见烈性变异占比 in tier1\"\n   3. \"罕见纯合变异占比 in tier1\"\n   4. \"纯合变异占比 in all\"\n   5. 循环 -\u003e chr\n      1. \"chr\"+chr+\"纯合变异占比\"\n   6. \"SNVs_all\"\n   7. \"SNVs_tier1\"\n   8. \"Small insertion（包含 dup）_all\"\n   9. \"Small insertion（包含 dup）_tier1\"\n   10. \"Small deletion_all\"\n   11. \"Small deletion_tier1\"\n   12. \"exon CNV_all\"\n   13. \"exon CNV_tier1\"\n   14. \"large CNV_all\"\n   15. \"large CNV_tier1\"\n3. `addQCSheet`\n   1. `AddSheet(\"quality\")`\n   2. 遍历 `qualityColumn` -\u003e `key` 横向赋值\n      1. `row=sheet.AddRow`\n      2. `row.AddCell().SetString(key)`\n      3. 遍历 `qualitys` -\u003e `quality`\n         1. `row.AddCell().SetString(item[key])`\n\n## AES加密数据库\n\n### DATABASE\n\n| name         | version               | note |\n|--------------|-----------------------|------|\n| 全外疾病库        | 2023.Q1-2023.05.17    |      |\n| 基因突变谱        | V6.0.0.20230411       |      |\n| ACMGSF       | V2.0.2023.5           |      |\n| PrePregnancy | PP155-V5.1.2_20230427 |      |\n| NBSP         | V2.3.1.20230505       |      |\n| VIPHL        | 20230509.30952        |      |\n\n### 疾病库/基因频谱\n\n```shell\n#!/bin/bash\nwget -N https://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/non_alt_loci_set.txt\nstat non_alt_loci_set.txt\nbuildDb/buildDb \\\n  -prefix db/全外疾病库 \\\n  -key 'entry ID' \\\n  -input db/backup/全外疾病库2023.Q1-2023.05.17.xlsx \\\n  -sheet '更新后全外背景库（6272疾病OMIMID，4787个基因）' \\\n  -rowCount 8425 -keyCount 4787\n\n```\n\nor\n\n```shell\nsh buildDb/buildDb.sh 'db/backup/全外疾病库2023.Q1-2023.05.17.xlsx' '更新后全外背景库（6272疾病OMIMID，4787个基因）' 'entry ID' 8425 4787 db/全外疾病库\n```\n\n![buildDb.png](docs/buildDb.png)\n\n### ACMG Secondary Finding\n\n```shell\nsfCode=b7ea138a9842cbb832271bdcf4478310 # 替换成实际密钥\n../NB2xlsx/buildDB/buildDB -code $sfCode -extract \"#Chr,Start,Stop,Ref,Call,Gene Symbol,Transcript,cHGVS,证据项,致病等级,参考文献,关联疾病表型OMIM号,关联疾病英文名称,关联疾病中文名称,数据库时间\" -input db/backup/ACMG\\ 73基因23.05.23.xlsx -output db/ACMGSF.json.aes -keys etc/SF.key.txt\n```\n\n![img.png](docs/ACMGSF.png)\n\n#### 注释校验\n\n- 因本地无 `WES` 注释流程，上传 `db/ACMGSF.json.aes` 和 `db/ACMGSF.json.aes.mut.tsv` 到集群\n- 如本地有WES注释流程，所有操作均可在同一环境下完成\n- 确认 `致病性等级` 的 注释个数 一致\n\n```shell\nsh check/check.sf.sh\n```\n\n![img.png](docs/ACMGSF.check.1.png)\n![img.png](docs/ACMGSF.check.2.png)\n\n### VIPHL\n\n- 因原始 `VIPHL` 库文件格式不同，需要转换格式步骤\n- 因建库和注释校验均需 `WES` 注释流程，整个下面示例步骤统一在集群环境处理，注意非加密文件的保密\n- 如本地有 `WES` 注释流程，建议在本地进行处理\n- 核对 `差异`\n\n```shell\n## 生成 check/VIPHL/VIPHL.xlsx\nchmod 700 check/viphl位点唯一结果统计_total_20230509.xlsx\nsh check/build.hl.sh viphl位点唯一结果统计_total_20230509.xlsx\n\n## 构建加密库\nhlCode=6d276bc509883dbafe05be835ad243d7  # 替换成实际密钥\nsrc/NB2xlsx/buildDB/buildDB -code $hlCode -extract \"#Chr,Start,Stop,Ref,Call,Gene Symbol,Transcript,cHGVS,HLinterpretation,HLcriteria\" -input check/VIPHL/VIPHL.xlsx -keys etc/HL.key.txt -output db/VIPHL.json.aes -sheet Result\n\n## 注释校验\nsh check/check.hl.sh\n```\n\n![VIPHL format](docs/VIPHL.format.png)\n![VIPHL build](docs/VIPHL.png)\n![VIPHL.check.2](docs/VIPHL.check.2.png)\n\n## CHPO\n\n```shell\nbuildHPO -chpo chpo-2021.json -g2p genes_to_phenotype.txt -output db/gene2chpo.txt\n```\n\n## 注意\n\n### 特殊位点库\n\n**特殊位点库**根据华大内部流程`bgicg_anno.pl`注释结果中的`MutationName`查找是否特殊位点，\n所以以下情形可能发生库失效问题：\n\n1. 输入文件不包含以`MutationName`命名的特定注释结果\n2. 输入文件`MutationName`与流程`bgicg_anno.pl`结果不一致\n3. 流程`bgicg_anno.pl`有更新，但是数据库未同步更新\n4. 位点用与现有配置不同的数据库注释导致的注释结果不一致\n\n## 特性\n\n### 结合性处理\n\n#### 格式转换\n\n```go\nfunc zygosityFormat(zygosity string) string {\nzygosity = strings.Replace(zygosity, \"het-ref\", \"Het\", -1)\nzygosity = strings.Replace(zygosity, \"het-alt\", \"Het\", -1)\nzygosity = strings.Replace(zygosity, \"hom-alt\", \"Hom\", -1)\nzygosity = strings.Replace(zygosity, \"hem-alt\", \"Hemi\", -1)\nzygosity = strings.Replace(zygosity, \"hemi-alt\", \"Hemi\", -1)\nreturn zygosity\n}\n```\n\n#### Het-\u003eHom修正\n\n```go\nfunc homRatio(item map[string]string, threshold float64) {\nvar aRatio = strings.Split(item[\"A.Ratio\"], \";\")\nvar zygositys = strings.Split(item[\"Zygosity\"], \";\")\nif len(aRatio) \u003c= len(zygositys) {\nfor i := range aRatio {\nvar zygosity = zygositys[i]\nif zygosity == \"Het\" {\nvar ratio, err = strconv.ParseFloat(aRatio[i], 64)\nif err != nil {\nratio = 0\n}\nif ratio \u003e= threshold {\nzygositys[i] = \"Hom\"\n}\n}\n}\n}\nitem[\"Zygosity\"] = strings.Join(zygositys, \";\")\n}\n```\n\n#### Hom-\u003eHemi修正\n\n```go\nfunc hemiPAR(item map[string]string, gender string) {\nvar chromosome = item[\"#Chr\"]\nif isChrXY.MatchString(chromosome) \u0026\u0026 isMale.MatchString(gender) {\nstart, e := strconv.Atoi(item[\"Start\"])\nsimpleUtil.CheckErr(e, \"Start\")\nstop, e := strconv.Atoi(item[\"Stop\"])\nsimpleUtil.CheckErr(e, \"Stop\")\nif !inPAR(chromosome, start, stop) \u0026\u0026 withHom.MatchString(item[\"Zygosity\"]) {\nzygosity := strings.Split(item[\"Zygosity\"], \";\")\ngenders := strings.Split(gender, \",\")\nif len(genders) \u003c= len(zygosity) {\nfor i := range genders {\nif isMale.MatchString(genders[i]) \u0026\u0026 isHom.MatchString(zygosity[i]) {\nzygosity[i] = strings.Replace(zygosity[i], \"Hom\", \"Hemi\", 1)\n}\n}\nitem[\"Zygosity\"] = strings.Join(zygosity, \";\")\n} else {\nlog.Fatalf(\"conflict gender[%s]and Zygosity[%s]\\n\", gender, item[\"Zygosity\"])\n}\n}\n}\n}\n```\n\n## 复用拼接字段\n\n因下游数据库结构新增字段开发工作量大，部分额外字段拼接进已有字段内\n\n| 复用字段                   | 拼接字段         | 连接字符串 | 是否可选拼接 |\n|------------------------|--------------|-------|--------|\n| `ClinVar Significance` | `CLNSIGCONF` | ':'   | 是      |\n| `flank`                | `HGVSc`      | ' '   | 是      |\n\n## 注意\n\n- exon cnv输入文件不存在时仅log报错，不中断\n\n## UTIL\n\n### `tier1tags`\n\nWGS 使用anno2xlsx过滤后，进行spliceAI注释和过滤，然后重新进行Tier1判断、\"遗传相符\"和\"筛选标签\"\n参考 [README.md](util/tier1tags/README.md)\n\n#### TO-DO\n\n- [ ] Tier1判断去冗余\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fliserjrqlxue%2Fanno2xlsx","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fliserjrqlxue%2Fanno2xlsx","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fliserjrqlxue%2Fanno2xlsx/lists"}