{"id":51879679,"url":"https://github.com/19pine-ai/namerank","last_synced_at":"2026-07-25T11:01:27.436Z","repository":{"id":371481631,"uuid":"1235612598","full_name":"19PINE-AI/namerank","owner":"19PINE-AI","description":"NameRank: measuring LLM-mediated recognition as the post-bibliometric impact channel","archived":false,"fork":false,"pushed_at":"2026-07-15T06:41:50.000Z","size":40019,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-15T08:23:07.281Z","etag":null,"topics":["ai-research","benchmark","bibliometrics","large-language-models","llm","llm-evaluation","nlp","recognition"],"latest_commit_sha":null,"homepage":"https://01.me/research/namerank","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/19PINE-AI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-05-11T13:43:48.000Z","updated_at":"2026-07-15T06:41:58.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/19PINE-AI/namerank","commit_stats":null,"previous_names":["19pine-ai/namerank"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/19PINE-AI/namerank","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/19PINE-AI%2Fnamerank","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/19PINE-AI%2Fnamerank/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/19PINE-AI%2Fnamerank/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/19PINE-AI%2Fnamerank/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/19PINE-AI","download_url":"https://codeload.github.com/19PINE-AI/namerank/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/19PINE-AI%2Fnamerank/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35877013,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-25T02:00:06.922Z","response_time":64,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-research","benchmark","bibliometrics","large-language-models","llm","llm-evaluation","nlp","recognition"],"created_at":"2026-07-25T11:01:26.663Z","updated_at":"2026-07-25T11:01:27.427Z","avatar_url":"https://github.com/19PINE-AI.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# NameRank\n\n**The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank**\n*by Bojie Li (Pine AI) and Noah Shi (University of Washington).*\n\nNameRank is a continuous $[0,1]$ recognition score for people and named artifacts\nin the LLM era. Each entity is probed with one open-ended question across a\n36-model frontier panel; an independent judge returns a **binary recognition\nverdict** against a curated gold answer — *did the model state a specific,\nnon-guessable fact about this exact entity?* — so hallucination, context echo,\nand lucky guesses earn nothing. NameRank is the fraction of the panel that\nrecognizes the entity. It operationalizes the recognition-variance residual that\nbibliometrics cannot explain (Li 2026, IKP §5.7).\n\n- **Paper:** [arXiv:2607.12520](https://arxiv.org/abs/2607.12520) · [HTML](https://arxiv.org/html/2607.12520) · [PDF](https://arxiv.org/pdf/2607.12520) (47 pages) · local copy [`paper/main.pdf`](paper/main.pdf) · source: [`paper/main.tex`](paper/main.tex)\n- **Companion site:** [`site/`](site/) — an interactive React explainer (mechanism walkthrough, findings dashboards, and a 4,685-entity explorer with every model's verbatim answer and verdict). Live at \u003chttps://01.me/research/namerank\u003e.\n- **Scale:** 4,685 entities across 54 cohorts, a 36-model panel, and the record-level recognition verdicts in [`experiments/t6_v2_protocol/outputs/recognition_final.jsonl.gz`](experiments/t6_v2_protocol/outputs/recognition_final.jsonl.gz) (gzipped, ~12 MB).\n- **Robustness:** [`experiments/`](experiments/) — 18 self-contained follow-up audits, all folded into the paper. See [§ Robustness experiments](#robustness-experiments-experiments).\n\n## Headline findings\n\n1. **The credential treadmill — and its marquee inversion.** Every Olympic-style\n   credential sits *below* a working-researcher baseline (0.40): IMO gold 0.12,\n   Rhodes 0.15, MSRA PhD Fellowship 0.18 — because no named artifact ships with\n   the medal. Yet the ranking flips at the very top, where Nobel (0.98), Turing\n   (0.97), and Fields (0.96) laureates saturate the panel.\n2. **Artifact \u003e creator.** For independent creators the tool out-ranks its maker\n   (e.g. Tianshou 0.78 vs. its author 0.22); the credential that *does* propagate\n   is a named method or an awarded paper. Being one of many named authors on a\n   flagship model report or system card earns almost nothing — recognition\n   attaches to the artifact's distinctive name, not the roster behind it.\n3. **No bibliometric predicts recognition well.** Under the recognition verdict,\n   $\\log(h\\text{-index})$ and $\\log(\\text{citations})$ each explain only\n   $R^2\\approx0.22$; neither is a strong instrument for corpus presence.\n4. **Corpus-density gradient.** Top-density institutions out-recognize peers at\n   matched citations; US CS faculty average 0.64 vs. 0.33 for China.\n5. **Attention ≠ recognition, and models can't introspect it.** On 258 news\n   events, recognition loads on peak salience, not persistence; a self-report\n   probe shows a model's \"what do I know?\" reads a corpus prior, not its own\n   knowledge. Two Chinese-language sites with enormous user bases make the\n   boundary concrete: tuixue.online (peak audience ~1M) scores 0.39 and\n   icourse.club 0.14, both below nanoGPT (0.78), whose audience is far smaller —\n   because those sites spread *functionally*, with no name attached in indexable\n   text.\n\n## Repository layout\n\n```\nnamerank/\n├── paper/                         # LaTeX sources + compiled PDF\n│   ├── main.tex  appendix.tex  references.bib\n│   ├── arxiv.sty                  # single-column tech-report style\n│   ├── main.pdf                   # compiled paper (47pp)\n│   └── figures/\n│       ├── _data.py               # loads the recognition run for every figure\n│       ├── compute_all_numbers.py # → computed_numbers.json (every number the paper cites)\n│       ├── make_fig*.py           # regenerate each figure PDF\n│       └── fig_*.pdf\n├── site/                          # React/Vite companion site (source of the live explainer)\n│   ├── src/{sections,components,lib,data}/   scripts/build_data.py\n│   └── public/data/               # generated chart + entity data\n├── docs/                          # reproducibility docs (probe/judge prompts, cohorts, worked examples)\n├── data/\n│   ├── inputs/                    # entities, gold answers, model set, probe/judge templates\n│   └── analysis/                  # derived recognition tables (see data/analysis/README.md)\n├── code/\n│   ├── run_probe.py               # shared probe → judge → embedding harness\n│   ├── build_release_tables.py    # regenerate data/analysis/ from the recognition run\n│   ├── country_affiliation.py     # institution → country map (country gradient)\n│   └── _paths.py                  # repo-relative path helpers\n├── experiments/                   # 18 robustness audits + t6_v2_protocol (the recognition run)\n├── tool/                          # self-contained NameRank measurement CLI\n├── requirements.txt\n└── LICENSE\n```\n\n## Reproducing the paper\n\nThe record-level source of truth is\n`experiments/t6_v2_protocol/outputs/recognition_final.jsonl.gz` — one binary\n`recognized` verdict per (entity, model), shipped gzipped (~12 MB; 234,574\nrecords). Everything downstream is a deterministic function of that file.\n`build_release_tables.py` reads the `.gz` directly, so no manual\ndecompression is needed.\n\n```bash\npip install -r requirements.txt\n\n# 1. Rebuild the released aggregate tables from the recognition run.\n#    Cross-checks every value against the paper and refuses to write on drift.\npython3 code/build_release_tables.py\n\n# 2. Recompute every number the paper cites, and regenerate the figures.\ncd paper/figures\npython3 compute_all_numbers.py          # → computed_numbers.json\nfor f in make_fig*.py; do python3 \"$f\"; done\n\n# 3. Compile the paper.\ncd .. \u0026\u0026 pdflatex main.tex \u0026\u0026 bibtex main \u0026\u0026 pdflatex main.tex \u0026\u0026 pdflatex main.tex\n```\n\nThe paper uses the bundled `arxiv.sty`; required LaTeX packages are all in\n`texlive-latex-extra` / `texlive-fonts-extra`. Figures and numbers have no\nhard-coded values — they load the recognition records through\n`paper/figures/_data.py`.\n\n### Re-running the probe pipeline (~$2.5K end-to-end)\n\n```bash\nexport OPENROUTER_API_KEY=...   # probed-model calls\nexport GEMINI_API_KEY=...       # judge\n\n# The recognition run's orchestration lives in experiments/t6_v2_protocol/scripts/\n# (probe → open-book recognition judge → uniform final pass). run_probe.py is the\n# shared probe/judge harness the experiments import; it is resumable.\n```\n\n### Companion site\n\n```bash\ncd site\nnpm install\nnpm run data     # regenerate JSON assets from ../data and ../experiments\nnpm run dev      # local preview\nnpm run build    # → dist/ (base './', works at any mount path)\n```\n\n`site/scripts/build_data.py` reads the recognition sources\n(`recognition_final.jsonl`, `computed_numbers.json`, the appendix tables). The\nper-entity answer shards (`site/public/data/answers/`, ~157 MB) are regenerated\nat build time and are gitignored; `answers_index.json` is committed.\n\n## Data schemas\n\n- **`experiments/t6_v2_protocol/outputs/recognition_final.jsonl.gz`** — one JSON object per (entity, model): `{dataset, entity_id, model_id, recognized (0/1), rationale}` (gzipped; `zcat` to read).\n- **`data/analysis/namerank_per_entity.csv`** — one row per entity: `entity_id, entity_name, cohort, n_models, namerank, namerank_sd, refusal_rate, embedding_sim_mean`.\n- **`data/analysis/namerank_matrix.json`** — `{entity_id: {model_id: recognized}}` for the 4,730-entity × 36-model main run.\n- **`data/analysis/cohort_summary.csv` / `credential_ladder.csv` / `per_model_summary.csv` / `cs_faculty_by_country.csv` / `cross_language_per_entity.csv` / `attribution_pairs_v2.csv`** — derived recognition tables; see [`data/analysis/README.md`](data/analysis/README.md).\n- **`data/inputs/`** — `pilot_entities.json` (entities + disambiguating context), `gold_answers.json`, `model_set.json`, `probe_template_{en,zh}.txt`, `judge_prompt.txt`. The recognition run's exact v2 inputs (gold answers, contexts, the open-book judge) live under `experiments/t6_v2_protocol/inputs/`.\n\n## Robustness experiments (`experiments/`)\n\nEach subdirectory is self-contained (`analyze.py` / scripts + derived outputs +\n`README.md` with the headline and reproduction command). The large raw\nper-(entity, model) probe dumps are gitignored; the derived summaries the\nanalyses consume are committed.\n\n| Experiment | Question | Headline |\n|---|---|---|\n| `t1_1_gold_length` | Does gold-answer length drive the credential gap? | No; treadmill ordering survives length adjustment. |\n| `t1_2_context_ab` | Does the disambiguating context leak the answer? | No; the corpus-density gradient survives context ablation. |\n| `t1_3_synthetic_null` | What does a never-seen name score? | Floor near zero; verdicts track the entity, not the model. |\n| `t1_4_wikipedia` | Is NameRank just \"has a Wikipedia page\"? | No; Wikipedia explains little of the variance. |\n| `t2_6_prompt_sensitivity` | Robust to probe wording? | Ordering robust (Pearson 0.93–0.98). |\n| `t2_7_artifact_mediation` | Is named-artifact amplification causal? | Injecting the artifact into context lifts recognition, via retrieval-deficient models. |\n| `t2_8_gendered_names` | Is there a gender bias? | Small reproducible man-coded lift, surviving controls. |\n| `t2_9_fractional_citations` | Is the bibliometric signal attribution density? | No — it is name-recurrence across distinct works. |\n| `t2_10_cross_judge` | A single-judge artifact? | No; findings hold across Gemini/GPT-5/Claude judges. |\n| `t3_1_cutoff_gradient` | Is the silent zone a corpus-timing artifact? | No — intrinsic; matched DiD ≈ 0. |\n| `t4_1_news_events` | Attention vs. recognition on 258 events. | Recognition loads on peak salience, not persistence. |\n| `t5_1_award_ladder` | Where does the credential ladder invert? | At the marquee tier (Nobel/Turing/Fields). |\n| `t5_2_attention_baseline` | Is NameRank just Wikipedia-pageview rank? | No — undefined for most entities; flow vs. stock. |\n| `t5_3_noi_medal_tiers` | Gold vs. silver vs. bronze. | A gold-vs-non-gold cliff; silver ≈ bronze. |\n| `t5_4_self_report` | Can you just ask the model what it knows? | No — self-report reads a corpus prior, not its own state. |\n| `t5_4_university_baseline` | University-faculty baseline. | Institution-density gradient at matched citations. |\n| `t5_5_llm_area` | Method originators vs. their methods. | The named method out-ranks its originator. |\n| `t5_5_noi_boundary_rd` | Recognition at a naming boundary. | Unique, English-documented names propagate; utility names do not. |\n| `t6_v2_protocol` | The recognition run itself. | Probe → open-book recognition judge → uniform final pass. |\n\n## Citation\n\n```bibtex\n@article{li2026namerank,\n  title={The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank},\n  author={Li, Bojie and Shi, Noah},\n  journal={arXiv preprint arXiv:2607.12520},\n  year={2026},\n  eprint={2607.12520},\n  archivePrefix={arXiv},\n  primaryClass={cs.AI},\n  note={Code: https://github.com/19PINE-AI/namerank}\n}\n```\n\n## License\n\nDual-licensed — see [`LICENSE`](LICENSE). Code: MIT. Data (probe specifications,\ngold answers, recognition verdicts, analysis tables): CC BY 4.0.\n\n## Status\n\nAccompanying the arXiv preprint [arXiv:2607.12520](https://arxiv.org/abs/2607.12520).\nThe metric is intended for re-runs at each major\nfrontier-model release cycle; this repo is tagged at preprint freeze and re-tagged\nfor each subsequent NameRank revision.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2F19pine-ai%2Fnamerank","html_url":"https://awesome.ecosyste.ms/projects/github.com%2F19pine-ai%2Fnamerank","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2F19pine-ai%2Fnamerank/lists"}