{"id":51999463,"url":"https://github.com/integration-automation/thesisagent","last_synced_at":"2026-07-31T10:01:08.688Z","repository":{"id":358273612,"uuid":"1240690025","full_name":"Integration-Automation/ThesisAgent","owner":"Integration-Automation","description":"Keyword-driven academic paper search across 11 sources (arXiv, Semantic Scholar, OpenAlex, PubMed, IEEE, ACM, DBLP, Crossref, OpenAIRE, Springer, Scholar) → thesis-style PowerPoint + Excel + BibTeX in one CLI call. Includes an MCP server. 14-language i18n.","archived":false,"fork":false,"pushed_at":"2026-06-14T18:51:39.000Z","size":1790,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-06-14T19:16:14.777Z","etag":null,"topics":["academic-papers","arxiv","bibtex","cli","crossref","i18n","ieee-xplore","literature-review","mcp","mcp-server","openalex","paper-search","powerpoint","pptx","pubmed","python","research-tools","semantic-scholar","springer-nature","thesis"],"latest_commit_sha":null,"homepage":"https://autopapertoppt.readthedocs.io/en/latest/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Integration-Automation.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"docs/contributing.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-05-16T12:53:33.000Z","updated_at":"2026-06-14T17:38:35.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/Integration-Automation/ThesisAgent","commit_stats":null,"previous_names":["integration-automation/autopapertoppt","integration-automation/thesisagent"],"tags_count":7,"template":false,"template_full_name":null,"purl":"pkg:github/Integration-Automation/ThesisAgent","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Integration-Automation%2FThesisAgent","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Integration-Automation%2FThesisAgent/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Integration-Automation%2FThesisAgent/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Integration-Automation%2FThesisAgent/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Integration-Automation","download_url":"https://codeload.github.com/Integration-Automation/ThesisAgent/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Integration-Automation%2FThesisAgent/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36114483,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-31T02:00:06.731Z","response_time":112,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["academic-papers","arxiv","bibtex","cli","crossref","i18n","ieee-xplore","literature-review","mcp","mcp-server","openalex","paper-search","powerpoint","pptx","pubmed","python","research-tools","semantic-scholar","springer-nature","thesis"],"created_at":"2026-07-31T10:01:07.515Z","updated_at":"2026-07-31T10:01:08.652Z","avatar_url":"https://github.com/Integration-Automation.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# ThesisAgents\n\n[![CI](https://github.com/Integration-Automation/ThesisAgents/actions/workflows/ci.yml/badge.svg)](https://github.com/Integration-Automation/ThesisAgents/actions/workflows/ci.yml)\n[![Release](https://github.com/Integration-Automation/ThesisAgents/actions/workflows/release.yml/badge.svg?branch=main)](https://github.com/Integration-Automation/ThesisAgents/actions/workflows/release.yml)\n[![PyPI](https://img.shields.io/pypi/v/thesisagents.svg)](https://pypi.org/project/thesisagents/)\n[![Python](https://img.shields.io/pypi/pyversions/thesisagents.svg)](https://pypi.org/project/thesisagents/)\n[![License: MIT](https://img.shields.io/github/license/Integration-Automation/ThesisAgents.svg)](https://github.com/Integration-Automation/ThesisAgents/blob/main/LICENSE)\n[![Docs](https://readthedocs.org/projects/thesisagents/badge/?version=latest)](https://thesisagents.readthedocs.io/en/latest/)\n\n\u003e **Languages**: **English** · [繁體中文](readmes/README.zh-TW.md) · [简体中文](readmes/README.zh-CN.md) · [日本語](readmes/README.ja.md) · [Español](readmes/README.es.md) · [Français](readmes/README.fr.md) · [Deutsch](readmes/README.de.md) · [한국어](readmes/README.ko.md) · [Português](readmes/README.pt.md) · [Русский](readmes/README.ru.md) · [Italiano](readmes/README.it.md) · [Tiếng Việt](readmes/README.vi.md) · [हिन्दी](readmes/README.hi.md) · [Bahasa Indonesia](readmes/README.id.md)\n\u003e **Documentation**: [thesisagents.readthedocs.io](https://thesisagents.readthedocs.io/en/latest/)\n\nA keyword-driven paper search assistant that fetches results from arXiv,\nSemantic Scholar, OpenAlex, PubMed, ACM (via Crossref), IEEE Xplore,\nDBLP, generic Crossref, OpenAIRE, Springer Nature, Europe PMC, DOAJ,\nHAL, CORE, and Google Scholar,\nnormalises them into one record shape, and exports the deduplicated set\nas a **thesis-style PowerPoint deck**, an **Excel workbook**, and a\n**BibTeX file** — all from one CLI call or one MCP tool call. Optionally\nenriches each paper by reading its PDF and producing a structured\nper-paper summary either in-context (LLM-as-agent flow) or via the\nAnthropic API (Python pipeline flow).\n\n## For AI agents driving this project\n\n**If you are an LLM (Claude Code, Cursor, Aider, Codex CLI, …) reading\nthis README to figure out what to do — start here.** Everything below\nis also pinned in [`AGENTS.md`](AGENTS.md) (cross-agent mirror) and\n[`CLAUDE.md`](CLAUDE.md) (canonical reference); read those before\nmaking non-trivial changes.\n\n### What the user actually wants\n\nThe default deliverable is **a thesis-style rich `.pptx` per paper**,\nnot the lightweight abstract-only deck the CLI produces by default. The\nlightweight emit is an intermediate artefact — your job is to upgrade\nit.\n\n### Decision tree\n\n1. Is `ANTHROPIC_API_KEY` set in the environment? → just run the CLI;\n   it auto-enriches via the Anthropic API. You are done.\n2. No key, but you (an LLM) are driving the session → **you produce\n   the rich summary yourself**. Read each PDF, hand-author a\n   `PaperSummary` with rich-tier fields (`pain_points`,\n   `research_question`, `contributions_detailed`, `headline_metrics`,\n   `technique_table`, `method_sections`, `evaluation_sections`,\n   `system_flow`, `research_questions`, `rq_results`,\n   `core_observation`, `limitations`, `future_work`), drop a\n   `scripts/regen_\u003cquery\u003e.py`, run it. **Do not tell the user to set\n   the API key** — you are the LLM that would have written the\n   summary.\n3. No LLM in the loop (CI / cron / unattended) → lightweight is\n   acceptable.\n\n### 6-step MCP workflow\n\n```\n1. (optional) list_sources()                              # see which plugins are enabled\n2. search(keywords, sources, top_tier_only=true, ...)\n3. (optional) download_pdfs(papers, out_dir=\"./exports/...\")\n4. fetch_pdf_text(pdf_url=paper.pdf_url)                  # per paper\n5. (you read each PDF and produce a structured summary dict)\n6. export(papers=[{...paper, \"summary\": {...}}], language=\"zh-tw\", ...)\n```\n\nAll thirteen MCP tools (including `list_sources`, `list_exports`,\n`download_pdfs`, `pptx_inspect` / `pptx_review` / `pptx_update_slide` /\n`pptx_add_slide` / etc.) are\ndocumented in [`docs/mcp.md`](docs/mcp.md).\n\n### Mandatory: URL / DOI verification before shipping\n\nPublisher URL paths **cannot be guessed** — AAAI uses numeric IDs\n(`v40i5.37389`), IEEE uses an opaque `arnumber`, ACM uses opaque DOIs.\nWhen you hand-author a `Paper`, **copy `url` / `doi` / `arxiv_id`\nverbatim from the search xlsx that produced this run** — never from\nmemory, never constructed from the title.\n\nThe xlsx is written to `exports/\u003crun\u003e/\u003cslug\u003e-\u003ctimestamp\u003e.xlsx` with\ncolumn 7 = DOI, column 8 = URL. Audit your regen script when you\nfinish:\n\n```python\nfrom openpyxl import load_workbook\nfrom scripts.regen_\u003crun\u003e import ALL_PAPERS\nreal = {sh.cell(row=r, column=2).value: sh.cell(row=r, column=8).value\n        for sh in [load_workbook(\"exports/\u003crun\u003e/\u003cslug\u003e-\u003cts\u003e.xlsx\")[\"Papers\"]]\n        for r in range(2, sh.max_row + 1)}\nfor p in ALL_PAPERS:\n    actual = next((u for t, u in real.items() if p.title[:30] in (t or \"\")), None)\n    if actual and not (p.url == actual\n                       or p.url.split(\"v\")[0] == actual.split(\"v\")[0]):\n        print(f\"! {p.bibtex_key()} authored {p.url} vs real {actual}\")\n```\n\nTwo fabrications caught this way in production: wrong AAAI volume\n(`v39i23.34521` vs real `v39i22.34537`) and invented author-slug path\n(`view/fang2026` instead of `v40i5.37389`).\n\n### Mandatory: prune irrelevant downloads before shipping\n\nSearch keyword matching is keyword-based, so off-topic papers will\nslip in: a \"Claude code\" query returned a Viterbi-decoder paper\nbecause both contain \"code\"; \"LLM code review\" matched an\nobject-detection literature review. Once you read the abstracts and\nclassify a paper off-topic for the user's actual intent, prune the\nrun dir:\n\n```python\nfrom pathlib import Path\nrun = Path(\"exports/\u003crun\u003e\")\nirrelevant_keys = (\"key-of-off-topic-paper-1\", \"key-of-off-topic-paper-2\")\nfor key in irrelevant_keys:\n    for path in (run / \"pdfs\" / f\"{key}.pdf\", run / f\"{key}.pptx\"):\n        if path.exists():\n            path.unlink()\n```\n\nDelete `exports/\u003crun\u003e/pdfs/\u003ckey\u003e.pdf` + `exports/\u003crun\u003e/\u003ckey\u003e.pptx`.\n**Keep** the aggregate `\u003cslug\u003e-\u003ctimestamp\u003e.xlsx` / `.bib` — those are\nthe honest record of what the search returned. Borderline cases get\na rich summary; better to over-include than to silently drop a\npossible match.\n\n### Worked example\n\n[`scripts/regen_llm_security_batch.py`](scripts/regen_llm_security_batch.py)\nships 8 hand-authored rich summaries built exactly this way. Use it as\nthe template for any multi-paper search. The zh-tw companion is at\n[`scripts/regen_llm_security_batch_zh_tw.py`](scripts/regen_llm_security_batch_zh_tw.py).\n\n### Don'ts\n\n- **Don't** end a multi-paper search by telling the user \"set\n  `ANTHROPIC_API_KEY` for a rich deck\" — you are the LLM that could\n  have written the summaries.\n- **Don't** treat the per-paper lightweight `.pptx` as the\n  deliverable.\n- **Don't** stop after `download_pdfs` reports N PDFs saved — that's\n  the start of the rich-authoring phase, not the end.\n- **Don't** invent numbers, RQs, contributions, or limitations not in\n  the paper.\n- **Don't** fabricate URLs / DOIs / arXiv IDs — see the rule above.\n- **Don't** leave irrelevant downloads in the run directory. Keyword\n  search matches can include off-topic papers (a \"Claude code\" query\n  pulled in a Viterbi-decoder paper; \"LLM code review\" pulled in an\n  object-detection literature review). After classifying papers\n  off-topic, delete their `pdfs/\u003ckey\u003e.pdf` and lightweight\n  `\u003ckey\u003e.pptx`; keep the aggregate xlsx / bib as the honest record of\n  what the search returned.\n- **Don't** mention \"Claude\", \"Claude Code\", \"AI-generated\", \"GPT\",\n  \"Copilot\", or any AI tool/model name in commit messages, PR\n  descriptions, code comments, or documentation.\n\n## Features\n\n- **Fifteen pluggable sources**: `arxiv`, `semantic_scholar`, `openalex`,\n  `pubmed`, `acm` (Crossref-scoped), `dblp`, `crossref` (unscoped),\n  `openaire`, `springer` (needs API key), `europepmc` (open, no key —\n  life-sciences + preprints + agriculture), `doaj` (open, no key —\n  open-access journals, usually with a direct PDF link), `hal` (open,\n  no key — France's CS / maths / physics archive with full-text PDFs),\n  `core` (needs free API key — largest open-access aggregator, 250M+\n  works), `ieee` (default-on via visible Chrome; API key adds official\n  Xplore API), `scholar` (default-on via visible Chrome). Each lives in\n  `sources/\u003cname\u003e/` behind a `Fetcher`\n  adapter. A top-tier-venue whitelist filters results to flagship CS\n  conferences/journals plus Nature/Science/PNAS by default; pass\n  `--all-venues` to disable.\n- **Single-paper mode**: paste an arXiv ID, arXiv URL, DOI, PMID, or IEEE\n  document URL — ThesisAgents resolves it via the right source and\n  emits the same export bundle. Useful for paper reading notes and thesis\n  defence prep.\n- **Local PDF mode** (`--pdf \u003cpath\u003e`): pass one PDF or a directory.\n  A heuristic extractor pulls **title, authors, year, arXiv ID, DOI, and\n  the real abstract** straight from each PDF's front matter (anchored\n  on the explicit `Abstract` / `ABSTRACT` / `摘要` header, not a blind\n  prefix). `--title` / `--authors` / `--year` / `--venue` / `--doi` /\n  `--arxiv-id` override on a single-PDF call; on a directory, per-file\n  extraction wins so every paper gets its own deck named after its\n  BibTeX key.\n- **Eight exporters**:\n  - `.pptx` — 16:9 widescreen, page-numbered, three rendering tiers\n    (lightweight abstract-only · enriched-flat · **thesis-style** with\n    pain-point quadrants, KPI callouts, technique-comparison tables,\n    per-RQ result tables, contribution summary, core observation,\n    limitations \u0026 future work, Q\u0026A, references). All template strings\n    are i18n'd across **14 languages**: English, 繁體中文, 简体中文,\n    日本語, Español, Français, Deutsch, 한국어, Português, Русский,\n    Italiano, Tiếng Việt, हिन्दी, Bahasa Indonesia.\n  - **Designed-deck visual identity** (not the default Calibri-on-white\n    look): per-language typography (Inter for Latin, Microsoft JhengHei\n    UI / YaHei UI / Yu Gothic UI / Malgun Gothic / Nirmala UI for CJK\n    + Hindi), programmatic accent geometry (top accent bar on every\n    content slide + left band on the cover), academic-style table\n    formatting (default grid stripped, navy header rule, soft inter-row\n    dividers, alternating row stripe, middle-vertical alignment, bold\n    row labels), and a five-colour palette discipline (navy / teal /\n    grey / light / white) with red **banned** for text (use bold +\n    teal `#0E7490` for emphasis instead).\n  - **Dark mode is the default render path.** Build runs on the light\n    palette, then a post-build pass swaps text + fill + cell-border\n    RGBs to a dark deck (slide bg `#12151B`, body text `#E5E7EB`,\n    accent teal swapped to a brighter `#2DD4BF`). OLED projectors and\n    low-light venues get the dark deck without authors having to think\n    about it; pass `--light-mode` (CLI), uncheck **Light mode** (GUI\n    Deck tab), or `ExportOptions(dark_mode=False)` (programmatic) to\n    opt out for projectors in well-lit rooms or for printed handouts.\n  - `.xlsx` — Papers sheet + Query provenance sheet, hyperlinked URL /\n    PDF, frozen header, auto column widths. Column 5 (**Source**) shows\n    the real publication venue (e.g. \"IEEE Access\"); column 6\n    (**Indexed via**) shows which fetcher returned the metadata\n    (e.g. \"openalex\"), so the two pieces of information never collide.\n  - `.md` — full source / title / abstract list.\n  - `.bib` — collision-free citation keys, LaTeX-escaped fields.\n  - `.json` — raw payload for downstream tooling.\n  - `.ris` — RIS interchange imported by Zotero / Mendeley / EndNote /\n    RefWorks (the BibTeX sibling for non-LaTeX reference managers).\n  - `.csv` — flat one-row-per-paper table for spreadsheets / quick grep\n    triage (RFC-4180 quoting, so commas in titles never shift columns).\n  - `.csl.json` — CSL-JSON for Pandoc / citeproc; render a bibliography\n    in any CSL style (APA, IEEE, Nature, …). The `.csl.json` extension\n    keeps it distinct from the plain `.json` dump.\n- **PPT editing toolkit**: `thesisagents.exporters.pptx_edit`\n  (inspect / update_slide / delete_slide / reorder_slides / add_slide)\n  works against any deck the exporter produces, plus the equivalent\n  `pptx_*` MCP tools so an LLM agent can iterate on a generated deck.\n- **MCP server**: 12 tools — `list_sources` + `list_exports`\n  (discovery), `search`, `fetch_paper`, `fetch_pdf_text`,\n  `download_pdfs`, `export`, and the five `pptx_*` editing tools. Lets\n  any MCP-aware LLM\n  (Claude Code, Claude Desktop, Cursor, …) drive the whole workflow.\n- **Two enrichment paths** for going beyond the abstract into a true\n  thesis-style deck:\n  - **LLM-as-agent (no API key)** — the calling LLM reads the PDF body\n    text via `fetch_pdf_text`, writes a structured summary in-context,\n    and passes it to `export`.\n  - **Python pipeline (`--enrich`)** — the CLI calls Anthropic's API\n    itself; default model `claude-opus-4-7`.\n- **Visible-Chrome publisher flows**: Scholar SERP, IEEE `/rest/search`,\n  and every paywalled-PDF download (ieeexplore / dl.acm / link.springer\n  / sciencedirect / wiley / oup / nature / science / …) run inside a\n  real visible Chrome session via `selenium`. The user solves captcha\n  / completes SSO in the live window once; `THESISAGENTS_CHROME_PROFILE_DIR`\n  persists the cookies across runs.\n- **LLM-as-agent flow** (`scripts/llm_*.py`): when the LLM in your editor\n  wants to drive the browser itself (rather than let `asyncio.gather` do\n  it), `scripts/llm_driven_search.py` opens Chrome on Scholar + IEEE,\n  `scripts/llm_download_pdfs.py` walks an xlsx and downloads every paper\n  in one Chrome session (IEEE / ACM / Springer / arXiv / ACL Anthology /\n  NeurIPS / OpenReview), and `scripts/regen_*.py` shows the worked\n  pattern for hand-authoring a rich `PaperSummary` per paper.\n- **OA PDF resolver**: post-dedup, every paper without `pdf_url`\n  goes through Unpaywall → S2 `openAccessPdf` → arXiv title search →\n  CORE.ac.uk (when keys are set). Typical lift on IEEE / ACM / Springer\n  / Elsevier-heavy queries: 40-70 percentage points.\n- **Safety by default**: HTTPS-only HTTP transport, per-source rate\n  limit (token bucket), `defusedxml` for any XML payload,\n  path-traversal-safe export paths, no `eval` / `exec` / `pickle` on\n  user input.\n- **zh-tw / zh-cn vocabulary guard**: ~244 regex patterns in\n  `tests/test_i18n.py::test_zh_tw_files_use_traditional_chinese_vocabulary`\n  catch Simplified-Chinese loan words rendered with Traditional hanzi\n  (e.g. `內存` → `記憶體`, `魯棒性` → `穩健性`, `軟件` → `軟體`,\n  `緩存` → `快取`). Same guard runs in reverse for zh-cn locale\n  strings. Full rule + the regex catalogue live in\n  `.claude/agents/rules/language-vocabulary-check.md`.\n\n## Quick start\n\n```powershell\ngit clone \u003crepo-url\u003e\ncd ThesisAgents\npython -m venv .venv\n.venv\\Scripts\\Activate.ps1            # Windows PowerShell\n# source .venv/bin/activate           # Linux / macOS\n\n# Install with dev extras (also pulls in MCP SDK and intelligence deps)\npip install -e .[dev]\n```\n\nSearch arXiv and export deck + workbook + BibTeX (default for `--query`):\n\n```powershell\npy -m thesisagents --query \"diffusion models\" --source arxiv --max 10 `\n                      --out .\\exports\\\n```\n\nFetch a single paper by URL — defaults to `.pptx + .bib` (the `.xlsx`\nmakes less sense for one row):\n\n```powershell\npy -m thesisagents --paper \"https://arxiv.org/abs/1706.03762\" `\n                      --filename-stem attention `\n                      --out .\\exports\\\n```\n\nRender the deck in 繁體中文:\n\n```powershell\npy -m thesisagents --paper \"https://arxiv.org/abs/1706.03762\" `\n                      --lang zh-tw --out .\\exports\\\n```\n\nLLM-pipeline enrichment (Python calls Anthropic itself — needs API key):\n\n```powershell\n$env:ANTHROPIC_API_KEY = \"sk-ant-...\"\npy -m thesisagents --paper \"https://arxiv.org/abs/1706.03762\" `\n                      --enrich --lang zh-tw --out .\\exports\\\n```\n\n## CLI flags\n\n| Flag | Purpose |\n|---|---|\n| `--query` / `-q` | Keywords (required unless `--paper`). |\n| `--paper` / `-p` | arXiv ID / URL, DOI, PMID, or IEEE document URL. Mutually exclusive with `--query`. |\n| `--source` / `-s` | Comma-separated source list. Default `arxiv`. |\n| `--max` / `-n` | Max results per source (1..200). Default 25. |\n| `--year-from` / `--year-to` | Inclusive year filter. |\n| `--export` / `-e` | Formats: any of `pptx,xlsx,md,bib,json,ris,csv,csl`. Default depends on mode (see below). |\n| `--out` / `-o` | Output directory. Default `./exports`. |\n| `--filename-stem` | Override the generated filename stem. |\n| `--no-abstract` | Omit abstract content from exports. |\n| `--lang` / `-l` | Deck language: one of 14 — `en`, `zh-tw`, `zh-cn`, `ja`, `es`, `fr`, `de`, `ko`, `pt`, `ru`, `it`, `vi`, `hi`, `id`. Default `en`. |\n| `--enrich` | Fail-loud variant of auto-enrich. Needs `ANTHROPIC_API_KEY` and `[intelligence]` extra. (Auto-enrich is default when the key is set.) |\n| `--lightweight` | Skip enrichment + force the abstract-only deck. Use only for quick / unattended runs; **when an LLM agent is driving, prefer the LLM-as-agent flow** below. |\n| `--llm-model` | Override default `claude-opus-4-7` for enrichment. |\n| `--no-pdf` | Skip the automatic PDF download. Also disables the per-paper PPT gate (no PDF → no full content). |\n| `--no-oa-resolve` | Skip the post-dedup OA PDF resolver (Unpaywall + S2 + arXiv + CORE.ac.uk). |\n| `--top-tier-only` | Restrict results to arXiv + a curated CS-flagship whitelist (S\u0026P, CCS, NDSS, USENIX Security, NeurIPS, ICML, ICSE, …). Off by default. |\n| `--paywall-threshold` | Fraction of paywalled results that triggers the confirmation prompt. Default 0.30. |\n| `--yes` | Skip the paywall prompt and proceed. |\n| `--max-slides` | Per-paper slide cap (default 25; pass 0 for unlimited). |\n| `--light-mode` | Render the pptx with a white background + navy text. Default is dark mode (dark background + near-white text) — pass this for projectors in well-lit rooms or when the deck will be printed. |\n| `--quiet` | Suppress per-paper printout. |\n\n### Environment variables\n\n| Variable | Used by | Purpose |\n|---|---|---|\n| `ANTHROPIC_API_KEY` | `--enrich` | LLM auth. Not needed for the LLM-as-agent path over MCP. |\n| `THESISAGENTS_LLM_MODEL` | `--enrich` | Override the default `claude-opus-4-7`. |\n| `THESISAGENTS_S2_API_KEY` | Semantic Scholar + OA resolver | Higher rate limit; also used by the OA resolver's S2 `openAccessPdf` step. Free key at \u003chttps://www.semanticscholar.org/product/api\u003e. |\n| `THESISAGENTS_NCBI_API_KEY` | PubMed | Raises NCBI's anonymous limit (3/s) to 10/s. Optional. |\n| `THESISAGENTS_CONTACT_EMAIL` | PubMed, ACM, Crossref, OpenAlex, **Unpaywall** | Polite-pool tag + enables the OA resolver's Unpaywall step (biggest PDF-coverage win for IEEE / ACM / Springer / Elsevier-paywalled papers; typical lift 40-70 pp). |\n| `THESISAGENTS_IEEE_API_KEY` | IEEE (API path) | Official IEEE Xplore API; surfaces `pdf_url` for in-scope papers. |\n| `THESISAGENTS_DISABLE_IEEE_SCRAPING` | IEEE | **IEEE is default-ON via visible Chrome.** Set `=1` to opt out (e.g. CI without Chrome). The httpx scrape branch only runs as a fallback when WebRunner is unavailable. |\n| `THESISAGENTS_CROSSREF_PLUS_TOKEN` | ACM, Crossref | Crossref Plus subscriber token (Bearer header). Optional. |\n| `THESISAGENTS_SPRINGER_API_KEY` | Springer | Required; free key from \u003chttps://dev.springernature.com/\u003e. Plugin raises `ConfigError` without it. |\n| `THESISAGENTS_DISABLE_SCHOLAR_SCRAPING` | Google Scholar | **Scholar is default-ON via visible Chrome.** Set `=1` to opt out (Google's ToS forbids automated access — default-on for coverage, opt-out to avoid captcha / IP-block risk). |\n| `THESISAGENTS_CHROME_PROFILE_DIR` | Scholar + IEEE + paywalled-PDF downloads | Persistent Chrome `--user-data-dir`. Set this and complete VPN / SSO / Google sign-in once; subsequent runs inherit the cookies so IEEE returns paywalled metadata and Scholar serves un-throttled SERPs. |\n| `THESISAGENTS_DISABLE_WEBRUNNER` | Scholar + IEEE + paywalled-PDF downloads | `=1` forces the httpx paths instead of driving real Chrome. Useful for CI / Docker without a Chrome binary; otherwise leave unset. |\n| `THESISAGENTS_CORE_API_KEY` | OA resolver + `core` search source | Free key from \u003chttps://core.ac.uk/services/api\u003e. Enables the CORE.ac.uk OA-lookup step (200M+ institutional / regional OA items) **and** the `core` search source. Without it, the `core` source is silently skipped and the other OA strategies (Unpaywall, S2, arXiv) still run. |\n| `THESISAGENTS_PDF_COOKIES_FILE` | PDF downloader | Netscape `cookies.txt`. Off by default. Use only with publishers you have institutional rights to. |\n| `THESISAGENTS_LOG_LEVEL` | logger | `INFO` default; `DEBUG` for verbose tracing. |\n\nDefaults: `--query` → `pptx,xlsx,bib`. `--paper` → `pptx,bib`. Always\noverridable with explicit `--export`.\n\n## LLM-as-agent flow\n\nWhen an LLM in your editor (Claude Code, Cursor, Aider, Codex CLI, …)\nwants to drive the publisher browser itself — pick URLs, inspect the\nreturned DOM, decide which papers to dig into — five scripts under\n`scripts/` cover the canonical path:\n\n| Script | What it does |\n|---|---|\n| `scripts/llm_driven_search.py \"\u003cquery\u003e\"` | Boots visible Chrome, navigates Scholar SERP for the query, JS-fetches IEEE `/rest/search` from inside the IEEE origin, dumps SERP HTML + IEEE JSON to `exports/_llm_scratch/`. |\n| `scripts/llm_parse_results.py` | Reads the dumped artefacts, runs the project's parsers, dedups + ranks + exports `.xlsx` + `.md` for the LLM to inspect. |\n| `scripts/llm_download_pdfs.py \u003cxlsx\u003e` | Walks the xlsx, dispatches each row to the right per-publisher downloader (IEEE / ACM / Springer / arXiv / ACL Anthology / NeurIPS / OpenReview) in ONE Chrome session. Idempotent: papers with a valid `\u003cid\u003e.pdf` already on disk skip immediately. |\n| `scripts/llm_download_{ieee,acm,springer}_pdf.py \u003cid\u003e` | Single-paper variants for iterating on selectors / debugging one entry. |\n| `scripts/regen_*.py` | Worked example of hand-authored rich `PaperSummary` per paper → rich-tier `.pptx`. Look at `scripts/regen_speculative_decoding_zh_tw.py` for the canonical shape. |\n\nFull end-to-end runbook (search → rich deck) lives in\n`.claude/agents/tasks/paper-summary-author.md` — open it before starting a\nnew query so the LLM can run the flow without pausing for user input.\n\n## MCP server\n\nRegister with Claude Code:\n\n```powershell\nclaude mcp add thesisagents -- \".venv\\Scripts\\python.exe\" -m thesisagents.mcp\n```\n\nOr write to your settings file:\n\n```json\n{\n  \"mcpServers\": {\n    \"thesisagents\": {\n      \"command\": \".venv\\\\Scripts\\\\python.exe\",\n      \"args\": [\"-m\", \"thesisagents.mcp\"]\n    }\n  }\n}\n```\n\nTools:\n\n| Tool | Purpose |\n|---|---|\n| `list_sources` | Enumerate every plugin + report whether each is enabled in the current env. Call this once before `search`. |\n| `search` | Keywords → list of papers. Accepts `top_tier_only`, `min_citations`; defaults to the full no-API-key source mix. |\n| `fetch_paper` | arXiv / DOI / PMID / IEEE identifier → single paper. |\n| `fetch_pdf_text` | Download one PDF, return extracted body text. **The MCP path to \"I read the paper\".** |\n| `download_pdfs` | Batch-download a papers list's PDFs into `{out_dir}/pdfs/`. Returns per-paper results keyed by BibTeX key. |\n| `export` | Papers list + formats → writes `.pptx/.xlsx/.md/.bib/.json/.ris/.csv/.csl.json`. Accepts a `summary` field per paper for the rich thesis-style schema, `max_slides_per_paper` (default 25), and `dark_mode` (default `false` — the project default is the light navy-band deck, pass `true` for the dark OLED / low-light post-pass). |\n| `pptx_inspect` | Read slide / shape structure of an existing deck. |\n| `pptx_review` | Audit a deck in one call — overflow + colour contracts + `paper_rule` section completeness. Auto-detects the deck language; also the CLI `python -m thesisagents review \u003cdeck.pptx\u003e`. |\n| `pptx_update_slide` | Replace `title` / `body` / `meta` (by shape name) or arbitrary shapes by index. |\n| `pptx_delete_slide` | Remove a slide and its part relationship. |\n| `pptx_reorder_slides` | Permute slides via `sldIdLst`. |\n| `pptx_add_slide` | Append or insert a new title / body / meta slide. |\n\nLLM-as-agent flow (no `ANTHROPIC_API_KEY` needed — the LLM is the agent):\n\n```\n1. (optional) list_sources()                       # discover enabled plugins\n2. search(keywords=..., sources=[...], top_tier_only=true)\n3. (optional) download_pdfs(papers, out_dir=\"./exports/...\")  # persist PDFs\n4. fetch_pdf_text(pdf_url=paper.pdf_url)           # per paper\n5. (the LLM reads body text, produces a structured `summary` dict)\n6. export(papers=[{...paper, \"summary\": {pain_points: [...], rq_results: [...]}}],\n          language=\"zh-tw\", formats=[\"pptx\",\"bib\"], dark_mode=true, ...)\n```\n\nFull reference in [`docs/mcp.md`](docs/mcp.md).\n\n## Project layout\n\n```\nThesisAgents/\n├── thesisagents/                 # main package\n│   ├── core/                        # Paper / PaperSummary / RqResult / dedup / ranking / pipeline\n│   ├── fetchers/                    # HTTPS-only async client, token-bucket rate limit\n│   ├── exporters/                   # pptx (thesis-style) · xlsx · bib · md · json · ris · csv · csl · pptx_edit · i18n\n│   ├── intelligence/                # PDF fetch + Anthropic summariser  ([intelligence] extra)\n│   ├── mcp/                         # FastMCP server (12 tools)\n│   ├── utils/                       # logging, path safety\n│   ├── cli.py                       # argparse CLI\n│   └── __main__.py\n├── sources/                         # plugin folders: arxiv, semantic_scholar,\n│                                    #   openalex, pubmed, acm, ieee, scholar,\n│                                    #   dblp, crossref, openaire, springer,\n│                                    #   europepmc, doaj, hal, core\n├── tests/                           # pytest suite + recorded fixtures (no live HTTP)\n├── docs/                            # Sphinx (14 language trees)\n├── scripts/                         # one-off regen scripts\n└── pyproject.toml                   # ruff, bandit, build, optional extras\n```\n\n## Definition of Done\n\n```powershell\n.venv\\Scripts\\python.exe -m pytest tests/\n.venv\\Scripts\\python.exe -m ruff check .\n.venv\\Scripts\\python.exe -m bandit -c pyproject.toml -r thesisagents/ sources/\n```\n\nThe `-c` flag on bandit is required — without it bandit ignores the\nproject skip config. When touching the pptx exporter, also run an\noverflow check (see `CLAUDE.md` \"Slide Deck Rules\").\n\n## Desktop GUI (PySide6)\n\nA native desktop interface ships behind the `[gui]` extra:\n\n```powershell\npip install thesisagents[gui]\nthesisagents-gui                 # or: thesisagents gui\n```\n\nThe window has four tabs — **Search**, **Settings** (persists API keys\nvia QSettings), **Enrich** (drives the LLM-as-agent / Python-pipeline\nenrichment over a `collection_ready` signal), and **Deck** (the Light\nmode toggle + slide-cap + max-figures controls flow through to\n`ExportOptions`). The Windows release zip ships the Nuitka-compiled\nbundle with PySide6 included, so `thesisagents.exe gui` works\nwithout a separate Python install.\n**UI ships in all 14 languages** (English, 繁體中文, 简体中文,\n日本語, Español, Français, Deutsch, 한국어, Português, Русский,\nItaliano, Tiếng Việt, हिन्दी, Bahasa Indonesia) — first run picks\nthe language from your OS locale, then **Settings → Interface\nlanguage** lets you change it. The deck output language is a\nseparate dropdown so you can run the UI in one language and emit\nslides in another. The layout is responsive: every form sits in\na `QScrollArea` and the window resizes down to 900×600 (still fits\n720p), with HiDPI scaling on by default.\n\nFull reference: [`docs/gui.md`](docs/gui.md).\n\n## Packaging as a standalone executable\n\nTwo packagers are documented for shipping a single-file binary that\nruns without Python installed:\n\n- **[`docs/packaging-pyinstaller.md`](docs/packaging-pyinstaller.md)**\n  — fast build (under a minute), 200–300 MB output, 2–4 s startup.\n  Best when you iterate on the build script.\n- **[`docs/packaging-nuitka.md`](docs/packaging-nuitka.md)** —\n  slow build (5–15 minutes), 80–150 MB output, sub-second startup,\n  some bytecode protection. Best when end users run the binary\n  many times.\n\nBoth docs cover the project-specific gotcha — the dynamic source\nplugins under `sources/\u003cname\u003e/` — and ship a verified command for\nthe CLI and the MCP server entry points.\n\n## Continuous integration \u0026 releases\n\nTwo GitHub Actions workflows live under `.github/workflows/`:\n\n- **`ci.yml`** runs on every push and PR to `main`. Matrix is Ubuntu +\n  Windows × Python 3.12 / 3.13 / 3.14 (6 jobs). Each job runs\n  `ruff check`, `bandit -c pyproject.toml`, and `pytest`.\n- **`release.yml`** waits for `ci.yml` to complete on `main`\n  (`workflow_run` trigger). It runs only if CI succeeded. **Every\n  CI-success push to `main` is a release** — the workflow auto-bumps\n  the patch version in `pyproject.toml`, commits the bump back to\n  `main` as `chore: bump version to X.Y.Z`, and pipelines:\n  1. **`bump-version`** — read current `X.Y.Z` from `pyproject.toml`,\n     increment to `X.Y.(Z+1)`, commit + push back to `main` using the\n     workflow `GITHUB_TOKEN`. That push does NOT re-trigger CI (per\n     GitHub's rule that `GITHUB_TOKEN`-driven pushes can't start new\n     workflow runs), so the cycle terminates naturally.\n  2. **`publish-pypi`** — build sdist + wheel, `twine check`,\n     `twine upload` via `PYPI_API_TOKEN`.\n  3. **`create-draft-release`** — open a *draft* GitHub release at\n     tag `v\u003cversion\u003e` with auto-generated notes.\n  4. **`build-nuitka`** — compile a Nuitka standalone bundle on a\n     Windows runner (entry point: `python -m thesisagents` via\n     `--python-flag=-m`), smoke-test it, zip the resulting\n     `thesisagents.dist/` folder, and attach the zip + a `.sha256`\n     checksum to the draft release. Standalone (not onefile) by\n     design: onefile self-extracts to `%TEMP%` on every launch,\n     adding startup latency and tripping antivirus heuristics on\n     locked-down machines. Windows-only by design too: Linux / macOS\n     users install from PyPI. Build cache keyed on `pyproject.toml`\n     cuts warm builds from ~70 min cold to ~5–10 min.\n  5. **`publish-release`** — unmark the draft once the Nuitka asset\n     is uploaded, so users never see a half-finished release.\n\n  **Skipping a release.** Include `[skip release]` anywhere in the\n  commit message and the bump + every downstream job is skipped — use\n  this for docs-only / typo / refactor commits that shouldn't burn a\n  version number.\n\nTo enable PyPI publishing + release executables:\n\n1. Generate a project-scoped API token at\n   \u003chttps://pypi.org/manage/account/token/\u003e.\n2. In the GitHub repo: `Settings → Secrets and variables → Actions →\n   New repository secret`. Name it `PYPI_API_TOKEN` and paste the\n   token value.\n3. Allow GitHub Actions to push to `main`: `Settings → Actions →\n   General → Workflow permissions → Read and write permissions`. The\n   bump commit is pushed by the workflow's `GITHUB_TOKEN`.\n4. Cut releases by merging PRs into `main`. The pipeline takes\n   ~3–5 min to publish to PyPI and ~50–70 min more (cold) or ~5–10 min\n   (warm Nuitka cache) for the Windows zip to attach.\n\nThe `publish-pypi` job intentionally does NOT attach a GitHub\nEnvironment, so each run surfaces as a Release entry (with its\nNuitka `.exe` attached) rather than as a \"Deployment\" sidebar\nwidget on the repo home — releases get their own dedicated page\nand a Deployment entry on top would just be redundant noise.\n\n## License\n\nSee `LICENSE`. The arXiv API is used under arXiv's API terms of use\n(\u003chttps://info.arxiv.org/help/api/tou.html\u003e) — observe the 1 request\nper 3 seconds soft limit; the bundled fetcher already enforces this via\nits token bucket.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fintegration-automation%2Fthesisagent","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fintegration-automation%2Fthesisagent","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fintegration-automation%2Fthesisagent/lists"}