A document conversion and publishing toolkit for an org-centric knowledge workflow.
memex-kb started as a Google Docs → Denote knowledge base converter and has grown into a broader toolbox for turning legacy or platform-bound content into plain-text, version-controlled, and AI-friendly formats.
Today the repository combines three layers:
- Knowledge ingestion — Google Docs, Threads, Confluence, GitHub Stars, and blog exports
- Structured document pipelines — proposal workflows, Org/ODT/HWP-oriented transformations
- Reusable publishing templates — paper, presentation, and PowerPoint template injection workflows
The guiding idea is simple:
Legacy content → structured text → reproducible artifacts → human + AI collaboration
memex-kb is useful when you want to:
- convert documents into Markdown, Org-mode, BibTeX, ODT, DOC, PDF, or PPTX-adjacent workflows
- preserve structure well enough for search, versioning, and AI-assisted editing
- standardize output with Denote-style naming, rule-based classification, and template-driven publishing
- keep the whole workflow reproducible with Nix Flakes and CLI-first tooling
| Backend / Source | Status | Main entry point | Output |
|---|---|---|---|
| Google Docs | Stable | scripts/gdocs_md_processor.py, ./run.sh gdocs-export |
Markdown, DOCX, PDF, HTML, TXT |
| Threads | Stable | scripts/threads_exporter.py, ./run.sh threads-export |
Org-mode + images |
Confluence export (.doc MIME HTML) |
Stable | scripts/confluence_to_markdown.py |
Clean Markdown |
| GitHub Stars | Stable | scripts/gh_starred_to_bib.sh, ./run.sh github-starred-export |
BibTeX |
| Naver Blog | Active | scripts/naver_blog_crawler.py, ./run.sh naver-* |
Denote-style Org + assets |
| Anthropic Distill HTML papers | Active | scripts/anthropic_paper_to_org.py, ./run.sh paper2org / paper2org-html / paper2org-pdf |
Org (math/figure/citation-aware) → citeproc HTML + ArXiv-style acmart PDF |
| HWPX / OWPML related workflows | Active | hwpx2org/, orgadoc2odt/, proposal-pipeline/ |
Org, ODT, DOC, HWP-oriented outputs |
| Template / Pipeline | Purpose |
|---|---|
templates/arxiv-acm/ |
Org-mode → ACM acmart → PDF / ArXiv-ready source workflow |
templates/presentation/ |
Quarto / Reveal.js HTML presentation template |
templates/presentation-pptx/ |
org2pptx: inject Org-mode content into an existing PPTX template while preserving layout/design |
proposal-pipeline/ |
Google Docs → Markdown → Org-mode → ODT/DOC proposal workflow |
.claude/skills/syndicate/ + scripts/syndicate.py / ./run.sh syndicate |
ROSSE syndication (issue #4): one garden canonical note → one copy-paste bundle per surface (raw / full / summary classes) for Facebook, LinkedIn, Naver Blog, Tistory, Threads, X, Bluesky, Instagram. Copy-paste / browser-Claude first, not full API automation |
.claude/skills/scanbook/ |
Repo-local operating manual for scanned-book work. New agents should read this first for MinerU server checks, per-book config, correction strategy, and EPUB gotchas |
./run.sh mineru-setup / mineru-parse |
Scanned PDF → MinerU VLM Markdown + content_list.json + images, using the remote gpu2i vLLM server through an SSH tunnel |
scripts/mineru2org.py + scripts/corrections/*.json |
MinerU Markdown → clean Org: structure recovery, footnotes, images, LaTeX, EPUB metadata, and book-specific corrections |
./run.sh diff-review |
Engine-agnostic QA helper for comparing two transcriptions and surfacing only conflicts |
./run.sh org2epub-build |
Org → clean EPUB 3.0 using the maintained local ~/repos/gh/ox-epub fork directly; supports images, LaTeX→SVG math, tables, footnotes, TOC, Korean, and epubcheck validation |
scanpdf2org/ |
Older page-render + vision-transcription surface; kept as a fallback/oracle path, not the primary scanned-book pipeline |
memex-kb/
├── README.md
├── AGENTS.md
├── BACKENDS.md
├── DEVELOPMENT.md
├── DENOTE-RULES.md
├── run.sh # Primary command entry point
├── flake.nix # Reproducible dev environment
├── .claude/skills/syndicate/ # Repo-local ROSSE syndication operating manual
├── .claude/skills/scanbook/ # Repo-local scanned-book → EPUB operating manual
├── .claude/skills/anthropic-paper2org/ # Anthropic Distill paper capture/export manual
├── .pi/settings.json # Loads repo-local skills for pi sessions
├── config/ # Local env/config templates
├── scripts/ # Main backend and utility scripts
│ ├── adapters/
│ ├── gdocs_md_processor.py
│ ├── threads_exporter.py
│ ├── confluence_to_markdown.py
│ ├── gh_starred_to_bib.sh
│ ├── md_to_gdocs.py
│ ├── md_to_gdocs_html.py
│ ├── mineru2org.py
│ └── naver_blog_crawler.py
├── mineru-client/ # Thin local client for remote gpu2i MinerU vLLM
├── templates/
│ ├── arxiv-acm/
│ ├── presentation/
│ └── presentation-pptx/
├── proposal-pipeline/ # Proposal authoring and export pipeline
├── scanpdf2org/ # Older scanned PDF → page render → vision fallback
├── epub2org/ # EPUB → Org (reverse direction)
├── hwpx2org/ # HWPX/Org-related conversion utilities
├── orgadoc2odt/ # AsciiDoc/ODT conversion utilities
├── office/ # Real project working materials and samples
├── docs/ # Converted output and project notes
└── logs/ # Execution logs
scripts/: the main place for backend integrations and conversion entry pointstemplates/: reusable starter templates for papers and presentationsproposal-pipeline/: the most opinionated end-to-end workflow in the repo.claude/skills/scanbook/: the durable operating guide for scanned-book → EPUB work; read before touchingscanpdf/work/<book>/.claude/skills/anthropic-paper2org/: the durable operating guide for Anthropic Distill HTML paper → Org / HTML / PDF workmineru-client/,scripts/mineru2org.py,scripts/corrections/*.json: the current MinerU → Org → EPUB pathoffice/: practical working examples and proposal artifactshwpx2org/andorgadoc2odt/: lower-level format conversion experiments and tools
This project uses Nix Flakes.
Use one of the following:
# interactive shell
nix develop
# one-off command
nix develop --command python scripts/threads_exporter.py --download-images
# recommended for regular work
direnv allow- reproducible dependencies
- no ad-hoc
pip install - consistent Python / Pandoc / CLI tooling
- easier agent automation
./run.sh./run.sh gdocs-export <DOC_ID>
./run.sh gdocs-export <DOC_ID> --format md
./run.sh gdocs-export <DOC_ID> --format docx --depth 0./run.sh threads-export --download-images
./run.sh threads-export --max-posts 5 --download-images./run.sh confluence-convert document.doc
./run.sh confluence-batch ./input-dir ./output-dir./run.sh github-starred-export
./run.sh github-starred-export ~/org/resources/github-starred.bib./run.sh proposal-build --export-md
./run.sh proposal-merge --strip-hwpx-idx --org-tables
./run.sh proposal-export-odt./run.sh arxiv-build
./run.sh arxiv-build templates/arxiv-acm/sample.orgURL="https://transformer-circuits.pub/2026/workspace/index.html"
./run.sh paper2org "$URL" --name jspace --fetch
./run.sh paper2org-html "$URL" --name jspace
./run.sh paper2org-pdf "$URL" --name jspaceA complete sample for:
- Org-mode authoring
- ACM
acmartLaTeX export - PDF generation suitable for paper drafting / ArXiv submission workflows
See: templates/arxiv-acm/README.md
A Quarto / Reveal.js presentation starter for browser-based slide decks.
See: templates/presentation/README.md
A newer org2pptx pipeline for teams that must submit or reuse a branded PowerPoint template.
Instead of rendering slides from scratch, it:
- parses an Org file
- injects content into an existing
.pptxtemplate - preserves original slide backgrounds, logos, layouts, and branding
This is especially useful when pandoc --reference-doc or layout-name-based approaches fail on localized corporate templates.
See: templates/presentation-pptx/README.md
- Need Google Docs tabs exported cleanly → use
gdocs-export - Need social writing archived into Org → use
threads-export - Need legacy Confluence exports cleaned up → use
confluence-convert - Need citation-ready GitHub Stars → use
github-starred-export - Need proposal submission artifacts → use
proposal-pipeline/ - Need a paper PDF from Org → use
templates/arxiv-acm/ - Need an Anthropic Distill paper in Org/HTML/PDF → use
paper2org,paper2org-html, andpaper2org-pdf - Need HTML slides → use
templates/presentation/ - Need content injected into an existing company PPTX → use
templates/presentation-pptx/
| File | Purpose |
|---|---|
AGENTS.md |
Working guidance for coding agents and maintainers |
BACKENDS.md |
Backend-specific notes and usage details |
DEVELOPMENT.md |
Development guidance for extending the project |
DENOTE-RULES.md |
Naming and structuring rules for Denote-style output |
proposal-pipeline/README.md |
Detailed proposal workflow documentation |
office/README.md |
Real-world working context and example materials |
See CHANGELOG.md for CalVer snapshots. Recent highlights include:
- Anthropic Distill paper capture: HTML → Org → citeproc HTML / acmart PDF
- ROSSE syndication bundles for copy-paste publishing surfaces
- MinerU / OCR scanned-book pipelines and clean EPUB generation
- Template workflows for acmart papers, Reveal.js presentations, and PPTX injection
MIT