# π Research Digest
A local-first arXiv feed with AI-written plain-language summaries β running entirely on your own hardware.
Curate research interests, and Research Digest keeps a local corpus of matching papers, each with a plain-English summary, a layman explanation, topical tags, and semantically-related papers. Desktop grid for deep reading, mobile feed for quick scrolling, full-text filtering on both. No API keys, no cloud. Summaries and embeddings are generated by turbolab, a self-hosted OpenAI-compatible model server.
# β¨ Features
- π§ Real AI summaries β summary, layman explanation, difficulty, and tags from your own LLM via turbolab. Nothing leaves your network.
- π Search & relate β client-side keyword filter on every page; βrelated papersβ precomputed from e5 embeddings.
- π Corpus-first β SQLite is the source of truth. The site renders from the DB, so a failed arXiv fetch never wipes good output.
- π‘ Resilient pipeline β 429 backoff, upsert-never-delete, and a render step that refuses to publish an empty digest.
- π± Desktop + mobile β multi-column grid and a full-screen swipeable feed.
- βοΈ Configurable β JSON interests, keyword scoring, look-back window.
# π§© How it works
SQLite holds the corpus. The pipeline is a chain of independent, idempotent stages β and only the first one touches the network:
| Stage | Network? | Does |
|---|---|---|
fetch.py |
arXiv | Upsert new papers (original abstract). 429 backoff; never deletes. |
summarize.py |
turbolab | Fill missing summaries/layman/difficulty/tags via /v1/chat/completions. |
embed.py |
turbolab | Fill missing vectors via /v1/embeddings (e5, with passage: prefix). |
relate.py |
β | Precompute nearest-neighbour papers (cosine) for βrelatedβ. |
render.py |
β | Build the static site from the DB. Atomic writes; refuses to publish empty. |
run.sh chains them; a fetch failure is logged and the rest still run on the existing corpus.
Because everything after fetch is offline, a multi-day arXiv rate limit just means βno new
papersβ β the site stays fully live.
# π Quick Start
You need a reachable turbolab server (chat model for summaries, e5 model for embeddings).
git clone https://github.com/usr-wwelsh/research-digest.git
cd research-digest
python3 -m venv venv && ./venv/bin/pip install -r requirements.txt
# point at your turbolab server (kept out of git)
echo 'export TURBOLAB_URL=http://YOUR_HOST:7860' > .env
./run.sh # fetch -> summarize -> embed -> relate -> render
# or, no network:
./run.sh --offline # rebuild the site from the existing corpus
Open index.html (landing), latest.html (digest), archive.html, or feed.html (mobile).
Migrating from v1? Recover your old backlog from the archived HTML with zero arXiv calls:
./venv/bin/python migrate_from_html.py # parses arxiv_archive/*.html into digest.db
./run.sh --offline # summarize + embed + render the backlog
(The original abstracts arenβt in old HTML, so salvaged papers are flagged
needs_abstract_backfill; ./venv/bin/python fetch.py --backfill refetches them in batches.)
# βοΈ Configuration
{
"interests": {
"Efficient ML / Edge AI": {
"query": "cat:cs.LG OR cat:cs.CV OR cat:cs.CL",
"keywords": ["efficient", "edge", "quantization", "distillation"]
}
},
"settings": {
"papers_per_interest": 25,
"recent_days": 7,
"fetch_multiplier": 3
},
"turbolab": {
"url": "http://localhost:7860",
"passage_prefix": "passage: ",
"query_prefix": "query: "
}
}
The turbolab URL is normally set per-deployment via the gitignored .env (TURBOLAB_URL),
which overrides config.json β so a private LAN address never lands in the repo.
| Setting | Default | Description |
|---|---|---|
papers_per_interest |
25 | Papers kept per interest per fetch |
recent_days |
7 | Look-back window (0 = all time) |
fetch_multiplier |
3 | Over-fetch, then keyword-rank, then trim |
arXiv query syntax: combine category codes with OR/AND, e.g. cat:cs.LG OR cat:cs.AI
(full taxonomy).
# π§ Self-hosted deployment (Proxmox LXC)
From your Proxmox host:
bash <(curl -sL https://raw.githubusercontent.com/usr-wwelsh/Research-Digest/main/create-lxc.sh)
This creates a Debian LXC, installs Caddy + cloudflared, sets up the venv, configures the
weekly cron (Monday 8am), and serves on :8080. After it finishes, set TURBOLAB_URL in
/opt/research-digest/.env, edit config.json, and run sudo -u www-data /opt/research-digest/run.sh.
Idle footprint is ~50β80MB (Caddy + cloudflared) β the heavy ML lives in turbolab on another host, so the digest container stays tiny.
# π Project structure
research-digest/
βββ config.json # interests, settings, turbolab block
βββ db.py # SQLite schema + data access (source of truth)
βββ fetch.py # arXiv ingest (backoff, upsert, --backfill)
βββ summarize.py # turbolab summaries
βββ embed.py # turbolab e5 embeddings
βββ relate.py # nearest-neighbour "related papers"
βββ render.py # static site builder (atomic, refuses empty)
βββ migrate_from_html.py # one-time v1 backlog salvage
βββ turbolab.py # turbolab client (chat + embeddings)
βββ templates/ # Jinja2 templates (autoescaped)
βββ run.sh # pipeline runner (cron entrypoint)
βββ setup.sh # LXC bootstrap
βββ create-lxc.sh # Proxmox LXC creator
βββ Caddyfile # static file server
βββ digest.db # the corpus (gitignored)
# π οΈ Requirements
- Python 3.9+ Β· deps:
requests,jinja2,numpy(no torch β the ML is in turbolab) - A reachable turbolab server
- Internet only for the
fetchstage
# π License
MIT β see LICENSE.
# π Acknowledgments
Built for researchers who want to stay current β on their own hardware, with no gatekeepers.