# Botdocs
Convert markdown documentation into beautiful static sites with AI-powered semantic search — no backend required.
# Features
- Markdown to HTML - Converts
.mdfiles into polished static sites - Semantic Search - Client-side vector search using Transformers.js
- Dark Mode - Built-in theme switching
- Deep Links - Search results link directly to sections
- No Backend - Everything runs in the browser
- Fast - Syntax highlighting with Shiki
- Live Preview -
--watchrebuilds on save and serves the site locally - SEO - Optional Open Graph/Twitter tags,
sitemap.xmlandrobots.txtviabaseUrl - Agent Friendly - Raw
.mdsource published next to every page, plus anllms.txtindex
# Installation
Install globally via npm:
npm install -g botdocs
# Usage
# Generate site from markdown
botdocs ./docs
# Disable chatbot
botdocs ./docs --no-chat
# Custom output directory
botdocs ./docs -o ./public
# Verbose logging
botdocs ./docs -v
# Use a specific theme
botdocs ./docs -t material
# Custom config file
botdocs ./docs -c ./my-config.json
# Combine multiple options
botdocs ./docs -o ./public -t slate -v
# Live preview: rebuild on save + local server
botdocs ./docs --watch
# CLI Options
| Option | Alias | Description | Default |
|---|---|---|---|
--output <dir> |
-o |
Output directory for generated site | output |
--no-chat |
Disable AI chatbot functionality | false |
|
--config <file> |
-c |
Path to config file | botdocs.config.json |
--theme <theme> |
-t |
Theme to use | classic |
--verbose |
-v |
Enable verbose logging | false |
--watch |
-w |
Rebuild on changes and serve the site for live preview | false |
--port <number> |
-p |
Port for the preview server (with --watch) |
3000 |
# Available Themes
- classic - Clean, professional theme (default)
- material - Material Design theme
- minimal - Clean, minimalist theme
- slate - Dark slate theme
- modern - Modern documentation theme
# Configuration
Create botdocs.config.json in your docs directory:
{
"title": "My Documentation",
"description": "Project docs",
"theme": "classic",
"customCss": "custom.css",
"attribution": true,
"baseUrl": "https://example.com/docs/",
"chat": { "enabled": true },
"build": {
"chunkSize": 500,
"chunkOverlap": 50,
"minChunkSize": 15,
"topK": 3,
"minScore": 0.75
}
}
# Configuration Options
| Option | Type | Default | Description |
|---|---|---|---|
title |
string | "Documentation" |
Site title |
description |
string | "Project documentation" |
Site description |
theme |
string | "classic" |
Theme to use (classic, material, minimal, slate, modern) |
customCss |
string | none | Path to a CSS file, resolved relative to the config file’s directory. Appended after theme CSS in bundle.css, so same-specificity selectors override the theme without !important |
attribution |
boolean | true |
Show “Built with Botdocs” footer link |
baseUrl |
string | none | Canonical URL where the site is hosted. When set, pages get rel=canonical and Open Graph/Twitter card tags, a sitemap.xml and robots.txt are generated, and llms.txt links become absolute |
chat.enabled |
boolean | true |
Enable AI chatbot |
chat.welcomeMessage |
string | "Ask me anything about the docs!" |
Chatbot welcome message |
build.chunkSize |
number | 500 |
Text chunk size for embeddings |
build.chunkOverlap |
number | 50 |
Overlap between chunks |
build.minChunkSize |
number | 15 |
Chunks smaller than this (estimated tokens) get folded into a neighboring chunk instead of becoming a standalone, low-signal search result |
build.topK |
number | 3 |
Number of results to return |
build.minScore |
number | 0.75 |
Minimum vector similarity (0-1) a result must reach to be returned at all, regardless of topK — filters out weak/off-topic matches instead of always padding results. The e5 embedding model has a fairly high similarity floor even for unrelated text, so this needs to sit well above 0.5 to actually gate anything |
# Front Matter
---
title: Getting Started
description: Quick start guide
---
# Your content here
# How It Works
- Build: Parses markdown → generates embeddings → creates
vector-db.json - Runtime: User query → embed → search vector DB → return relevant chunks
- No LLM: Pure semantic search, not AI text generation
- Consent: On first use, visitors are asked before the embedding model downloads to their browser, with a disclosure of what runs locally
# Architecture
- Embedding Model:
e5-small-v2(384-dim vectors, 2.2x faster than all-MiniLM-L6-v2) - Search: Hybrid — vector cosine similarity fused with BM25 keyword scoring (Reciprocal Rank Fusion), gated by a minimum similarity threshold, client-side only
- Browser Bundle: ~825KB (includes Transformers.js)
- Deployment: Fully static, works on any host
# Development
Building from source:
git clone https://github.com/usr-wwelsh/botdocs.git
cd botdocs
npm install
npm run build && npm run build:client
botdocs ./test-docs
# Retrieval eval
npm run eval scores search against a golden set of queries (eval/golden.json) over a frozen corpus snapshot (eval/corpus/). It compares five retrievers: grep, BM25, dense, hybrid (RRF), and the shipped hybrid with its relevance gates. It reports recall@5 and MRR on answerable queries, plus a false-positive rate on queries that should return nothing.
Each golden entry has an id, a query, a split (tune or holdout), and expect: labels like turbolab/README.md#memory (one section) or omniMux/README.md (any section of the page). An empty expect means nothing should match. The id prefix sets the query’s kind (identifier-, paraphrase-, buried-, cross-, negative-), and answerable queries are also scored per kind. Tune thresholds against tune only; holdout keeps the numbers honest.
Embeddings are cached in eval/.cache/, keyed by corpus, chunker, embedder, and build config, so only the first run embeds the corpus.
# License
MIT © usr-wwelsh