# Botdocs

View on GitHub

Read about the commits

npm version

Convert markdown documentation into beautiful static sites with AI-powered semantic search — no backend required.

# Features

  • Markdown to HTML - Converts .md files into polished static sites
  • Semantic Search - Client-side vector search using Transformers.js
  • Dark Mode - Built-in theme switching
  • Deep Links - Search results link directly to sections
  • No Backend - Everything runs in the browser
  • Fast - Syntax highlighting with Shiki
  • Live Preview - --watch rebuilds on save and serves the site locally
  • SEO - Optional Open Graph/Twitter tags, sitemap.xml and robots.txt via baseUrl
  • Agent Friendly - Raw .md source published next to every page, plus an llms.txt index

# Installation

Install globally via npm:

npm install -g botdocs

# Usage

# Generate site from markdown
botdocs ./docs

# Disable chatbot
botdocs ./docs --no-chat

# Custom output directory
botdocs ./docs -o ./public

# Verbose logging
botdocs ./docs -v

# Use a specific theme
botdocs ./docs -t material

# Custom config file
botdocs ./docs -c ./my-config.json

# Combine multiple options
botdocs ./docs -o ./public -t slate -v

# Live preview: rebuild on save + local server
botdocs ./docs --watch

# CLI Options

Option Alias Description Default
--output <dir> -o Output directory for generated site output
--no-chat Disable AI chatbot functionality false
--config <file> -c Path to config file botdocs.config.json
--theme <theme> -t Theme to use classic
--verbose -v Enable verbose logging false
--watch -w Rebuild on changes and serve the site for live preview false
--port <number> -p Port for the preview server (with --watch) 3000

# Available Themes

  • classic - Clean, professional theme (default)
  • material - Material Design theme
  • minimal - Clean, minimalist theme
  • slate - Dark slate theme
  • modern - Modern documentation theme

# Configuration

Create botdocs.config.json in your docs directory:

{
  "title": "My Documentation",
  "description": "Project docs",
  "theme": "classic",
  "customCss": "custom.css",
  "attribution": true,
  "baseUrl": "https://example.com/docs/",
  "chat": { "enabled": true },
  "build": {
    "chunkSize": 500,
    "chunkOverlap": 50,
    "minChunkSize": 15,
    "topK": 3,
    "minScore": 0.75
  }
}

# Configuration Options

Option Type Default Description
title string "Documentation" Site title
description string "Project documentation" Site description
theme string "classic" Theme to use (classic, material, minimal, slate, modern)
customCss string none Path to a CSS file, resolved relative to the config file’s directory. Appended after theme CSS in bundle.css, so same-specificity selectors override the theme without !important
attribution boolean true Show “Built with Botdocs” footer link
baseUrl string none Canonical URL where the site is hosted. When set, pages get rel=canonical and Open Graph/Twitter card tags, a sitemap.xml and robots.txt are generated, and llms.txt links become absolute
chat.enabled boolean true Enable AI chatbot
chat.welcomeMessage string "Ask me anything about the docs!" Chatbot welcome message
build.chunkSize number 500 Text chunk size for embeddings
build.chunkOverlap number 50 Overlap between chunks
build.minChunkSize number 15 Chunks smaller than this (estimated tokens) get folded into a neighboring chunk instead of becoming a standalone, low-signal search result
build.topK number 3 Number of results to return
build.minScore number 0.75 Minimum vector similarity (0-1) a result must reach to be returned at all, regardless of topK — filters out weak/off-topic matches instead of always padding results. The e5 embedding model has a fairly high similarity floor even for unrelated text, so this needs to sit well above 0.5 to actually gate anything

# Front Matter

---
title: Getting Started
description: Quick start guide
---

# Your content here

# How It Works

  1. Build: Parses markdown → generates embeddings → creates vector-db.json
  2. Runtime: User query → embed → search vector DB → return relevant chunks
  3. No LLM: Pure semantic search, not AI text generation
  4. Consent: On first use, visitors are asked before the embedding model downloads to their browser, with a disclosure of what runs locally

# Architecture

  • Embedding Model: e5-small-v2 (384-dim vectors, 2.2x faster than all-MiniLM-L6-v2)
  • Search: Hybrid — vector cosine similarity fused with BM25 keyword scoring (Reciprocal Rank Fusion), gated by a minimum similarity threshold, client-side only
  • Browser Bundle: ~825KB (includes Transformers.js)
  • Deployment: Fully static, works on any host

# Development

Building from source:

git clone https://github.com/usr-wwelsh/botdocs.git
cd botdocs
npm install
npm run build && npm run build:client
botdocs ./test-docs

# Retrieval eval

npm run eval scores search against a golden set of queries (eval/golden.json) over a frozen corpus snapshot (eval/corpus/). It compares five retrievers: grep, BM25, dense, hybrid (RRF), and the shipped hybrid with its relevance gates. It reports recall@5 and MRR on answerable queries, plus a false-positive rate on queries that should return nothing.

Each golden entry has an id, a query, a split (tune or holdout), and expect: labels like turbolab/README.md#memory (one section) or omniMux/README.md (any section of the page). An empty expect means nothing should match. The id prefix sets the query’s kind (identifier-, paraphrase-, buried-, cross-, negative-), and answerable queries are also scored per kind. Tune thresholds against tune only; holdout keeps the numbers honest.

Embeddings are cached in eval/.cache/, keyed by corpus, chunker, embedder, and build config, so only the first run embeds the corpus.

# License

MIT © usr-wwelsh

🔒 This site's search assistant runs a local AI model in your browser — no cloud AI, no tracking. See the ℹ️ in the chat panel for details.