Local Document Search & Summarization
A local-first MCP server for document management and semantic search, with a web dashboard and REST API for AI agents.
Why it matters
Empower your local document management with intelligent semantic search and AI-powered summarization. Quickly find and understand information within your documents without relying on external cloud services.
Outcomes
What it gets done
Ingest and index local documents for semantic search.
Perform AI-driven searches for contextual understanding.
Extract key information and generate summaries.
Manage document versions and metadata locally.
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/vb-mcp-documentation-server | bash Capabilities
Tools your agent gets
Add a document with title, content, and metadata to the local store.
Show a list of saved documents and their metadata.
Get a full document by id from the local store.
Delete a document, its chunks, and associated source files.
Convert files in uploads folder to documents with chunking, embeddings, and backup.
Return the absolute path to the uploads folder.
Show files in the uploads folder.
AI-powered search using Gemini for advanced document analysis and contextual understanding.
Overview
MCP Documentation Server
A local-first MCP server for document management and semantic search over uploaded text, Markdown, and PDF files, pairing MCP tools with a REST API and web dashboard. Use when an AI coding agent needs a private, offline knowledge base it can search semantically without loading MCP schemas into context.
What it does
MCP Documentation Server is a local-first Model Context Protocol server for document management and semantic search, letting AI coding agents store, search, and retrieve knowledge without external databases, cloud APIs, or vendor lock-in. It ships with a full web dashboard for browsing, searching, uploading, and managing documents from a browser, and exposes every MCP tool as a REST API as well.
When to use - and when NOT to
Use this when an AI coding agent such as Claude Code, OpenCode, Gemini CLI, or Cursor needs a local knowledge base it can search semantically - adding documents, processing uploaded .txt/.md/.pdf files, and running hybrid full-text plus vector search - entirely offline. The REST API is the recommended integration path for agents specifically because it avoids loading MCP tool schemas into the conversation context; only the response JSON enters. A ready-made Agent Skill at skills/documentation-server/SKILL.md teaches an agent every REST endpoint with examples, installable via npx skills add.
Capabilities
Twelve tools across three groups. Document management: add_document, list_documents, get_document, delete_document. File processing: process_uploads (chunk and embed everything in the uploads folder), get_uploads_path, list_uploads_files, get_ui_url. Search: search_documents (vector search within one document), search_all_documents (hybrid full-text plus vector search across all documents), get_context_window (neighboring chunks for broader LLM context), and search_documents_with_ai (Gemini-powered analysis, requires GEMINI_API_KEY). Search uses parent-child chunking: documents are split into large parent chunks that preserve context, then into small child chunks for precise vector matching, with results deduplicated by parent at query time. A web dashboard starts automatically on port 3080 alongside the MCP server, offering a document overview, browsing, an add-document form, search-all and search-in-document views, AI search, drag-and-drop file uploads, and a context-window explorer. Local embeddings run through Transformers.js with an LRU cache; the default model is Xenova/all-MiniLM-L6-v2 (384 dimensions), with Xenova/paraphrase-multilingual-mpnet-base-v2 (768 dimensions) recommended for best multilingual quality - changing models requires re-adding all documents since embeddings from different models are incompatible.
How to install
Published on the MCP Registry, so no clone is needed - install via npx:
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"]
}
}
}
The web UI opens automatically at http://localhost:3080. All further configuration is via optional environment variables: MCP_BASE_DIR for the data directory (default ~/.mcp-documentation-server), MCP_EMBEDDING_MODEL, GEMINI_API_KEY to enable AI search, START_WEB_UI to disable the dashboard, and WEB_HOST/WEB_PORT to change its binding. Without GEMINI_API_KEY, only the local embedding-based search tools are available.
Who it's for
Developers using AI coding agents who want a private, offline knowledge base - project docs, specs, references - searchable both by the agent via MCP or REST and by a human through the built-in dashboard, with no data leaving the machine unless Gemini AI search is explicitly enabled, released under the MIT license.
Source README
MCP Documentation Server
Local-first document management and semantic search for AI coding agents. No external databases, no cloud APIs, no vendor lock-in.
Unlike other MCP servers that are CLI-only, this one ships with a full web dashboard - browse, search, upload, and manage your knowledge base from your browser. Every MCP tool is also exposed as a REST API, giving AI agents a lean, schema-free interface.
- ๐ Runs fully offline - Orama vector DB with local AI embeddings (Transformers.js)
- ๐ Built-in Web UI - starts automatically on port 3080 alongside the MCP server
- ๐ Hybrid search - full-text + vector similarity with parent-child chunking
- ๐ค Optional AI search - Google Gemini for advanced document analysis (bring your own key)
- ๐ Drag & drop uploads -
.txt,.md,.pdfsupport - ๐ฆ Published on the MCP Registry - installable via npx, no clone needed
Quick Start
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"]
}
}
}
Open your browser at http://localhost:3080 - the web UI starts automatically.
๐ค Agent Skill (REST API) - recommended for AI agents
Every MCP tool is also accessible via the REST API on http://127.0.0.1:3080/api/. This is the recommended way to interact from AI agents (Claude Code, OpenCode, Gemini CLI, Cursor) because it avoids loading MCP tool schemas into the conversation context - only the response JSON enters.
curl -s http://127.0.0.1:3080/api/config
curl -s http://127.0.0.1:3080/api/documents
curl -s -X POST http://127.0.0.1:3080/api/search-all \
-H "Content-Type: application/json" \
-d '{"query": "your search", "limit": 5}'
A ready-to-use skill is included at skills/documentation-server/SKILL.md - it teaches your agent every endpoint with examples. Install it:
npx skills add https://github.com/andrea9293/mcp-documentation-server --skill documentation-server
Basic workflow
- Add documents using
add_documentor place.txt/.md/.pdffiles in the uploads folder and callprocess_uploads. - Search across everything with
search_all_documents, or within a single document withsearch_documents. - Use
get_context_windowto fetch neighboring chunks and give the LLM broader context.
Web UI
The web interface starts automatically on port 3080 when the MCP server launches. From the web UI you can:
- ๐ Dashboard - overview of all documents and stats
- ๐ Documents - browse, view, and delete documents
- โ Add Document - create documents with title, content, and metadata
- ๐ Search All - semantic search across all documents
- ๐ฏ Search in Doc - search within a specific document
- ๐ค AI Search - Gemini-powered analysis (if
GEMINI_API_KEYis set) - ๐ Upload Files - drag & drop files and process them into the knowledge base
- ๐ช Context Window - explore chunks around a specific index
Configure an MCP client
Minimal
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"]
}
}
}
With environment variables (all optional)
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"],
"env": {
"MCP_BASE_DIR": "/path/to/workspace",
"GEMINI_API_KEY": "your-api-key-here",
"MCP_EMBEDDING_MODEL": "Xenova/all-MiniLM-L6-v2",
"START_WEB_UI": "true",
"WEB_HOST": "127.0.0.1",
"WEB_PORT": "3080"
}
}
}
}
All environment variables are optional. Without GEMINI_API_KEY, only the local embedding-based search tools are available.
MCP Tools
The server registers the following tools (all validated with Zod schemas):
๐ Document Management
| Tool | Description |
|---|---|
add_document |
Add a document (title, content, optional metadata) |
list_documents |
List all documents with metadata and content preview |
get_document |
Retrieve the full content of a document by ID |
delete_document |
Remove a document, its chunks, database entries, and associated files |
๐ File Processing
| Tool | Description |
|---|---|
process_uploads |
Process all files in the uploads folder (chunking + embeddings) |
get_uploads_path |
Returns the absolute path to the uploads folder |
list_uploads_files |
Lists files in the uploads folder with size and format info |
get_ui_url |
Returns the Web UI URL (e.g. http://localhost:3080) - useful to open the dashboard or to locate the uploads folder from the browser |
๐ Search
| Tool | Description |
|---|---|
search_documents |
Semantic vector search within a specific document |
search_all_documents |
Hybrid (full-text + vector) cross-document search |
get_context_window |
Returns a window of chunks around a given chunk index |
search_documents_with_ai |
๐ค AI-powered search using Gemini (requires GEMINI_API_KEY) |
Configuration
Configure via environment variables or a .env file in the project root:
| Variable | Default | Description |
|---|---|---|
MCP_BASE_DIR |
~/.mcp-documentation-server |
Base directory for data storage |
MCP_EMBEDDING_MODEL |
Xenova/all-MiniLM-L6-v2 |
Embedding model name |
GEMINI_API_KEY |
- | Google Gemini API key (enables search_documents_with_ai) |
MCP_CACHE_ENABLED |
true |
Enable/disable LRU embedding cache |
START_WEB_UI |
true |
Set to false to disable the built-in web interface |
WEB_HOST |
127.0.0.1 |
Bind address for the web UI (use 0.0.0.0 to expose on all interfaces) |
WEB_PORT |
3080 |
Port for the web UI |
MCP_STREAMING_ENABLED |
true |
Enable streaming reads for large files |
MCP_STREAM_CHUNK_SIZE |
65536 |
Streaming buffer size in bytes (64KB) |
MCP_STREAM_FILE_SIZE_LIMIT |
10485760 |
Threshold to switch to streaming (10MB) |
Storage layout
~/.mcp-documentation-server/ # Or custom path via MCP_BASE_DIR
โโโ data/
โ โโโ orama-chunks.msp # Orama vector DB (child chunks + embeddings)
โ โโโ orama-docs.msp # Orama document DB (full content + metadata)
โ โโโ orama-parents.msp # Orama parent chunks DB (context sections)
โ โโโ migration-complete.flag # Written after legacy JSON migration
โ โโโ *.md # Markdown copies of documents
โโโ uploads/ # Drop .txt, .md, .pdf files here
Embedding Models
Set via MCP_EMBEDDING_MODEL:
| Model | Dimensions | Notes |
|---|---|---|
Xenova/all-MiniLM-L6-v2 |
384 | Default - fast, good quality |
Xenova/paraphrase-multilingual-mpnet-base-v2 |
768 | Recommended - best quality, multilingual |
Models are downloaded on first use (~80-420 MB). The vector dimension is determined automatically from the provider.
โ ๏ธ Important: Changing the embedding model requires re-adding all documents - embeddings from different models are incompatible. The Orama database is recreated automatically when the dimension changes.
Architecture
Server (FastMCP, stdio)
โโ Web UI (Express, port 3080)
โ โโ REST API โ DocumentManager
โโ MCP Tools
โโ DocumentManager
โโ OramaStore โ Orama vector DB (chunks DB + docs DB + parents DB), persistence, migration
โโ IntelligentChunker โ Parent-child chunking (code, markdown, text, PDF)
โโ EmbeddingProvider โ Local embeddings via @xenova/transformers
โ โโ EmbeddingCache โ LRU in-memory cache
โโ GeminiSearchService โ Optional AI search via Google Gemini
- OramaStore manages three Orama instances: one for document metadata/content, one for child chunks with vector embeddings, and one for parent chunks (context sections). All are persisted to binary files on disk and restored on startup.
- IntelligentChunker implements the Parent-Child Chunking pattern: documents are first split into large parent chunks that preserve full context (sections, paragraphs), then each parent is further split into small child chunks for precise vector search. At query time, results are deduplicated by parent so that the LLM receives both the matched fragment and the broader context.
- EmbeddingProvider lazily loads a Transformers.js model for local inference - no API calls needed.
Development
git clone https://github.com/andrea9293/mcp-documentation-server.git
cd mcp-documentation-server
npm install
npm run dev # FastMCP dev mode with hot reload
npm run build # TypeScript compilation
npm run inspect # FastMCP web UI for interactive tool testing
npm start # Direct tsx execution (MCP server + web UI)
npm run web # Run only the web UI (development)
npm run web:build # Run only the web UI (compiled)
Star History
FAQ
Common questions
Discussion
Questions & comments ยท 0
Sign In Sign in to leave a comment.

