- Explore MCP Servers
- DuckDB-RAG-MCP-Sample
Duckdb Rag Mcp Sample
What is Duckdb Rag Mcp Sample
DuckDB-RAG-MCP-Sample is a sample project that enables the extraction and vectorization of text from markdown documents, allowing for retrieval-augmented generation (RAG) using the MCP framework. It utilizes the Plamo-Embedding-1B for vectorization.
Use cases
Use cases include generating insights from markdown documentation, enhancing search functionalities in applications, and integrating with AI systems for improved data retrieval.
How to use
To use DuckDB-RAG-MCP-Sample, place the markdown files in a specified directory and run the command to convert them into a Parquet file. Then, configure the MCP server and client settings according to the desired client, specifying the path to the generated Parquet file.
Key features
Key features include text extraction and vectorization from markdown files, vector search using DuckDB, persistence of vector data via Parquet files, and vector search capabilities from MCP.
Where to use
DuckDB-RAG-MCP-Sample can be used in various fields such as data analysis, natural language processing, and any application requiring efficient document retrieval and processing.
Clients Supporting MCP
The following are the main client software that supports the Model Context Protocol. Click the link to visit the official website for more information.
Overview
What is Duckdb Rag Mcp Sample
DuckDB-RAG-MCP-Sample is a sample project that enables the extraction and vectorization of text from markdown documents, allowing for retrieval-augmented generation (RAG) using the MCP framework. It utilizes the Plamo-Embedding-1B for vectorization.
Use cases
Use cases include generating insights from markdown documentation, enhancing search functionalities in applications, and integrating with AI systems for improved data retrieval.
How to use
To use DuckDB-RAG-MCP-Sample, place the markdown files in a specified directory and run the command to convert them into a Parquet file. Then, configure the MCP server and client settings according to the desired client, specifying the path to the generated Parquet file.
Key features
Key features include text extraction and vectorization from markdown files, vector search using DuckDB, persistence of vector data via Parquet files, and vector search capabilities from MCP.
Where to use
DuckDB-RAG-MCP-Sample can be used in various fields such as data analysis, natural language processing, and any application requiring efficient document retrieval and processing.
Clients Supporting MCP
The following are the main client software that supports the Model Context Protocol. Click the link to visit the official website for more information.
Content
DuckDB RAG MCP Sample
markdown ドキュメントを埋め込みベクトル化して、MCP から RAG で解説できるようにするサンプルです。
ベクトル化には Plamo-Embedding-1B を使用しています。
機能
- markdown ファイルからテキスト抽出・ベクトル化
- DuckDB を使用したベクトル検索
- Parquet ファイルによるベクトルデータの永続化
- MCP からベクトル検索
使用方法
ベクトルデータ生成
最初に検索対象にしたい markdown ファイルを特定のディレクトリに配置し、以下のコマンドで Parquet ファイルに変換してください。
uv run main.py --directory ~/path/to/markdown/files --parquet vectors.parquet
MCP の設定
ビルド
以下のコマンドでシングルバイナリが dist/server として生成されます。
uv run pyinstaller --clean --strip --noconfirm --onefile server.py
MCP のクライアント設定
利用したいクライアントに応じて設定してください。
Claude Desktop の場合は以下のような感じです。
VECTOR_PARQUET は先ほど変換したファイルを指定してください。
uv run mcp install server.py -v VECTOR_PARQUET=/path/to/vectors.parquet
以下のように設定されます。
{ "mcpServers": { "DuckDB-RAG-MCP-Sample": { "command": "/path/to/dist/server", "env": { "VECTOR_PARQUET": "/path/to/vectors.parquet" } } } }
開発用サーバー起動
uv run mcp dev server.py
ライセンス
DuckDB RAG MCP Sampleは、Apache License, Version 2.0の下で提供されています。
Dev Tools Supporting MCP
The following are the main code editors that support the Model Context Protocol. Click the link to visit the official website for more information.










