Fast Whisper Mcp Server

6 MIT

FreeCommunity

AI Systems

# A high-performance speech recognition MCP server based on Faster Whisper, providing efficient audio transcription capabilities.

What is Fast Whisper Mcp Server

Fast-Whisper-MCP-Server is a high-performance speech recognition server based on Faster Whisper, designed to provide efficient audio transcription capabilities.

Use cases

Use cases include transcribing podcasts, generating subtitles for videos, converting audio notes into text, and processing large volumes of audio files efficiently.

How to use

To use Fast-Whisper-MCP-Server, clone the repository, set up a virtual environment, install dependencies, and start the server using ‘python whisper_server.py’ or ‘start_server.bat’ on Windows.

Key features

Key features include integration with Faster Whisper for efficient speech recognition, batch processing acceleration, automatic CUDA acceleration, support for multiple model sizes, various output formats (VTT, SRT, JSON), and model instance caching.

Where to use

Fast-Whisper-MCP-Server can be used in fields such as transcription services, content creation, accessibility tools, and any application requiring audio-to-text conversion.

Clients Supporting MCP

The following are the main client software that supports the Model Context Protocol. Click the link to visit the official website for more information.

Claude Desktop: Official desktop application from Anthropic, natively supports MCP protocol. claude.ai

Cherry Studio: Cross-platform desktop client supporting multiple LLM providers, built-in MCP server support. cherry-ai.com

LobeChat: Modern open-source ChatGPT/LLMs UI, supports MCP protocol integration. lobehub.com

DeepChat: Cross-platform desktop AI assistant, compatible with MCP protocol, focusing on privacy and efficiency. deepchat.thinkinai.xyz

5ire: Cross-platform open-source desktop intelligent assistant MCP client, supports local knowledge base and MCP server. 5ire.app

View More MCP Clients

Overview

What is Fast Whisper Mcp Server

Fast-Whisper-MCP-Server is a high-performance speech recognition server based on Faster Whisper, designed to provide efficient audio transcription capabilities.

Use cases

Use cases include transcribing podcasts, generating subtitles for videos, converting audio notes into text, and processing large volumes of audio files efficiently.

How to use

Key features

Where to use

Fast-Whisper-MCP-Server can be used in fields such as transcription services, content creation, accessibility tools, and any application requiring audio-to-text conversion.

Clients Supporting MCP

The following are the main client software that supports the Model Context Protocol. Click the link to visit the official website for more information.

Claude Desktop: Official desktop application from Anthropic, natively supports MCP protocol. claude.ai

Cherry Studio: Cross-platform desktop client supporting multiple LLM providers, built-in MCP server support. cherry-ai.com

LobeChat: Modern open-source ChatGPT/LLMs UI, supports MCP protocol integration. lobehub.com

DeepChat: Cross-platform desktop AI assistant, compatible with MCP protocol, focusing on privacy and efficiency. deepchat.thinkinai.xyz

5ire: Cross-platform open-source desktop intelligent assistant MCP client, supports local knowledge base and MCP server. 5ire.app

View More MCP Clients

Content

Whisper Speech Recognition MCP Server

中文文档

A high-performance speech recognition MCP server based on Faster Whisper, providing efficient audio transcription capabilities.

Features

Integrated with Faster Whisper for efficient speech recognition
Batch processing acceleration for improved transcription speed
Automatic CUDA acceleration (if available)
Support for multiple model sizes (tiny to large-v3)
Output formats include VTT subtitles, SRT, and JSON
Support for batch transcription of audio files in a folder
Model instance caching to avoid repeated loading
Dynamic batch size adjustment based on GPU memory

Installation

Dependencies

Python 3.10+
faster-whisper>=0.9.0
torch==2.6.0+cu126
torchaudio==2.6.0+cu126
mcp[cli]>=1.2.0

Installation Steps

Clone or download this repository
Create and activate a virtual environment (recommended)
Install dependencies:

pip install -r requirements.txt

PyTorch Installation Guide

Install the appropriate version of PyTorch based on your CUDA version:

CUDA 12.6:

pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu126

CUDA 12.1:

pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 --index-url https://download.pytorch.org/whl/cu121

CPU version:

pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cpu

You can check your CUDA version with nvcc --version or nvidia-smi.

Usage

Starting the Server

On Windows, simply run start_server.bat.

On other platforms, run:

python whisper_server.py

Configuring Claude Desktop

Open the Claude Desktop configuration file:
- Windows: %APPDATA%\Claude\claude_desktop_config.json
- macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
Add the Whisper server configuration:

{
  "mcpServers": {
    "whisper": {
      "command": "python",
      "args": [
        "D:/path/to/whisper_server.py"
      ],
      "env": {}
    }
  }
}

Restart Claude Desktop

Available Tools

The server provides the following tools:

get_model_info - Get information about available Whisper models
transcribe - Transcribe a single audio file
batch_transcribe - Batch transcribe audio files in a folder

Performance Optimization Tips

Using CUDA acceleration significantly improves transcription speed
Batch processing mode is more efficient for large numbers of short audio files
Batch size is automatically adjusted based on GPU memory size
Using VAD (Voice Activity Detection) filtering improves accuracy for long audio
Specifying the correct language can improve transcription quality

Local Testing Methods

Use MCP Inspector for quick testing:

mcp dev whisper_server.py

Use Claude Desktop for integration testing
Use command line direct invocation (requires mcp[cli]):

mcp run whisper_server.py

Error Handling

The server implements the following error handling mechanisms:

Audio file existence check
Model loading failure handling
Transcription process exception catching
GPU memory management
Batch processing parameter adaptive adjustment

Project Structure

whisper_server.py: Main server code
model_manager.py: Whisper model loading and caching
audio_processor.py: Audio file validation and preprocessing
formatters.py: Output formatting (VTT, SRT, JSON)
transcriber.py: Core transcription logic
start_server.bat: Windows startup script

License

MIT

Acknowledgements

This project was developed with the assistance of these amazing AI tools and models:

GitHub Copilot - AI pair programmer
Trae - Agentic AI coding assistant
Cline - AI-powered terminal
DeepSeek - Advanced AI model
Claude-3.7-Sonnet - Anthropic’s powerful AI assistant
Gemini-2.0-Flash - Google’s multimodal AI model
VS Code - Powerful code editor
Whisper - OpenAI’s speech recognition model
Faster Whisper - Optimized Whisper implementation

Special thanks to these incredible tools and the teams behind them.

Dev Tools Supporting MCP

The following are the main code editors that support the Model Context Protocol. Click the link to visit the official website for more information.

Zed: High-performance collaborative code editor, supports MCP protocol, providing a smooth programming experience. zed.dev

Cursor: AI code editor built on VS Code, supports MCP protocol for context-aware programming. cursor.com

Windsurf: AI code editor from Codeium, integrates MCP protocol to provide intelligent code assistance. windsurf.com

Continue: Open-source AI programming assistant plugin, supports VS Code and JetBrains, compatible with MCP protocol. continue.dev

Trae: AI-driven code editor, supports MCP protocol, focusing on enhancing developer programming experience. trae.ai

View More MCP Dev Tools

Tools

No tools

Comments

Recommend MCP Servers

Tavily MCP Server The Tavily MCP server provides: search, extract, map, crawl tools Real-time web search capabilities through the tavily-search tool Intelligent data extraction from web pages via the tavily-extract tool Powerful web mapping tool that creates a structured map of website Web crawler that systematically explores websites.

MCP Server Chart This is a TypeScript-based MCP server that provides chart generation capabilities. It allows you to create various types of charts through MCP tools. You can also use it in Dify.

GitHub MCP Server MCP Server for the GitHub API, enabling file operations, repository management, search functionality, and more.

Brave Search MCP Server Web and local search using Brave's Search API

Firecrawl MCP Server Advanced web scraping with JavaScript rendering, PDF support, and smart rate limiting

Context7 MCP LLMs rely on outdated or generic information about the libraries you use. You get:

Slack MCP server Channel management and messaging capabilities

Sequential Thinking MCP Server Dynamic and reflective problem-solving through thought sequences

Fetch MCP Server A Model Context Protocol server that provides web content fetching capabilities.

Playwright MCP A Model Context Protocol (MCP) server that provides browser automation capabilities using [Playwright](https://playwright.dev). This server enables LLMs to interact with web pages through structured accessibility snapshots, bypassing the need for screenshots or visually-tuned models.

View All MCP Servers