MCP serverio.github.fasuizu-br/speech-ai
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Overview
Score?
UNRATED 0.615
of what a free look can see, on 29 looks
Looks
36
last 2 hr ago
Tools
10
More info
URL
apim-ai-apis.azure-api.net/mcp/pronunciation/mcp
streamable-http
Says it is
speech-ai 2.14.5
protocol 2025-06-18
In the record since
32 days ago
Among servers18,413 with a card
0median 0.606 · this server 0.615 · highest on record 0.8561
Toolsfrom sha256:bc41b31275…86aadd
| Tool | Schema |
|---|---|
| assess_pronunciation Assess English pronunciation quality from audio.
Scores pronunciation at four levels: overall, sentence, word, and phoneme.
Each score is 0-100. Phonemes are returned in both IPA |
input · no output |
| check_pronunciation_service Check if the pronunciation assessment service is healthy and ready.
Returns:
dict with keys:
- status (str): 'healthy' or error state
- modelLoaded (bool): Whe |
input · no output |
| check_stt_service Check if the speech-to-text service is healthy and ready.
Returns:
dict with keys:
- status (str): 'healthy' or error state
- modelLoaded (bool): Whether the S |
input · no output |
| check_tts_service Check if the text-to-speech service is healthy and ready.
Returns:
dict with keys:
- status (str): 'healthy' or error state
- modelLoaded (bool): Whether the T |
input · no output |
| check_whisper_service Check if the Whisper STT Pro service is healthy and ready.
Returns:
dict with keys:
- status (str): 'healthy' or error state
- modelLoaded (bool): Whether the |
input · no output |
| get_phoneme_inventory Get the full phoneme inventory supported by the pronunciation scorer.
Returns a list of all English phonemes the engine can assess, including
ARPAbet symbol, IPA equivalent, examp |
input · no output |
| list_tts_voices List all available text-to-speech voices with metadata.
Returns:
dict with keys:
- voices (list): Available voices, each with id, name, gender, accent, grade
- |
input · no output |
| synthesize_speech Generate natural speech audio from English text.
Produces high-quality speech with 12 English voices.
Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata |
input · no output |
| transcribe_audio Transcribe audio to text with word-level timestamps.
Converts spoken English audio into text with optional word-level timestamps
and per-word confidence scores.
Args:
audio_b |
input · no output |
| transcribe_audio_pro Transcribe audio with Whisper Large V3 Turbo — multilingual STT.
Supports 99 languages with automatic language detection, word-level
timestamps, per-word confidence scores, and op |
input · no output |
Verify it yourself
npx teppi-check https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcpcurl -s https://api.teppi.xyz/v1/trust/mcp/mcs_01M1FZ2F9TE7BM8CKJ0591NWYN