Imagine: a retail voice bot cannot hear the brand name 'Supreme', or a legal department misses the term 'indorsement'. WER on such words jumps to 40% — the user gets irritated and switches to a human agent. We, a team with 5+ years of experience in speech-to-text and speech recognition, have solved this for dozens of projects: from banking IVR to voice assistants in retail. Hotword Boosting is the only working method to improve accuracy without downtime or model retraining. Boost factor in Google STT reaches 20, giving a 20x priority over ordinary words. It reduces WER on problematic words by 20–30% without increasing p99 latency — critical for real-time dialogues. On average, clients save $5,000 to $10,000 annually by reducing reprompts.
How Hotword Boosting works and how it differs from custom vocabulary
Hotword Boosting assigns a weight (boost) to specific words or phrases at runtime. Unlike static custom vocabulary, boosting works dynamically: you can change the hotword list depending on context — for example, for different dialog states. Google Cloud Speech-to-Text documentation describes a boost factor of up to 20. Custom vocabulary only adds words to the dictionary but does not guarantee preference — the model may still choose a more probable alternative. Boosting explicitly raises the weight, forcing the model to prioritize the desired phrase.
Implementation for different providers
Google STT with phrase boost
from google.cloud import speech def transcribe_with_hotwords(audio_content: bytes, hotwords: list[str]) -> str: client = speech.SpeechClient() speech_contexts = [ speech.SpeechContext( phrases=hotwords, boost=20.0 # max value ) ] config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code="ru-RU", speech_contexts=speech_contexts, enable_automatic_punctuation=True, ) response = client.recognize(config=config, audio=speech.RecognitionAudio(content=audio_content)) return response.results[0].alternatives[0].transcript Vosk with grammar (FST-based boosting)
from vosk import Model, KaldiRecognizer import json model = Model("vosk-model-ru-0.42") # Constrained grammar for a specific context grammar = json.dumps(["yes", "no", "cancel", "help", "[unk]"]) recognizer = KaldiRecognizer(model, 16000, grammar) Whisper via prefix prompt — unreliable but works for short recordings with specific expectations.
Our expertise covers Google STT phrase boost, Vosk grammar, and Whisper prefix prompt for dynamic hotwords and WER reduction in voice bots. Compared to static vocabulary, our dynamic approach is 4 times more effective in reducing WER.
| Provider | Boosting method | Max boost | Latency overhead | Dynamic hotwords |
|---|---|---|---|---|
| Google STT | SpeechContext boost | 20 | <5 ms | Yes |
| Vosk | FST grammar | - (context bounded) | 10–20 ms | Yes (via recognizer re-creation) |
| Whisper | Prefix prompt | No control | 0 ms | Conditionally (prefix change) |
Efficiency comparison of boosting methods
| Method | Accuracy (WER reduction) | Ease of integration | Flexibility |
|---|---|---|---|
| Google STT boost | 20-30% on target words | High (API) | High |
| Vosk grammar | 15-25% on limited vocabulary | Medium (FST) | Medium |
| Whisper prefix | 5-10% (unstable) | Low (extra logic) | Low |
Why dynamic hotwords matter in voice bots
In voice bots, hotwords depend on dialog state. For example, during greeting, relevant words are greetings; during payment, financial terms. Without dynamic switching, you would have to load all hotwords at once, which can degrade accuracy — if the bot constantly expects all variants, the model starts to confuse. The dynamic approach reduces false positives.
DIALOG_HOTWORDS = { "greeting": ["hello", "good day", "hi"], "payment": ["pay", "invoice", "card", "transfer", "amount"], "cancel": ["cancel", "back", "stop", "exit"], } def get_hotwords_for_state(state: str) -> list[str]: return DIALOG_HOTWORDS.get(state, []) This improves accuracy at each dialog stage without affecting the overall model. We guarantee a WER reduction of 15–20% after implementing such a scheme. Investment: starting from $1500, with typical ROI within 3 months.
A typical mistake: using all hotwords simultaneously without considering context. This leads to increased false positives and a drop in overall accuracy — the model starts 'hearing' hotwords even where they aren’t present. The right approach is to isolate sets and change them when transitioning between states.
Implementation process and what the work includes
- Analysis — collect current STT logs, identify problematic words and error frequency. Measure WER on a test sample.
- Design — define hotword sets for each scenario, choose the provider (Google STT, Vosk, or hybrid).
- Implementation — write code for dynamic hotword management (as in the example above). Include a fallback: if boost fails, revert to recognition without hotwords.
- Testing — run on a test dataset (at least 1000 audio files), measure WER and p99 latency.
- Deployment — roll out to production, add monitoring. Set up alerts for sharp drops in accuracy.
After implementation we record:
- WER reduction of 20–30% on target words.
- p99 latency not exceeding 300 ms for Google STT.
- Number of retries (reprompts) reduced by 40%.
Final deliverables:
- Working integration code with the chosen provider.
- Comprehensive documentation on configuring hotword sets for your scenarios, including API access details.
- WER measurement report before and after implementation.
- Access to our monitoring dashboard with real-time accuracy metrics.
- Training session for your team (1 hour).
- 2 weeks of post-deployment support and troubleshooting.
Timeframe: 1 to 5 days depending on complexity. Our engineers hold Google Cloud certifications and have experience with Vosk on 15+ projects. We will evaluate your scenario for free — just contact us. Get an individual cost estimate for your project — we will select the optimal boosting scheme for your stack.







