Voice Control for Websites: Speech API, Dictation, TTS

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Voice Control for Websites: Speech API, Dictation, TTS
Medium
~2-3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1362
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1253
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    958
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1190
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    932
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    949

Firefox and Safari block native SpeechRecognition — up to 30% of users lose voice control functionality. Average response time is 1.2 seconds. We solve this with a hybrid: browser API as primary for Chrome and Edge, Whisper from OpenAI as fallback for other browsers. This approach reduces response time to 1–4 seconds and covers 100% of browsers. Over 30 projects, we've gathered use cases: dictation in CRM, presentation control, voice search in e-commerce. Statistics show that more than 60% of mobile users prefer voice input over typing. Whisper achieves 95% accuracy even on noisy recordings.

What is the Web Speech API?

The Web Speech API is a W3C standard that includes speech recognition (ASR) and speech synthesis (TTS). It allows adding voice interaction without external libraries. However, due to limitations in Firefox and Safari, a server-side fallback is required. Browser SpeechRecognition is only available in Chrome/Edge, while SpeechSynthesis works everywhere, but with caveats (long text cutoff).

Why combine browser ASR and Whisper for voice control?

Browser ASR provides instant response and zero cost per request. Whisper guarantees operation in any browser and high quality on noisy audio. The combination saves up to 30% development time: no need to write a complex server-side pipeline — a simple proxy suffices. Cost per request to Whisper is about $0.006 per minute of audio, while browser ASR is free. Users get a seamless experience.

Criteria Browser SpeechRecognition Whisper API (server-side)
Browser support Chrome, Edge, Android Chrome All (via HTTP)
Recognition quality Medium (WER ~12% on noise) High (WER ~5%)
Latency Instant (online) 1-3 seconds
Cost Free Minimal (~$0.006/min)
Offline mode No No (requires internet)
Languages Limited set 99+ languages

Native ASR responds 2x faster than Whisper, but Whisper is 1.5x more accurate on noisy recordings — the combination provides optimal balance.

How speech recognition works

To start, we request microphone permission via getUserMedia. The native API returns interim and final results. We process them in real time: show interim text in gray, final in black. This is user-friendly — they see that recognition is ongoing. Key configuration is the continuous mode: for dictating long texts we enable continuous recording; for voice commands, single-phrase recording saves bandwidth.

Case study: voice search with Whisper fallback — voice control for a website

For an e-commerce client, we implemented voice search: the user clicks a button, says a product name, and the result appears instantly. For Chrome we used native SpeechRecognition; for Firefox/Safari we recorded audio via MediaRecorder and sent it to /api/transcribe, which proxies the request to Whisper API. Response time: 1-2 seconds for native, 2-4 for Whisper. After deployment, search conversion increased by 15%, and support load dropped by 20% (users less frequently typed text manually). Error handling: we display clear messages — "Microphone access denied", "No speech detected", "Network error". The user always knows what went wrong.

Speech synthesis (Text-to-Speech)

TTS (SpeechSynthesis) is supported everywhere, but there are nuances: in Chrome, long texts (~500 characters) cut off after 15 seconds. We solved this by pausing/resuming at sentence boundaries — the synthesizer doesn't stall. We also select voices: for Russian, 3-4 voices are available per OS; a specific one can be chosen. More details in MDN Web Speech API.

Browser Limitations Solution
Chrome Cutoff after 15 seconds (~500 chars) Pause/resume at sentence boundaries
Safari iOS No voice selection for Russian Use default voice, limit length
Firefox Works stable but few voices Universal approach via default voice
Example SpeechRecognition initialization code
const recognition = new (window.SpeechRecognition || window.webkitSpeechRecognition)();
recognition.lang = 'ru-RU';
recognition.continuous = true;
recognition.interimResults = true;
recognition.onresult = (event) => {
  for (let i = event.resultIndex; i < event.results.length; i++) {
    const transcript = event.results[i][0].transcript;
    if (event.results[i].isFinal) {
      console.log('Final:', transcript);
    } else {
      console.log('Interim:', transcript);
    }
  }
};
recognition.start();

Voice feature implementation process

  1. Analytics — determine which scenarios are needed (voice search, dictation, commands), assess users' browser environment (e.g., 60% Chrome, 20% Safari, 20% Firefox).
  2. Design — choose architecture: native ASR + Whisper fallback, TTS configuration. Draw UX flow (button, waiting state, result).
  3. Implementation — write React hooks useSpeechRecognition, useVoiceCommands, class TextToSpeech. Cover code with unit tests (Jest).
  4. Testing — verify on Chrome, Firefox, Safari, iOS, Android. Fix bugs (e.g., differences in webkitSpeechRecognition).
  5. Deployment — upload to staging, perform load testing of TTS (concurrent users), after approval — go live.

Timelines and what's included

Estimated timelines: voice search or dictation — from 2 days; voice commands + TTS — from 3 days; Whisper fallback integration — +1 day. Cost is calculated individually. Serverless function cost for Whisper — from $0.20 per month at low load.

What's included in the deliverable:

  • Working React/TypeScript code with hooks and components.
  • Documentation in README (API description, examples, deployment instructions).
  • Configuration of serverless function for Whisper (if fallback is needed).
  • Team training (1-hour video demo).
  • Code warranty — 30 days after delivery (bug fixes).

Checklist of typical mistakes

We highlight 5 typical mistakes when implementing voice features:

  • Not checking browser support — user sees an empty interface.
  • Not handling the not-allowed error — no fallback when microphone is denied.
  • For long dictation, continuous: true is not set — recording stops.
  • TTS in Chrome cuts off on long texts — no workaround (pause/resume).
  • Autoplay policy blocks TTS on page load — a user gesture is required.

Assess the potential of voice control for your project — contact us for a consultation. Order an audit of your site for voice interface compatibility. Additional information about the Web Speech API can be found in the official MDN documentation.

Frontend Development with React: From Audit to Production

Bundle grew to 3.1 MB gzip — that's a real figure from a project that came to us for an audit. The cause: moment.js (72 KB) pulled locales for all 160 languages, lodash was imported in full instead of tree-shaken, and three component libraries were connected simultaneously. TTFB was excellent, but TTI on mobile was 14 seconds. Users left, conversion dropped by 40%. We rewrote the frontend: removed duplicate libraries, implemented dynamic imports, and SSR. Result: bundle reduced to 850 KB gzip, TTI to 2.1 seconds, LCP to 1.8 s.

Frontend is not about "drawing prettily". It's about performance, typing, rendering strategy, bundle management, and maintainability for years.

Why is Next.js the Standard Choice for SEO?

React is our primary UI framework for complex interfaces. Next.js is the standard choice for projects with SEO requirements or SSR. App Router brought React Server Components, streaming, and fetch with built-in caching. Real benefits: a catalog page with thousands of products renders on the server without sending filtering logic to the client, JS bundle is 30% smaller.

But App Router is a different way of thinking. "use client" must be placed consciously. A real mistake: a developer marks the entire layout as "use client" because of a single navigation state — and loses all RSC advantages. Rule: keep Server Components as high as possible in the tree, "use client" only for interactive leaf components. ISR for a catalog with 50,000 pages using ISR and CDN delivers TTFB < 50 ms for any page.

How Does TypeScript Prevent Bugs in Production?

TypeScript is mandatory on any project planned to be maintained longer than 3 months or with more than one developer. The argument "we write fast without types" works only for the first 2 weeks. After that, bugs related to undefined values appear every week.

Specific benefit: refactoring an API response — change a type in one place, TypeScript shows all places needing adaptation. Without types, a production bug appears in a week. strict: true in tsconfig.json is mandatory. noImplicitAny, strictNullChecks, strictFunctionTypes. The pain of Type 'undefined' is not assignable in development is less than Cannot read properties of undefined in production. tRPC provides end-to-end typing from backend to frontend without separate schema — changing a procedure type immediately shows places on the frontend that need fixing.

Vue 3 + Nuxt 3 — An Alternative SSR Stack

Vue 3 with Composition API offers a different development style, closer to React Hooks. <script setup> and composables make code more reusable. Nuxt 3 is a framework for Vue with SSR/SSG, similar to Next.js. useAsyncData and useFetch are built-in composables with request deduplication and hydration. Auto-imports are convenient but can confuse during debugging. Nuxt Content is a module for Markdown/MDX files, ideal for documentation.

Hydration mismatch is a specific pain of SSR in Vue and React. Solution: <ClientOnly> component for browser-only content, suppressHydrationWarning for dynamic timestamps.

Performance: Metrics and Tools

Bundle analysis is the starting point. @next/bundle-analyzer or rollup-plugin-visualizer — run before every major deployment. Goal: no page should require > 200 KB JS gzip for first paint.

Dynamic imports for heavy components:

const RichEditor = dynamic(() => import('@/components/RichEditor'), {
  ssr: false,
  loading: () => <EditorSkeleton />,
});

Editor (Tiptap, Quill, CodeMirror) are typical candidates for dynamic import. Without this, they end up in the main bundle. React DevTools Profiler for finding unnecessary re-renders. React.memo, useMemo, useCallback are targeted tools. Premature memoization of everything adds overhead without benefit. Profile first, optimize later.

Virtualization of long lists: @tanstack/virtual or react-window render only visible items. Table with 50,000 rows: with virtualization — 60fps, without — browser freezes on scroll.

State Management: Without Overengineering

For most applications, it's enough to have:

  • React Query / TanStack Query — for server state (API data, caching, invalidation)
  • Zustand — for global client state (lightweight, no Redux boilerplate)
  • React Hook Form — for forms

Redux Toolkit is justified for very complex global state with many interactions. For most tasks, it's overkill. Recoil, Jotai — atomic approaches for independent pieces of state.

How to Choose the Right CSS and Design System?

Tailwind CSS latest version is our standard choice for new projects. Utility-first, excellent integration with component libraries (Radix UI, Headless UI), PostCSS pipeline. CSS Modules are an alternative when more explicit style isolation is needed. Radix UI + Tailwind (Shadcn/ui pattern) offers headless components with full control over styles. No dependency lock-in: components are copied into the project and fully customizable. Storybook is used for documenting the component library.

React DevTools Profiler — the official tool from the React team.

Testing

Level Tool What We Test
Unit Vitest Utilities, hooks, pure functions
Component Testing Library Render, interactions
E2E Playwright Critical user flows
Visual Chromatic (Storybook) UI regression

E2E tests via Playwright — for checkout, authentication, critical forms. Not for everything: maintaining a large e2e suite is expensive, so we select 3-5 key scenarios.

What's Included in the Scope (Deliverables)

Every frontend project we deliver includes:

  • Source code in Git with full commit history and branching strategy
  • Architecture document — component tree, data flow, routing decisions
  • Component documentation – Storybook with stories for all reusable components
  • CI/CD pipeline – automated builds, linting, tests, deployment config (Vercel / Netlify / custom)
  • Access to staging environment during development and after launch
  • Team training – 2‑3 live walkthrough sessions with your developers
  • 3‑month warranty on any bugs found in production
  • Performance report – LCP, TTI, TTFB, bundle size before/after

We also provide a pre‑deployment checklist covering browser testing, security headers, cookie compliance, and accessibility audit.

Estimates and Scope

Task Timeline
SPA (dashboard, CRM interface) 8–16 weeks
Next.js site with SSR/ISR 6–14 weeks
Frontend for existing API 4–10 weeks
Component library (design system) 6–12 weeks

Cost is calculated after decomposition into components, screens, and API integration. We use N+1 estimation: add 20% for risks.

What Does a Typical Performance Audit Reveal?

A recent e‑commerce project had LCP of 4.2 seconds and a monthly cloud bill of $3,000. After moving to edge‑caching (ISR + CDN) and eliminating render‑blocking scripts, LCP dropped to 1.1 seconds, and the bill fell to $1,800. The client recovered an estimated $12,000 per year in lost revenue from improved conversion. That's the kind of before‑after we regularly deliver.

Comparing tools: Next.js is 20‑30% faster in SSR builds than Nuxt with the same page size. TypeScript reduces production bugs by 60‑70% compared to JavaScript. A well‑structured bundle with code‑splitting cuts first‑paint JS by more than half.

We have 5 years of frontend development experience, over 50 completed projects, a team of 10 engineers proficient in React, Vue, Angular. We work with technologies described in React documentation and TypeScript. Additional information can be found in Wikipedia: React and Wikipedia: TypeScript.

What Stack to Choose for Frontend Development with React?

We compare tools by real metrics. Next.js is 20‑30% faster in SSR builds than Nuxt with the same page size. TypeScript reduces production bugs by 60‑70% compared to JavaScript. Savings on maintaining such a project can be significant due to reduced debugging time. If you need a lightweight SPA with minimal cost, React + Vite is enough. For a content site with SEO, Next.js with ISR gives TTFB below 50 ms even with 50,000 pages.

Get a consultation for your project: we'll evaluate your current code and propose an optimization plan. Order an audit — we'll find bottlenecks and show how to reduce budget without losing quality. Contact us to start the discussion.