AI Transcription with Web Interface: Development & Fine-Tuning

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
AI Transcription with Web Interface: Development & Fine-Tuning
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

You upload a meeting recording — the system instantly detects the language, launches faster-whisper on GPU, and delivers a ready transcript with speaker diarization in 5-10 minutes. But that's only half the job. Without a web interface, you can't correct errors, add annotations, or export to the required format. We build a complete solution: upload, background recognition, interactive editor, and export. Such a product can be used as an internal service or launched as a SaaS.

How AI transcription with a web interface works

Backend on FastAPI receives the file, queues the task via Celery and Redis, a worker with faster-whisper on GPU processes the audio, and the result is stored in PostgreSQL. Frontend on React polls the status and displays the transcript. The entire pipeline — from upload to export — takes as little as 5 minutes for a one-hour recording. The key component is faster-whisper, which provides a 4x inference speedup over the base Whisper.

Problems we solve: speed, accuracy, confidentiality

The first pain point — poor quality on noisy recordings or multi-speaker conversations. We use faster-whisper with noise suppression, which reduces WER (Word Error Rate) by 15–20% compared to base Whisper. The second issue — latency under growing load. The Celery and Redis architecture enables horizontal scaling of workers. The third — security: audio may contain sensitive data, so we encrypt everything at rest and in transit. faster-whisper is the key component, delivering up to 4x inference acceleration — 4 times faster than the original Whisper model.

What fine-tuning Whisper gives

The base Whisper model achieves 92–95% accuracy on general Russian speech. If your domain is medicine, law, or technical documentation, fine-tuning on your data pushes accuracy to 98%. We fine-tune the model on a dataset of 100–500 hours of labeled audio recordings. The average cost of such a dataset varies, but it typically pays off in 2–3 months due to reduced manual corrections. Our experience shows that after fine-tuning, errors in specific terms drop by 2–3 times.

How we build the system

We design the solution according to your load. Below is the stack we use in most projects:

  • Backend: FastAPI + Celery + Redis
  • Frontend: React + TypeScript + Tailwind
  • STT: faster-whisper (GPU) + cloud fallback
  • Storage: S3 (MinIO for on-premise)
  • DB: PostgreSQL

Backend API (example):

from fastapi import FastAPI, UploadFile, BackgroundTasks
from celery import Celery
import uuid

app = FastAPI()
celery = Celery('transcription', broker='redis://localhost:6379/0')

@app.post("/api/transcription/upload")
async def upload_audio(
    file: UploadFile,
    language: str = "ru",
    speakers: int = None,
    user_id: str = Depends(get_current_user)
):
    job_id = str(uuid.uuid4())
    file_path = await save_to_storage(file, job_id)
    job = await db.transcription_jobs.insert_one({
        "id": job_id,
        "user_id": user_id,
        "status": "queued",
        "file_path": file_path,
        "language": language,
        "created_at": datetime.utcnow()
    })
    celery.send_task(
        'transcribe_audio',
        args=[job_id, file_path, language, speakers]
    )
    return {"job_id": job_id, "status": "queued"}

@app.get("/api/transcription/{job_id}")
async def get_transcription(job_id: str, user_id = Depends(get_current_user)):
    job = await db.transcription_jobs.find_one({"id": job_id, "user_id": user_id})
    if not job:
        raise HTTPException(404)
    return job

React upload component:

const TranscriptionUploader: React.FC = () => {
  const [status, setStatus] = useState<'idle'|'uploading'|'processing'|'done'>('idle');
  const [jobId, setJobId] = useState<string>();
  const [transcript, setTranscript] = useState<string>();

  const handleUpload = async (file: File) => {
    setStatus('uploading');
    const form = new FormData();
    form.append('file', file);
    form.append('language', 'ru');

    const { job_id } = await api.post('/transcription/upload', form);
    setJobId(job_id);
    setStatus('processing');

    const interval = setInterval(async () => {
      const job = await api.get(`/transcription/${job_id}`);
      if (job.status === 'completed') {
        setTranscript(job.transcript);
        setStatus('done');
        clearInterval(interval);
      }
    }, 3000);
  };

  return (
    <div>
      <FileDropzone onFile={handleUpload} accept="audio/*,video/*" />
      {status === 'processing' && <ProgressSpinner jobId={jobId} />}
      {transcript && <TranscriptEditor text={transcript} jobId={jobId} />}
    </div>
  );
};

Comparison: self-hosted vs cloud STT

Parameter Self-hosted (faster-whisper) Cloud (e.g., Speech-to-Text)
Price per 1 hour of audio significantly cheaper (electricity only) more expensive (from $2)
Confidentiality full control data leaves the server
Latency p99 <10 sec 50–200 ms + network
Scaling limited by hardware elastic
Custom model yes (fine-tuning) no

The self-hosted option pays off at volumes from 500 hours per month: savings reach 90–95%. At 1000 hours per month, cloud STT would cost $2000, while self-hosted only the cost of electricity and GPU depreciation (T4 or A10). This saves up to $1900 per month. Guarantee of stable operation — our implementation experience in 12 projects.

Transcript editor and export

In the editor you can correct words, reassign speakers, add annotations. Highlighting low-confidence words (confidence < 0.7) speeds up review. Export formats:

Format Use case
SRT Subtitles for video
VTT Web subtitles (HTML5)
DOCX Documentation, reports
JSON Integration with CRM

Project stages

  1. Requirements analysis — gather scenarios, accuracy requirements, volumes, security needs.
  2. Architecture design — select optimal stack for your infrastructure.
  3. Backend development — implement upload API, task queue, processing via faster-whisper.
  4. Frontend development — upload interface, status bar, transcript editor.
  5. Integration and testing — connect CRM, perform load testing.
  6. Deployment and support — deploy on servers, documentation, team training.

Deliverables: what is included in the result

  • Architecture diagram
  • Repository with backend and frontend
  • CI/CD and infrastructure setup (Docker Compose / Kubernetes)
  • Integration with corporate portal or CRM
  • Technical documentation and user manual
  • Team training (1–2 sessions)
  • 2 weeks of post-launch support

Estimated timelines

  • MVP with upload and basic interface — 2–3 weeks
  • Full system with editor, team features, and fine-tuning — 1.5–2 months

How to avoid common mistakes

  • Poor audio quality (noise, overlap) — solved with denoising preprocessing.
  • High latency — optimize batch processing and GPU utilization.
  • Missing speaker diarization — apply voice embedding clustering.
  • Data leakage risk — encryption and on-premise deployment.

Contact us for a project assessment. Order a custom transcription system development. Get a consultation today.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.