Introduction
Podcasters spend hours on manual transcription and shownotes preparation. An average hour-long episode contains about 10,000 words of text. Even with modern ASR systems, Word Error Rate (WER) can reach 20% on multi-speaker recordings. We use Whisper large-v3 from OpenAI: a model with 1,550 million parameters trained on 680,000 hours of multilingual data. It reduces WER to 4–8% on clean studio recordings, and after fine-tuning — to 3–5%. Combined with GPT-4o, we get ready-made shownotes with timestamps in 5–10 minutes.
How Whisper large-v3 handles noise?
Whisper large-v3 outperforms previous versions thanks to an encoder-decoder architecture with attention over 128 token context. On noisy recordings — street noise, echo, cross-dialogues — the model is more robust due to training on synthetic noises. For specific accents or radio interference, we apply fine-tuning: we adapt the model on 1–2 hours of your data using LoRA adapters. This boosts accuracy by 10–15% without retraining the entire model.
Why automate summarization?
Manual writing of shownotes for a single podcast can take 2–3 hours. GPT-4o with a proper chain-of-thought prompt does it in 30 seconds, extracting up to 10 key topics and generating a brief description. Cost savings on editing — up to 80% compared to hiring a copywriter. Quality is not compromised: the model accounts for timestamps and thematic transitions.
Comparison of transcription models
| Model | WER (clean audio) | Speed (1 hour on GPU) | Features |
|---|---|---|---|
| Whisper large-v3 | 4–8% | 3–4 min | Best accuracy, open-source |
| Google Speech-to-Text | 10–15% | 2–3 min | Good GCP integration |
| Wav2Vec 2.0 | 12–18% | 1–2 min | Requires language fine-tuning |
Whisper large-v3 is twice as accurate as Wav2Vec 2.0 in WER and processes audio up to 12 hours without context loss. Unlike Google API, the model can be deployed locally — full data control and privacy.
Detailed processing pipeline
- Upload audio file or RSS feed. For RSS, monitoring is configured to poll the feed every 6 hours.
- Preprocessing: loudness normalization (LUFS -16) and spectral noise reduction via the noisereduce library.
- Transcription with Whisper large-v3: language="ru", word_timestamps=True.
- Speaker diarization via pyannote-audio: separation into voices, alignment with segments.
- Shownotes generation via GPT-4o with a prompt containing transcript (up to 6000 tokens) and timestamps.
- Forming an RSS feed with new items and publishing via your CMS API.
import whisper
from openai import AsyncOpenAI
async def transcribe_and_summarize_podcast(audio_path: str) -> dict:
# Transcription
model = whisper.load_model("large-v3")
result = model.transcribe(
audio_path,
language="ru",
task="transcribe",
verbose=False,
word_timestamps=True
)
transcript = result["text"]
segments = result["segments"] # [{start, end, text}, ...]
# Generate shownotes via GPT-4o
client = AsyncOpenAI()
response = await client.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "system",
"content": "Create shownotes for a podcast: a brief episode description (3-5 sentences), key topics as a list, timestamps for main topics in MM:SS format."
}, {
"role": "user",
"content": transcript[:6000]
}]
)
# Timestamps for key topics
chapters = extract_chapters(segments)
return {
"transcript": transcript,
"shownotes": response.choices[0].message.content,
"chapters": chapters,
"duration_sec": segments[-1]["end"] if segments else 0
}
def extract_chapters(segments: list) -> list[dict]:
"""Extract thematic blocks by pauses and semantics"""
chapters = []
# Look for pauses > 3 seconds as chapter boundaries
for i in range(1, len(segments)):
gap = segments[i]["start"] - segments[i-1]["end"]
if gap > 3.0:
chapters.append({
"timestamp": int(segments[i]["start"]),
"text": segments[i]["text"][:80]
})
return chapters
RSS feed integration
For podcasts with regular releases, we set up RSS monitoring. A new episode is automatically downloaded, transcribed, and shownotes are published on the site.
import feedparser
import httpx
async def process_podcast_feed(rss_url: str) -> list[dict]:
feed = feedparser.parse(rss_url)
results = []
for entry in feed.entries[:5]: # last 5 episodes
audio_url = next(
(enc.href for enc in entry.enclosures if enc.type.startswith("audio")),
None
)
if not audio_url:
continue
async with httpx.AsyncClient() as client:
audio_data = await client.get(audio_url)
with open(f"/tmp/{entry.id}.mp3", "wb") as f:
f.write(audio_data.content)
result = await transcribe_and_summarize_podcast(f"/tmp/{entry.id}.mp3")
result["title"] = entry.title
result["published"] = entry.published
results.append(result)
return results
What you get?
Full transcription and summarization pipeline, production-ready. Includes: content analysis, selection of optimal model (Whisper large-v3 or fine-tuned version), diarization setup, integration with your site via RSS or API, documentation in a repository, team training. Post-launch support — 1 month with a guaranteed stable WER below 10% after adaptation. Team experience — over 7 years in NLP and 50+ completed audio processing projects.
Typical mistakes and how to avoid them
Low recording quality is the main cause of high WER. Use studio microphones and avoid reverberation. For long episodes (over 2 hours), GPT-4o context window is limited to 128K tokens, so we split audio into 30-minute parts with a 5-second overlap for stitching. The chapter extraction algorithm based on pauses requires calibration: we adjust the silence threshold to your speech tempo — from 2 to 4 seconds.
Timeline and cost
Development of a typical pipeline takes 1 to 4 weeks. Cost is calculated individually after analyzing your recordings — accounting for duration, release frequency, and required integrations. Get a consultation: contact us for a free project assessment.
We guarantee stable performance and accuracy. We will assess your project and propose the optimal architecture — write to us, let's discuss the details.







