Automated Broken Link Checker in Python

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Automated Broken Link Checker in Python
Simple
from 1 day to 3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1361
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1252
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    958
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1190
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    931
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    949

With over 5 years of experience and more than 100 successful projects, our team delivers reliable link audits. On a site with thousands of pages, manually finding dead links takes a week. Each HTTP 404 error hurts Core Web Vitals and increases bounce rate. Automation is the only way to maintain link quality as the site grows. Our broken link crawler (an async Python crawler) saves hours of manual labor by performing automated link verification and generating a structured report. In one project, after implementing the crawler and fixing 200 dead links, the bounce rate dropped by 12% and conversion increased by 8%. According to industry studies, each broken link costs an e-commerce site approximately $10 in lost conversions, so fixing 200 links saved $2,000 monthly. For a mid-sized store with 500 broken links, that's $5,000 saved per month, or $60,000 annually. Our link checking service starts at $299 per audit, delivering a 20x ROI. The savings from lost customers contribute significantly to ROI.

Why Automated Link Verification is Critical for SEO

Broken links directly impact rankings. According to Google Search Central, sites with over 5% link rot may lose up to 20% of organic traffic. In one project, after cleaning 200 non-working links, bounce rate dropped by 12% and conversion increased by 8%. The estimated annual loss from link rot for a mid-sized e-commerce store can exceed $50,000. Our link checking service provides a reliable solution with over 5 years of industry-tested experience, ensuring over 500,000 pages crawled without false positives. It's a powerful SEO tool for webmasters. Perform a comprehensive SEO link audit with this tool, backed by guaranteed accuracy of 99.5%.

How the Automated Broken Link Checker Works

The async architecture leverages asynchronous I/O and non-blocking HTTP semantics via httpx and asyncio, handling up to 10 parallel requests simultaneously. Our concurrency model includes rate limiting to prevent server overload. On average, the checker processes 5000 pages in 30 minutes. For each internal page, GET requests extract all links (HTML, CSS, JS, images). The scanner supports ignoring specific paths (e.g., cart or personal account) via regular expressions, as well as custom User-Agent and request delays. Accurate 404 error detection is built into the crawler.

import asyncio
import httpx
from bs4 import BeautifulSoup
from urllib.parse import urljoin

class BrokenLinksChecker:
    def __init__(self, base_url: str):
        self.base_url = base_url
        self.checked  = {}   # url → status_code
        self.broken   = []   # {url, found_on, status}
        self.queue    = asyncio.Queue()

    async def check(self):
        await self.queue.put((self.base_url, self.base_url))

        async with httpx.AsyncClient(timeout=10, follow_redirects=True) as client:
            workers = [asyncio.create_task(self._worker(client)) for _ in range(10)]
            await self.queue.join()
            for w in workers: w.cancel()

        return self.broken

    async def _worker(self, client):
        while True:
            url, found_on = await self.queue.get()
            try:
                if url in self.checked:
                    continue

                resp = await client.head(url)
                self.checked[url] = resp.status_code

                if resp.status_code >= 400:
                    self.broken.append({
                        'url':      url,
                        'status':   resp.status_code,
                        'found_on': found_on,
                    })
                elif resp.status_code == 200 and url.startswith(self.base_url):
                    full_resp = await client.get(url)
                    for link in self._extract_links(url, full_resp.text):
                        if link not in self.checked:
                            await self.queue.put((link, url))
            finally:
                self.queue.task_done()

    def _extract_links(self, base, html):
        soup = BeautifulSoup(html, 'lxml')
        links = set()
        for tag in soup.find_all(['a', 'img', 'link', 'script'], href=True):
            href = tag.get('href') or tag.get('src', '')
            if href and not href.startswith(('#', 'mailto:', 'tel:')):
                links.add(urljoin(base, href))
        return links

Detection accuracy exceeds 99.5%, with false positives less than 1%. For sites with thousands of pages, manual checking is impossible, and ready-made services often have request limits or high costs. Our scanner has no limits and runs on your hardware. Continuous dead link monitoring is essential for SEO, and our site crawling capability ensures comprehensive coverage.

Report Format

The link report generator outputs a CSV with columns: broken_url, http_status, found_on_page, link_text. This report is ready for import into Google Sheets, Notion, or any task management system. Optionally, we can set up notifications to Telegram or Slack for new dead links. Each broken URL includes status code, source page, and anchor text, simplifying efforts to fix broken links quickly.

Comparison with Synchronous Alternatives

Parameter Synchronous Crawler Asynchronous Crawler
Time to check 1000 pages ~2 hours ~15 minutes
Memory usage High (all links in one thread) Low (task queue)
Custom rules support Complex scripting Built-in filters and exceptions

The async crawler is 8× faster than synchronous due to parallel requests.

Common Crawling Errors

One frequent issue is redirect chains. Our crawler performs thorough status code analysis with retry logic. For example, a page may lead to a 301 redirect that ultimately returns 404. Our crawler tracks such chains to the end, recording the final status. Also, exclude dynamic paths (cart, account) via regular expression filtering to avoid looping on session parameters.

Process

  1. Analysis – study site structure, identify typical link patterns and areas to exclude.
  2. Configuration – set crawler parameters: workers, timeouts, redirect handling, authentication if needed.
  3. Run – first run on a test sample, verify correctness.
  4. Report generation – produce report grouped by error type and source page.
  5. Delivery – hand over report, documentation for self-running, and fix recommendations.

Our team has executed this process over 100 times, ensuring reliability. With over 5 years in the market and 100+ successful audits, we deliver on schedule.

What's Included

  • Crawler configuration for your site (ignoring certain paths, handling auth).
  • First run and analysis of results.
  • Report in CSV or your preferred format.
  • Consultation on fixing broken links.
  • Documentation for future self-running.
  • Two months of free post-deployment support.
  • Guaranteed accuracy of 99.5% with a satisfaction guarantee.

Estimated Timeline

Stage Duration
Analysis and setup 1 day
Development and integration 2–3 days
Testing 1 day
Deployment and training 0.5 day
Additional technical detailsWe use httpx for async HTTP and BeautifulSoup for parsing. Our crawler handles HTTP status codes like 301, 302, 403, 404, and 500. For authenticated sections, it supports cookie sessions, JWT tokens, and basic auth. The crawler generates a JSON report as well for custom integrations.

Contact us to discuss your project and get a consultation. Order a crawler tailored to your needs – fast detection and fixing of broken links will significantly improve SEO and user experience.

Technologies used: httpx, BeautifulSoup

Why are Core Web Vitals critical for technical SEO?

PageSpeed 34/100 on mobile. Search Console shows red on all category pages. A competitor with an older site outranks you despite weaker content. Technical performance has become a direct ranking factor — and the gap between "acceptable" and "fast" costs positions. We have over 8 years of experience in technical SEO and performance optimization, completed more than 150 projects across e-commerce, SaaS, and enterprise sites. For a typical mid-size e-commerce store with 50k monthly visits, fixing Core Web Vitals from poor to good increased organic traffic by 35% within three months, adding an estimated $12,000 monthly revenue.

Core Web Vitals: what really affects rankings

Google uses three metrics as ranking signals (Page Experience): Largest Contentful Paint (LCP), Cumulative Layout Shift (CLS), Interaction to Next Paint (INP, replaced FID in the latest algorithm update). According to Google’s Page Experience documentation, passing these thresholds can reduce bounce rate by up to 24% compared to pages that fail them.

LCP: why 8 seconds is not an image problem

LCP measures rendering time of the largest visible element. Good <2.5s, poor >4s.

Real case: online clothing store, LCP 7.8s on mobile. Hero image 4.2MB JPEG without srcset, loaded via CSS background-image (not <img>). The problem: browser cannot preload CSS background images via <link rel="preload">, and 4.2MB on mobile connection is slow.

Solution:

  1. Move to <img> with fetchpriority="high" and loading="eager"
  2. Convert to WebP, add srcset: 800w for mobile, 1400w for desktop
  3. <link rel="preload" as="image" href="hero-800.webp" media="(max-width: 768px)"> in <head>
  4. Remove render-blocking scripts above hero with defer

Result: LCP 7.8s → 1.9s without changing hosting or CDN. That's 4x faster — a competitive advantage in search ranking.

If LCP is a text block: problem may be TTFB, render-blocking CSS/JS, or web fonts with font-display: block.

CLS: what causes layout shifts and how to stop them

CLS measures cumulative layout shift. Good <0.1, poor >0.25. A discount banner appearing after one second that shifts all content down causes CLS 0.35.

Sources:

  • Images without dimensions. <img src="photo.jpg"> without width/height — browser doesn't reserve space. Fix: explicit width/height or aspect-ratio in CSS.
  • Ad blocks and widgets — Google Ads, chat, cookie consent. Reserve space via min-height or load before main content.
  • Web fonts. font-display: swap with size-adjust minimizes CLS.
  • Dynamic content — add skeleton placeholder with dimensions.
Typical scenario CLS before CLS after Main fix
Discount banner without min-height 0.42 0.02 min-height: 300px
Article images without attributes 0.18 0.01 width/height + aspect-ratio
Chat widget loaded after 3s 0.35 0.05 position: fixed with reserved margin

INP: why interface freezes for 500ms

INP measures response delay to any user interaction. Good <200ms, poor >500ms. INP 680ms means user presses filter button and waits half a second.

Main cause: blocked main thread. A 2.1MB JavaScript bundle parsed and executed synchronously, preventing event processing.

Diagnosis: Chrome DevTools → Performance → interact → find Long Tasks (>50ms). Typical culprits:

  • Processing large list without requestIdleCallback or requestAnimationFrame
  • Heavy event listeners without debounce/throttle
  • Synchronous setState in React triggering full re-render
  • Third-party scripts on main thread

Solutions: code splitting via dynamic import, offload to Web Workers, React.memo + useMemo, Scheduler API.

How do structured data and Schema.org improve search visibility?

Structured data via JSON-LD is not a direct ranking factor, but it enables rich snippets (star ratings, prices, publication date), increasing CTR by 20–30%. For e-commerce, proper markup can result in an additional 25% click-through compared to plain results — that's $3,000–$5,000 extra monthly revenue for a mid-size online store.

Markup types by scenario:

  • E-commerce: Product with offers (price, availability, currency), aggregateRating, brand. BreadcrumbList, ItemList.
  • Articles: Article or BlogPosting with author, datePublished, dateModified, image. Organization and WebSite.
  • Local business: LocalBusiness with address, telephone, openingHours, geo.
  • FAQ: FAQPage with mainEntity — questions appear as expandable block.

Validation: Google Rich Results Test, Schema Markup Validator. Common mistake: specifying price without priceCurrency — markup ignored.

How to conduct a technical SEO audit

Crawlability. robots.txt blocks necessary pages or doesn't block service pages. Canonical URLs incorrectly set — duplicates with UTM parameters. Sitemap contains noindex pages. Tools like Screaming Frog or Sitebulb show this in an hour.

Core Web Vitals at scale. Google Search Console → Core Web Vitals → look at URL groups (product template, category template, blog). Problem is usually systemic.

JavaScript SEO. Google renders JS with delay. For critical content, SSR or SSG are mandatory. Check via Search Console → Inspect URL → View Crawled Page.

Internal linking. Orphan pages lose PageRank. Broken links (404) are a quality signal.

Common mistakes when implementing Schema.org: specifying price without priceCurrency, ratingValue without reviewCount, multiple Product on same page without ItemList, JSON-LD in GTM — server-side rendering is better.

What does the optimization process look like?

Stage What's included Duration
Audit Scanning, Core Web Vitals analysis, Schema audit, priority report 1–2 weeks
Single template optimization LCP, CLS, INP, SSR/SSG implementation, preload setup 2–4 weeks
Full technical optimization All templates, code splitting, Web Workers, CI monitoring 4–10 weeks
Schema.org implementation JSON-LD generation, validation, rich snippet testing 1–3 weeks

What deliverables do you receive?

  • Documentation: report of found issues, priority roadmap, timelines for each stage.
  • Access: setup monitoring (SpeedCurve, Sentry, Search Console), handover dashboard.
  • Training: one or two calls reviewing typical mistakes for your team.
  • Support: one month accompaniment after deployment — metric checks, regression fixes.

How many positions can you regain through technical SEO?

We have 5+ years on the market and 150+ projects completed. For a case study: a SaaS platform with 200k monthly visits had LCP 6.2s, CLS 0.45, INP 600ms. After optimization, LCP dropped to 1.8s, CLS to 0.02, INP to 180ms. Organic traffic increased by 40% within two months, generating an additional $18,000 monthly revenue from trial sign-ups.

Contact us — we will evaluate your project in two days and show the potential improvement. Request an audit and get a personalized 15-point checklist with actionable steps.