You launched an online store, competitors' prices change daily, and manual monitoring eats up hours of your managers' time. Without automated collection, you lose profit: you fail to react to a competitor's price drop or miss new products in their assortment. A competitor catalog parser is a tool that daily collects current prices, availability, and characteristics into your database. No more manually checking sites: the system automatically crawls the catalog, records changes, and sends alerts. Our experience — over 10 years in developing such solutions, dozens of successful end-to-end projects.
Why manual collection is ineffective?
Manual monitoring of three competitors with 500 products each takes 2–3 hours per day. Errors, omissions, outdated data. An automated parser solves these problems: collects data in minutes, works 24/7, never gets tired. Time savings — up to 90% compared to manual collection. Pays for itself in 2–3 months.
Site analysis before development
Before writing code — analysis of the target site:
- Catalog URL structure: pagination via
?page=N, infinite scroll, or tree navigation by categories - Rendering: static HTML (fast and simple) or data loaded via XHR/fetch (needs interception or headless)
- Protection: Cloudflare, rate limiting, authorization
- Data update frequency on the site — how quickly new products appear and prices change
| Site type | Parsing difficulty | Collection speed (1000 products) | Reliability |
|---|---|---|---|
| Static HTML | Low | 1–2 minutes | High |
| SPA with XHR (API) | Medium | 3–5 minutes | Very high |
| SPA without API (Client-side render) | High | 5–10 minutes | High (with proper delays) |
Typical minimum field set: SKU / article, title, price (regular + sale), availability, category, product page URL, scraping date. For some niches, important fields include: rating, number of reviews, weight/dimensions, brand.
Technical implementation
For static sites — httpx + parsel (or Cheerio for Node.js). Async requests, connection pool of 10–20 workers, delay of 1–3 seconds between requests to the same domain.
import httpx
import asyncio
import random
from parsel import Selector
UA_POOL = [
'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36',
]
async def fetch_page(session: httpx.AsyncClient, url: str) -> str:
headers = {
'User-Agent': random.choice(UA_POOL),
'Accept-Language': 'ru-RU,ru;q=0.9',
}
resp = await session.get(url, headers=headers, timeout=15)
resp.raise_for_status()
return resp.text
async def parse_catalog_page(html: str, base_url: str) -> list[dict]:
sel = Selector(html)
products = []
for item in sel.css('.product-card'):
price_raw = item.css('.price::text').get('').strip()
price = int(''.join(c for c in price_raw if c.isdigit())) if price_raw else None
products.append({
'title': item.css('.product-title::text').get('').strip(),
'price': price,
'sku': item.attrib.get('data-sku'),
'url': base_url + item.css('a::attr(href)').get(''),
'in_stock': bool(item.css('.in-stock')),
'image_url': item.css('img::attr(src)').get(),
})
return products
For SPA with XHR — intercept API requests via Playwright. Many modern online stores, when opening a page, make a fetch request to their own API that returns JSON with product data:
from playwright.async_api import async_playwright
import json
async def intercept_catalog_api(catalog_url: str) -> list[dict]:
products = []
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
async def handle_response(response):
if '/api/catalog' in response.url and response.status == 200:
try:
data = await response.json()
if 'products' in data:
products.extend(data['products'])
except Exception:
pass
page.on('response', handle_response)
await page.goto(catalog_url, wait_until='networkidle')
await browser.close()
return products
If the API returns JSON directly — we can call it directly bypassing the browser, which is 10–20 times faster. To find the endpoint — use DevTools Network tab while manually browsing the catalog.
How does SPA with XHR parsing work?
In an SPA, the main challenge is not the HTML but the API requests that load data. We intercept these requests via Playwright and get clean JSON. This is more reliable than parsing dynamically generated DOM. If the API is open — we call it directly, saving resources.
Pagination and full crawl
For pagination via ?page=N — sequential crawl until an empty page:
async def scrape_full_catalog(base_url: str) -> list[dict]:
all_products = []
page_num = 1
async with httpx.AsyncClient() as session:
while True:
url = f'{base_url}?page={page_num}'
html = await fetch_page(session, url)
products = await parse_catalog_page(html, base_url)
if not products:
break
all_products.extend(products)
page_num += 1
await asyncio.sleep(random.uniform(1.5, 3.0)) # polite delay
return all_products
For category tree — first recursively collect all category URLs, then crawl each category with pagination.
Storage and incremental updates
CREATE TABLE competitor_products (
id SERIAL PRIMARY KEY,
source VARCHAR(100) NOT NULL, -- 'competitor_a', 'competitor_b'
external_id VARCHAR(255) NOT NULL,
title TEXT NOT NULL,
price DECIMAL(10,2),
price_sale DECIMAL(10,2),
in_stock BOOLEAN DEFAULT TRUE,
category VARCHAR(500),
url TEXT NOT NULL,
image_url TEXT,
attributes JSONB DEFAULT '{}',
first_seen TIMESTAMPTZ DEFAULT NOW(),
last_seen TIMESTAMPTZ DEFAULT NOW(),
UNIQUE(source, external_id)
);
CREATE TABLE competitor_price_history (
id BIGSERIAL PRIMARY KEY,
product_id INT REFERENCES competitor_products(id),
price DECIMAL(10,2),
price_sale DECIMAL(10,2),
in_stock BOOLEAN,
scraped_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE INDEX ON competitor_price_history(product_id, scraped_at DESC);
On subsequent crawls — INSERT ... ON CONFLICT (source, external_id) DO UPDATE SET last_seen = NOW(), price = EXCLUDED.price, .... History entry is made only if price or availability changed (compare with last entry via LAG() or store price in the main table).
Scheduling and alerts
Celery Beat or Node.js cron. Recommended frequency for a competitor's catalog — every 4–12 hours, depending on price dynamics in the niche. For marketplaces with fast-changing prices — every hour for top positions.
Alert when a competitor's price drops below yours — SQL query or PostgreSQL trigger with notification to Slack/Telegram via webhook. Example query:
SELECT cp.title, cp.price AS competitor_price, mp.price AS my_price
FROM competitor_products cp
JOIN my_products mp ON mp.sku = cp.external_id
WHERE cp.source = 'competitor_a'
AND cp.price < mp.price
AND cp.in_stock = TRUE
ORDER BY (mp.price - cp.price) DESC;
How to set up alerts for competitor price drops?
- Set a threshold:
SELECT ... WHERE cp.price < mp.price * 0.95— alert on a 5% drop. - Configure a webhook in Telegram/Slack.
- Run the SQL query after each crawl and send the result.
We implement this logic as part of the parser: you receive a notification in messenger with a table of products where the competitor became cheaper.
How to ensure uninterrupted parser operation?
Competitors' sites change — the parser periodically breaks. We set up monitoring: alert if in the last run less than 50% of the average product count is collected. When the structure changes, updating usually takes 2–4 hours. We guarantee support and adaptation to new site versions.
What's included in the work?
- Exhaustive analysis of the target site (structure, protection, API)
- Development of the parser with pagination, categories, incremental updates
- Database setup for storing price and assortment history
- Scheduling (cron) and alert configuration (Telegram/Slack)
- Documentation for operation and access
- Training of your staff to use the system
- Warranty support for 1 month and response to failures within 2–4 hours
We will evaluate your project — contact us, we will offer the optimal end-to-end solution. Order parser development and get a tool that will bring real benefits in competitive struggle. Wikipedia: Web scraping







