Review collection and import from marketplaces: parser development
We build parsers for automatic review collection from marketplaces. A client lost 30% conversion due to outdated reviews — manual collection took hours, data went unrefreshed for weeks. Our team (5+ years experience, 50+ projects) created a solution that collects reviews from Wildberries, Ozon, Yandex.Market, and other platforms, normalizes them, and imports them into your store's database. Time savings — up to 70%. Project payback — 2–3 months. Order parser development for your online store.
Typical scenario: a manager manually copies reviews from 5 marketplaces for 200 products — that's 8 hours a day. Data grows stale, customers see no new reviews for weeks, trust drops. A parser grabs reviews every 2 hours, updating product pages in real time. Result: conversion growth by 15% in the first month.
Technically, each marketplace is its own headache. Wildberries offers a JSON API, but with a limit of 1000 reviews per session. Ozon is an SPA on Nuxt, where data is loaded via GraphQL — we have to emulate a browser with Playwright. Yandex.Market changes its structure every six months, so we use adaptive parsing with a fallback strategy. Our stack: Python for high-load platforms, PHP (Laravel) for integration with your site.
Supported marketplaces and methods
| Platform | Method | Notes |
|---|---|---|
| Wildberries | JSON API | Open API, pagination, up to 1000 reviews per session |
| Ozon | Playwright | SPA, needs authorization, 50x slower than JSON |
| Yandex.Market | Unofficial API | Rate limiting, requires proxy rotation |
| Google Reviews | Places API | Paid, official, up to 5 reviews per request (requires business subscription) |
| Otzovik.com | HTML parsing | CAPTCHA on mass requests — we use solvers |
| iHerb | HTML / JSON API | Structured HTML, easiest |
Method comparison: JSON-API (Wildberries) processes 1000 reviews in 10 seconds, while Selenium on Ozon takes 5 minutes for the same 1000. That's a 30 times difference. For performance, we choose JSON wherever possible.
How the Wildberries parser works
We use asynchronous httpx to iterate through pagination. Example:
Click to expand Python code example
# scraper/reviews/wildberries.py
import httpx
import asyncio
from dataclasses import dataclass
from typing import Optional
@dataclass
class Review:
external_id: str
product_nm_id: int
author: str
rating: int
text: str
pros: Optional[str]
cons: Optional[str]
date: str
photos: list[str]
helpful_count: int
class WildberriesReviewScraper:
REVIEWS_URL = "https://feedbacks2.wb.ru/feedbacks/v2/{nm_id}"
def __init__(self):
self.client = httpx.AsyncClient(
headers={
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)",
"Origin": "https://www.wildberries.ru",
"Referer": "https://www.wildberries.ru/",
}
)
async def get_reviews(self, nm_id: int, take: int = 100) -> list[Review]:
all_reviews = []
skip = 0
while True:
url = self.REVIEWS_URL.format(nm_id=nm_id)
params = {
"immt": nm_id,
"skip": skip,
"take": take,
"order": "dateDesc",
}
resp = await self.client.get(url, params=params)
resp.raise_for_status()
data = resp.json()
feedbacks = data.get("feedbacks", [])
if not feedbacks:
break
for fb in feedbacks:
all_reviews.append(self._normalize(nm_id, fb))
skip += take
await asyncio.sleep(1.0)
# Limit: no more than 1000 reviews per session
if skip >= 1000:
break
return all_reviews
def _normalize(self, nm_id: int, raw: dict) -> Review:
photos = []
for photo in raw.get("photos", []):
if full_url := photo.get("fullSize"):
photos.append(full_url)
return Review(
external_id=raw.get("id", ""),
product_nm_id=nm_id,
author=raw.get("wbUserDetails", {}).get("name", "Customer"),
rating=raw.get("productValuation", 0),
text=raw.get("text", ""),
pros=raw.get("pros"),
cons=raw.get("cons"),
date=raw.get("createdDate", ""),
photos=photos,
helpful_count=raw.get("feedbackValuation", 0),
)
Parsing HTML reviews
For sites without an API, we parse HTML using PHP and Symfony Crawler. Example for iHerb:
// app/Services/ReviewScraper/HtmlReviewScraper.php
class HtmlReviewScraper
{
public function scrapeIherb(string $productUrl, int $pages = 5): array
{
$reviews = [];
for ($page = 1; $page <= $pages; $page++) {
$html = $this->fetch("{$productUrl}?p={$page}&is=1&s=6");
$crawler = new Crawler($html);
$items = $crawler->filter('[itemprop="review"]');
if (!$items->count()) break;
$items->each(function (Crawler $node) use (&$reviews) {
$reviews[] = [
'external_id' => $node->attr('data-review-id'),
'author' => trim($node->filter('[itemprop="author"]')->text('')),
'rating' => (int) $node->filter('[itemprop="ratingValue"]')->attr('content'),
'date' => $node->filter('[itemprop="datePublished"]')->attr('content'),
'title' => trim($node->filter('[itemprop="name"]')->text('')),
'text' => trim($node->filter('[itemprop="reviewBody"]')->text('')),
'helpful' => (int) $node->filter('.helpful-yes')->text('0'),
'verified' => $node->filter('.verified-buyer')->count() > 0,
];
});
sleep(rand(2, 4));
}
return $reviews;
}
}
Deduplication and import into Laravel
After collection, reviews go through a Job with deduplication by external_id and source. Example:
// app/Jobs/ImportProductReviews.php
class ImportProductReviews implements ShouldQueue
{
public int $tries = 3;
public int $backoff = 120;
public function handle(ReviewImportService $service): void
{
$mapping = ProductReviewMapping::where('product_id', $this->productId)
->where('source', $this->source)
->firstOrFail();
$reviews = $this->scrape($mapping->external_id);
$imported = 0;
$skipped = 0;
foreach ($reviews as $reviewData) {
$exists = ProductReview::where([
'source' => $this->source,
'external_id' => $reviewData['external_id'],
])->exists();
if ($exists) {
$skipped++;
continue;
}
$service->import($this->productId, $this->source, $reviewData);
$imported++;
}
Log::info("Reviews imported", [
'product_id' => $this->productId,
'source' => $this->source,
'imported' => $imported,
'skipped' => $skipped,
]);
}
}
Filtering and moderation
Stop-words (spam, ads, profanity), too-short reviews (<20 characters), and author anonymization — all customizable. Below is a service example:
// app/Services/ReviewImportService.php
class ReviewImportService
{
private array $stopWords = ['buy', 'discount', 'promocode', 'vk.com', 't.me'];
public function import(int $productId, string $source, array $data): ?ProductReview
{
if (mb_strlen($data['text']) < 20) return null;
foreach ($this->stopWords as $word) {
if (mb_stripos($data['text'], $word) !== false) return null;
}
return ProductReview::create([
'product_id' => $productId,
'source' => $source,
'external_id' => $data['external_id'],
'author' => $this->anonymizeAuthor($data['author']),
'rating' => max(1, min(5, (int) $data['rating'])),
'text' => $this->sanitize($data['text']),
'pros' => $this->sanitize($data['pros'] ?? ''),
'cons' => $this->sanitize($data['cons'] ?? ''),
'date' => $data['date'],
'is_verified' => $data['verified'] ?? false,
'helpful' => $data['helpful'] ?? 0,
'status' => 'pending',
]);
}
private function anonymizeAuthor(string $name): string
{
$parts = explode(' ', trim($name));
if (count($parts) >= 2) {
return $parts[0] . ' ' . mb_substr($parts[1], 0, 1) . '.';
}
return $name ?: 'Customer';
}
private function sanitize(string $text): string
{
return strip_tags(trim($text));
}
}
Structured data for SEO
After import, reviews are published in JSON-LD on the product page. This increases the chance of appearing in rich snippets and boosts CTR by 20-30%. Learn more about structured data.
// app/Http/Controllers/ProductController.php
public function show(string $slug): Response
{
$product = Product::withReviews()->findBySlug($slug);
$reviewSchema = $product->reviews->map(fn($r) => [
'@type' => 'Review',
'author' => ['@type' => 'Person', 'name' => $r->author],
'datePublished' => $r->date,
'reviewBody' => $r->text,
'reviewRating' => [
'@type' => 'Rating',
'ratingValue' => $r->rating,
'bestRating' => 5,
],
]);
$aggregateRating = [
'@type' => 'AggregateRating',
'ratingValue' => round($product->reviews->avg('rating'), 1),
'reviewCount' => $product->reviews->count(),
];
}
How to avoid blocks while scraping?
To avoid blocking, we use a pool of 50+ residential proxies, random delays from 1 to 4 seconds, rotate User-Agent (Chrome, Firefox, Safari). CAPTCHAs are solved via Anti-Captcha. Each session is limited to 1000 reviews.
Why review automation boosts SEO?
- Structured data (JSON-LD) helps search engines display ratings in snippets.
- UGC texts contain long-tail keywords.
- Regular updates signal an active store.
Comparison: manual collection of 100 reviews takes 2 hours, our parser — 2 minutes, 60 times faster. Budget savings on copywriters and moderators — up to 70%.
Work process
- Analysis: determine platforms, APIs, complexity.
- Development: write parser, moderation, deduplication.
- Testing: run on 1000+ reviews, check for bugs.
- Deployment: set up cron, error monitoring.
- Documentation: stack description, startup instructions.
| Stage | Duration |
|---|---|
| Requirements analysis | 1-2 days |
| Parser development | 2-3 days per platform |
| Testing | 1-2 days |
| Deployment and documentation | 1 day |
Starting price for a single platform parser is $500, with volume discounts available. We guarantee reliable performance, with over 5 years of experience and certified scraping practices. Our review parser bot handles review aggregation from multiple sources, ensuring all data is imported correctly.
What's included
- Parser for each marketplace (up to 5 platforms).
- Moderation (stop-words, length, anonymization).
- Deduplication (by external_id + source).
- Automatic updates on schedule (daily or every 4 hours).
- Structured data on product page.
- Logging and error notifications.
Timeline: from 3 to 5 business days per platform. For a complex project (3 platforms + moderation + SEO) — 7-10 days. Contact us to evaluate your project. Get a consultation on your project.







